Loading…

Speaker cluster-based speaker adaptive training for deep neural network acoustic modeling

A speaker cluster-based speaker adaptive training (SAT) method under deep neural network-hidden Markov model (DNN-HMM) framework is presented in this paper. During training, speakers that are acoustically adjacent to each other are hierarchically clustered using an i-vector based distance metric. DN...

Full description

Saved in:

Bibliographic Details
Main Authors:	Chu, Wei, Chen, Ruxin
Format:	Conference Proceeding
Language:	English
Subjects:	Adaptation models Clusters Conferences Decoding Deep Neural Network Electronics Hidden Markov models i-vector Neural networks Silicon Speaker Adaptive Training Speaker Clustering Speech Speech recognition Tasks Training Transforms
Online Access:	Request full text
Tags:	Add Tag No Tags, Be the first to tag this record!

Description
Summary:	A speaker cluster-based speaker adaptive training (SAT) method under deep neural network-hidden Markov model (DNN-HMM) framework is presented in this paper. During training, speakers that are acoustically adjacent to each other are hierarchically clustered using an i-vector based distance metric. DNNs with speaker dependent layers are then adaptively trained for each cluster of speakers. Before decoding starts, an unseen speaker in test set is matched to the closest speaker cluster through comparing i-vector based distances. The previously trained DNN of the matched speaker cluster is used for decoding utterances of the test speaker. The performance of the proposed method on a large vocabulary spontaneous speech recognition task is evaluated on a training set of 1500 hours of speech, and a test set of 24 speakers with 1774 utterances. Compared to a speaker independent DNN with a word error rate of 11.6%, a relative 6.8% improvement in performance is obtained from the proposed method.
ISSN:	2379-190X
DOI:	10.1109/ICASSP.2016.7472688