Loading…
Emotions, speech and the ASR framework
Automatic recognition & understanding of speech are crucial steps towards natural human-machine interaction. Apart from the recognition of the word sequence, the recognition of properties such as prosody, emotion tags, or stress tags may be of particular importance in this communication process....
Saved in:
Published in: | Speech communication 2003-04, Vol.40 (1-2), p.213-225 |
---|---|
Main Author: | |
Format: | Article |
Language: | English |
Citations: | Items that cite this one |
Online Access: | Get full text |
Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Summary: | Automatic recognition & understanding of speech are crucial steps towards natural human-machine interaction. Apart from the recognition of the word sequence, the recognition of properties such as prosody, emotion tags, or stress tags may be of particular importance in this communication process. This paper discusses the possibilities to recognize emotion from the speech signal, primarily from the viewpoint of automatic speech recognition (ASR). The general focus is on the extraction of acoustic features from the speech signal that can be used for the detection of the emotional state or stress state of the speaker. After the introduction, a short overview of the ASR framework is presented. Next, we discuss the relation between recognition of emotion & ASR, & the different approaches found in the literature that deal with the correspondence between emotions & acoustic features. The conclusion is that automatic emotional tagging of the speech signal is difficult to perform with high accuracy, but prosodic information is nevertheless potentially useful to improve the dialogue handling in ASR tasks on a limited domain. 58 References. Adapted from the source document |
---|---|
ISSN: | 0167-6393 |
DOI: | 10.1016/S0167-6393(02)00083-3 |