Feature analysis and neural network-based classification of speech under stress

被引：66

作者：

Hansen, JHL

Womack, BD

机构：

[1] Robust Speech Processing Laboratory, Department of Electrical Engineering, Duke University, Durham

来源：

IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING | 1996年 / 4卷 / 04期

关键词：

D O I：

10.1109/89.506935

中图分类号：

O42 [声学];

学科分类号：

070206 [声学]; 082403 [水声工程];

摘要：

It is well known that the variability in speech production due to task-induced stress contributes significantly to loss in speech processing algorithm performance. If an algorithm could be formulated that detects the presence of stress in speech, then such knowledge could be used to monitor speaker state, improve the naturalness of speech coding algorithms, or increase the robustness of speech recognizers. The goal in this study is to consider several speech features as potential stress-sensitive relayers using a previously established stressed speech database (SUSAS). The following speech parameters will be considered: mel, delta-mel, delta-delta-mel, auto-correlation-mel, and cross-correlation-mel cepstral parameters, Next, an algorithm for speaker-dependent stress classification is formulated for the 11 stress conditions: Angry, Clear, Cond50, Cond70, Fast, Lombard, Loud, Normal, Question, Slow, and Soft, It is suggested that additional feature variations beyond neutral conditions reflect the perturbation of vocal tract articulator movement under stressed conditions. Given a robust set of features, a neural network-based classifier is formulated based on an extended delta-bar-delta learning rule. Performance is considered for the following three test scenarios: monopartition (nontargeted) and tripartition (both nontargeted and targeted) input feature vectors.

引用

页码：307 / 313

页数：7

共 11 条

[1]

NONLINEAR-ANALYSIS AND CLASSIFICATION OF SPEECH UNDER STRESSED CONDITIONS [J].