Robust speech recognition using the modulation spectrogram

被引:157
作者
Kingsbury, BED
Morgan, N
Greenberg, S
机构
[1] Int Comp Sci Inst, Berkeley, CA 94704 USA
[2] Univ Calif Berkeley, Berkeley, CA 94720 USA
关键词
robust speech recognition; reverberation;
D O I
10.1016/S0167-6393(98)00032-6
中图分类号
O42 [声学];
学科分类号
070206 ; 082403 ;
摘要
The performance of present-day automatic speech recognition (ASR) systems is seriously compromised by levels of acoustic interference (such as additive noise and room reverberation) representative of real-world speaking conditions. Studies on the perception of speech by human listeners suggest that recognizer robustness might be improved by focusing on temporal structure in the speech signal that appears as low-frequency (below 16 Hz) amplitude modulations in subband channels following critical-band frequency analysis. A speech representation that emphasizes this temporal structure, the "modulation spectrogram", has been developed. Visual displays of speech produced with the modulation spectrogram are relatively stable in the presence of high levels of background noise and reverberation. Using the modulation spectrogram as a front end for ASR provides a significant improvement in performance on highly reverberant speech. When the modulation spectrogram is used in combination with log-RASTA-PLP (log RelAtive SpecTrAl Perceptual Linear Predictive analysis) performance over a range of noisy and reverberant conditions is significantly improved, suggesting that the use of multiple representations is another promising method for improving the robustness of ASR systems. (C) 1998 Elsevier Science B.V. All rights reserved.
引用
收藏
页码:117 / 132
页数:16
相关论文
共 32 条
[1]  
ARAI T, 1998, P 1998 IEEE INT C AC
[2]  
BOURLARD H, 1994, CONNECTIONIST SPEECH, P155
[3]   EFFECT OF TEMPORAL ENVELOPE SMEARING ON SPEECH RECEPTION [J].
DRULLMAN, R ;
FESTEN, JM ;
PLOMP, R .
JOURNAL OF THE ACOUSTICAL SOCIETY OF AMERICA, 1994, 95 (02) :1053-1064
[4]   Remaking speech [J].
Dudley, H .
JOURNAL OF THE ACOUSTICAL SOCIETY OF AMERICA, 1939, 11 (02) :169-177
[5]  
FURUI S, 1986, P IEEE INT C AC SPEE, P1991
[6]  
GREENBERG S, 1997, P IEEE INT C AC SPEE, P1647
[7]  
GREENBERG S, 1996, P 4 INT C SPOK LANG, pS24
[8]  
GREENBERG S, 1998, P JOINT M AC SOC AM
[9]  
Greenberg S., 1997, P ESCA WORKSH ROB SP, P23
[10]   CRITICAL BANDWIDTH AND FREQUENCY COORDINATES OF BASILAR MEMBRANE [J].
GREENWOOD, D .
JOURNAL OF THE ACOUSTICAL SOCIETY OF AMERICA, 1961, 33 (10) :1344-&