Toward a model for lexical access based on acoustic landmarks and distinctive features

被引:326
作者
Stevens, KN [1 ]
机构
[1] MIT, Elect Res Lab, Cambridge, MA 02139 USA
[2] MIT, Dept Elect Engn & Comp Sci, Cambridge, MA 02139 USA
关键词
D O I
10.1121/1.1458026
中图分类号
O42 [声学];
学科分类号
070206 ; 082403 ;
摘要
This article describes a model in which the acoustic speech signal is processed to yield a discrete representation of the speech stream in terms of a sequence of segments, each of which is described by a set (or bundle) of binary distinctive features. These distinctive features specify the phonemic contrasts that are used in the language, such that a change in the value of a feature can potentially generate a new word. This model is a part of a more general model that derives a word sequence from this feature representation, the words being represented in a lexicon by sequences of feature bundles. The processing of the signal proceeds in three steps: (1) Detection of peaks, valleys, and discontinuities in particular frequency ranges of the signal leads to identification of acoustic landmarks. The type of landmark provides evidence for a subset of distinctive features called articulator-free features (e.g., [vowel], [consonant], [continuant]). (2) Acoustic parameters are derived from the signal near the landmarks to provide evidence for the actions of particular articulators, and acoustic cues are extracted by sampling selected attributes of these parameters in these regions. The selection of cues that are extracted depends on the type of landmark and on the environment in which it occurs. (3) The cues obtained in step (2) are combined, taking context into account, to provide estimates of "articulator-bound" features associated with each landmark (e.g., [Lips], [high], [nasal]). These articulator-bound features, combined with the articulator-free features in (1), constitute the sequence of feature bundles that forms the output of the model, Examples of cues that are used, and justification for this selection, are given, as well as examples of the process of inferring the underlying features for a segment when there is variability in the signal due to enhancement gestures (recruited by a speaker to make a contrast more salient) or due to overlap of gestures from neighboring segments. (C) 2002 Acoustical Society of America.
引用
收藏
页码:1872 / 1891
页数:20
相关论文
共 69 条
[1]  
[Anonymous], 2000, THESIS MIT CAMBRIDGE
[2]  
[Anonymous], 1989, SPEAKING
[3]  
Browman C.P., 1991, Papers in Laboratory Phonology I: Between the Grammar and the Physics of Speech, P341, DOI DOI 10.1017/CBO9780511627736.019
[4]   Acoustic correlates of English and French nasalized vowels [J].
Chen, MY .
JOURNAL OF THE ACOUSTICAL SOCIETY OF AMERICA, 1997, 102 (04) :2360-2370
[5]  
CHEN MY, 2000, P 6 INT C SPOK LANG, V4, P636
[6]   CENTER OF GRAVITY EFFECT IN VOWEL SPECTRA AND CRITICAL DISTANCE BETWEEN THE FORMANTS - PSYCHO-ACOUSTICAL STUDY OF THE PERCEPTION OF VOWEL-LIKE STIMULI [J].
CHISTOVICH, LA ;
LUBLINSKAYA, VV .
HEARING RESEARCH, 1979, 1 (03) :185-195
[7]   Variation and universals in VOT: evidence from 18 languages [J].
Cho, T ;
Ladefoged, P .
JOURNAL OF PHONETICS, 1999, 27 (02) :207-229
[8]  
Choi J. Y., 1999, THESIS MIT CAMBRIDGE
[9]  
Chomsky Noam., 1968, The sound pattern of English
[10]  
Clements G.N., 1983, CV PHONOLOGY