Monaural speech separation and recognition challenge

被引：139

作者：

Cooke, Martin ^{[1
,2
]}

Hershey, John R. ^{[3
]}

Rennie, Steven J. ^{[3
]}

机构：

[1] Univ Basque Country, Dept Elect & Elect, Fac Ciencias & Tecnol, Leioa 48940, Spain

[2] Ikerbasque Basque Sci Fdn, Bilbao 48011, Bizkaia, Spain

[3] IBM Corp, Thomas J Watson Res Ctr, Yorktown Hts, NY 10598 USA

来源：

COMPUTER SPEECH AND LANGUAGE | 2010年 / 24卷 / 01期

关键词：

Speech recognition; Speech separation; Speaker identification; Simultaneous speech; Auditory scene analysis; Noise robustness; ROBUST; MASKING;

D O I：

10.1016/j.csl.2009.02.006

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Robust speech recognition in everyday conditions requires the solution to a number of challenging problems, not least the ability to handle multiple sound sources. The specific case of speech recognition in the presence of a competing talker has been studied for several decades, resulting in a number of quite distinct algorithmic solutions whose focus ranges from modeling both target and competing speech to speech separation using auditory grouping principles. The purpose of the monaural speech separation and recognition challenge was to permit a large-scale comparison of techniques for the competing talker problem. The task was to identify keywords in sentences spoken by a target talker when mixed into a single channel with a background talker speaking similar sentences. Ten independent sets of results were contributed, alongside a baseline recognition system. Performance was evaluated using common training and test data and common metrics. Listeners' performance in the same task was also measured. This paper describes the challenge problem, compares the performance of the contributed algorithms, and discusses the factors which distinguish the systems. One highlight of the comparison was the finding that several systems achieved near-human performance in some conditions, and one out-performed listeners overall. (C) 2009 Elsevier Ltd. All rights reserved.

引用

页码：1 / 15

页数：15

共 53 条

[1]

[Anonymous], SPRINGER HDB AUDITOR

[2]

[Anonymous], 2005, Speech Enhancement

[3]

[Anonymous], HTK BOOK HTK V2 0

[4]

[Anonymous], 2007, Speech Enhancement: Theory and Practice