Chinese word segmentation and named entity recognition: A pragmatic approach

被引:191
作者
Gao, JF [1 ]
Li, M [1 ]
Wu, A [1 ]
Huang, CN [1 ]
机构
[1] Microsoft Res Asia, Nat Language Comp Grp, Sigma Ctr, Beijing 100080, Peoples R China
关键词
D O I
10.1162/089120105775299177
中图分类号
TP18 [人工智能理论];
学科分类号
081104 [模式识别与智能系统]; 0812 [计算机科学与技术]; 0835 [软件工程]; 1405 [智能科学与技术];
摘要
This article presents a pragmatic approach to Chinese word segmentation. It differs from most previous approaches mainly in three respects. First, while theoretical linguists have defined Chinese words using various linguistic criteria, Chinese words in this study are defined pragmatically as segmentation units whose definition depends on how they are used and processed in realistic computer applications. Second, we propose a pragmatic mathematical framework in which segmenting known words and detecting unknown words of different types (i. e., morphologically derived words, factoids, named entities, and other unlisted words) can be performed simultaneously in a unified way. These tasks are usually conducted separately in other systems. Finally, we do not assume the existence of a universal word segmentation standard that is application-independent. Instead, we argue for the necessity of multiple segmentation standards due to the pragmatic fact that different natural language processing applications might require different granularities of Chinese words. These pragmatic approaches have been implemented in an adaptive Chinese word segmenter, called MSRSeg, which will be described in detail. It consists of two components: (1) a generic segmenter that is based on the framework of linear mixture models and provides a unified approach to the five fundamental features of word-level Chinese language processing: lexicon word processing, morphological analysis, factoid detection, named entity recognition, and new word identification; and (2) a set of output adaptors for adapting the output of (1) to different application-specific standards. Evaluation on five test sets with different standards shows that the adaptive system achieves state-of-the-art performance on all the test sets.
引用
收藏
页码:531 / 574
页数:44
相关论文
共 70 条
[1]
Aho Alfred V., 1986, ADDISON WESLEY SERIE
[2]
[Anonymous], COMPUTATIONAL LINGUI
[3]
[Anonymous], 1987, J CHINESE INFORM PRO
[4]
Baayen R.H., 1989, THESIS FREE U AMSTER
[5]
The psychology of reactions to environmental agents [J].
Berglund, B ;
Job, RFS .
ENVIRONMENT INTERNATIONAL, 1996, 22 (01) :1-1
[6]
Brill E, 1995, COMPUT LINGUIST, V21, P543
[7]
CHANG E, 2001, EUROSPEECH 2001, P2799
[8]
CHANG JS, 1997, INT J COMPUTATIONAL, V2, P97
[9]
CHEN A, 2003, 2 SIGHAN WORKSH CHIN
[10]
Chen K.J., 1998, INT J COMPUTATIONAL, V3, P27