Ad hoc classification of radiology reports

被引:45
作者
Aronow, DB [1 ]
Feng, FF [1 ]
Croft, WB [1 ]
机构
[1] Univ Massachusetts, Ctr Intelligent Informat Retrieval, Amherst, MA 01003 USA
关键词
D O I
10.1136/jamia.1999.0060393
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Objective: The task of ad hoc classification is to automatically place a large number of text documents into nonstandard categories that are determined by a user. The authors examine the use of statistical information retrieval techniques for ad hoc classification of dictated mammography reports. Design: The authors' approach is the automated generation of a classification algorithm based on positive and negative evidence that is extracted from relevance-judged documents. Test documents are sorted into three conceptual bins: membership in a user-defined class, exclusion from the user-defined class, and uncertain. Documentation of absent findings through the use of negation and conjunction, a hallmark of interpretive test results, is managed by expansion and tokenization of these phrases. Measurements: Classifier performance is evaluated using a single measure, the F measure, which provides a weighted combination of recall and precision of document sorting into true positive and true negative bins. Results: Single terms are the most effective text feature in the classification profile, with some improvement provided by the addition of pairs of unordered terms to the profile. Excessive iterations of automated classifier enhancement degrade performance because of overtraining. Performance is best when the proportions of relevant and irrelevant documents in the training collection are close to equal. Special handling of negation phrases improves performance when the number of terms in the classification profile is Limited. Conclusions: The ad hoc classifier system is a promising approach for the classification of large collections of medical documents. NegExpander tan distinguish positive evidence from negative evidence when the negative evidence plays an important role in the classification.
引用
收藏
页码:393 / 411
页数:19
相关论文
共 31 条
  • [1] ALLAN J, 1996, P 5 TEXT RETR C GAIT
  • [2] [Anonymous], 1995, P 4 TREC
  • [3] Aronow D B, 1995, Proc Annu Symp Comput Appl Med Care, P309
  • [4] ARONOW DB, 1995, P 8 WORLD C MED INF, P8
  • [5] ARONOW DB, 1997, DIGITAL LIBRARY JAN
  • [6] BUCKLEY C, 1994, ACM SIGIR 17 INT C, P292
  • [7] Callan J. P., 1992, DEXA 92. Database and Expert Systems Applications. Proceedings of the International Conference, P78
  • [8] CHARNIAK E, 1991, AI MAG, V12, P50
  • [9] An experiment comparing lexical and statistical methods for extracting MeSH terms from clinical free text
    Cooper, GF
    Miller, RA
    [J]. JOURNAL OF THE AMERICAN MEDICAL INFORMATICS ASSOCIATION, 1998, 5 (01) : 62 - 75
  • [10] DEESTRADA WD, 1997, 1997 AMIA ANN FALL S, P509