MeSH: a window into full text for document summarization

被引：29

作者：

Bhattacharya, Sanmitra ^{[1
]}

Viet Ha-Thuc ^{[1
]}

Srinivasan, Padmini ^{[1
,2
]}

机构：

[1] Univ Iowa, Dept Comp Sci, Iowa City, IA 52242 USA

[2] Univ Iowa, Dept Management Sci, Iowa City, IA 52242 USA

来源：

BIOINFORMATICS | 2011年 / 27卷 / 13期

基金：

美国国家科学基金会;

关键词：

BIOMEDICAL LITERATURE; CLASSIFICATION; MEDLINE;

D O I：

10.1093/bioinformatics/btr223

中图分类号：

Q5 [生物化学];

学科分类号：

071010 ; 081704 ;

摘要：

Motivation: Previous research in the biomedical text-mining domain has historically been limited to titles, abstracts and metadata available in MEDLINE records. Recent research initiatives such as TREC Genomics and BioCreAtIvE strongly point to the merits of moving beyond abstracts and into the realm of full texts. Full texts are, however, more expensive to process not only in terms of resources needed but also in terms of accuracy. Since full texts contain embellishments that elaborate, contextualize, contrast, supplement, etc., there is greater risk for false positives. Motivated by this, we explore an approach that offers a compromise between the extremes of abstracts and full texts. Specifically, we create reduced versions of full text documents that contain only important portions. In the long-term, our goal is to explore the use of such summaries for functions such as document retrieval and information extraction. Here, we focus on designing summarization strategies. In particular, we explore the use of MeSH terms, manually assigned to documents by trained annotators, as clues to select important text segments from the full text documents. Results: Our experiments confirm the ability of our approach to pick the important text portions. Using the ROUGE measures for evaluation, we were able to achieve maximum ROUGE-1, ROUGE-2 and ROUGE-SU4 F-scores of 0.4150, 0.1435 and 0.1782, respectively, for our MeSH term-based method versus the maximum baseline scores of 0.3815, 0.1353 and 0.1428, respectively. Using a MeSH profile-based strategy, we were able to achieve maximum ROUGE F-scores of 0.4320, 0.1497 and 0.1887, respectively. Human evaluation of the baselines and our proposed strategies further corroborates the ability of our method to select important sentences from the full texts.

引用

页码：I120 / I128

页数：9

共 35 条

[1]

Agarwal Shashank, 2009, AMIA Annu Symp Proc, V2009, P6

[2]

[Anonymous], 1971, The SMART Retrieval System-Experiments in Automatic Document Processing

[3]

[Anonymous], 2003, P 2003 C N AM CHAPT

[4]

[Anonymous], 2005, P JOENS LEARN INSTR

[5]

[Anonymous], P BIOCR 3 WORKSH BET

[6]

Aone C, 1999, ADVANCES IN AUTOMATIC TEXT SUMMARIZATION, P71

[7]

Bhattacharya S., 2010, Proceedings of the BioCreative III workshop, P55

[8] AUTOMATIC CONDENSATION OF ELECTRONIC PUBLICATIONS BY SENTENCE SELECTION [J].

BRANDOW, R ;

MITZE, K ;

RAU, LF .

INFORMATION PROCESSING & MANAGEMENT, 1995, 31 (05) :675-685

[9] GeneLibrarian: an effective gene-information summarization and visualization system [J].

Chiang, Jung-Hsien ;

Shin, Jyh-Wei ;

Liu, Heng-Hui ;

Chin, Chong-Liang .

BMC BIOINFORMATICS, 2006, 7 (1)

[10] Five-way smoking status classification using text hot-spot identification and error-correcting output codes [J].

Cohen, Aaron M. .

JOURNAL OF THE AMERICAN MEDICAL INFORMATICS ASSOCIATION, 2008, 15 (01) :32-35

← 1 2 3 4 →