Generating gene summaries from biomedical literature: A study of semi-structured summarization

被引:28
作者
Ling, Xu [1 ]
Jiang, Jing [1 ]
He, Xin [1 ]
Mei, Qiaozhu [1 ]
Zhai, Chengxiang [1 ]
Schatz, Bruce [1 ]
机构
[1] Univ Illinois, Dept Comp Sci, Inst Genom Biol, Urbana, IL 61801 USA
基金
美国国家科学基金会;
关键词
summarization; Genomics; Probabilistic language model;
D O I
10.1016/j.ipm.2007.01.018
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Most knowledge accumulated through scientific discoveries in genomics and related biomedical disciplines is buried in the vast amount of biomedical literature. Since understanding gene regulations is fundamental to biomedical research, summarizing all the existing knowledge about a gene based on literature is highly desirable to help biologists digest the literature. In this paper, we present a study of methods for automatically generating gene summaries from biomedical literature. Unlike most existing work on automatic text summarization, in which the generated summary is often a list of extracted sentences, we propose to generate a semi-structured summary which consists of sentences covering specific semantic aspects of a gene. Such a semi-structured summary is more appropriate for describing genes and poses special challenges for automatic text summarization. We propose a two-stage approach to generate such a summary for a given gene - first retrieving articles about a gene and then extracting sentences for each specified semantic aspect. We address the issue of gene name variation in the first stage and propose several different methods for sentence extraction in the second stage. We evaluate the proposed methods using a test set with 20 genes. Experiment results show that the proposed methods can generate useful semi-structured gene summaries automatically from biomedical literature, and our proposed methods outperform general purpose summarization methods. Among all the proposed methods for sentence extraction, a probabilistic language modeling approach that models gene context performs the best. (C) 2007 Elsevier Ltd. All rights reserved.
引用
收藏
页码:1777 / 1791
页数:15
相关论文
共 21 条
[1]  
[Anonymous], 2003, P 12 TEXT RETRIEVAL
[2]   FlyBase: genes and gene models [J].
Drysdale, RA ;
Crosby, MA .
NUCLEIC ACIDS RESEARCH, 2005, 33 :D390-D395
[3]   Summarizing text documents: Sentence selection and evaluation metrics [J].
Goldstein, J ;
Kantrowitz, M ;
Mittal, V ;
Carbonell, J .
SIGIR'99: PROCEEDINGS OF 22ND INTERNATIONAL CONFERENCE ON RESEARCH AND DEVELOPMENT IN INFORMATION RETRIEVAL, 1999, :121-128
[4]  
HEARST MA, 1996, P 19 ANN INT ACM SIG, P76
[5]   Overview of BioCreAtIvE task IB: normalized gene lists [J].
Hirschman, L ;
Colosimo, M ;
Morgan, A ;
Yeh, A .
BMC BIOINFORMATICS, 2005, 6 (Suppl 1)
[6]   Accomplishments and challenges in literature data mining for biology [J].
Hirschman, L ;
Park, JC ;
Tsujii, J ;
Wong, L ;
Wu, CH .
BIOINFORMATICS, 2002, 18 (12) :1553-1561
[7]  
Iliopoulos I, 2001, Pac Symp Biocomput, P384
[8]  
KRAAIJ W, 2001, P DUC2001 WORKSH NEW
[9]  
Kummamuru Krishna, 2004, P 13 INT C WORLD WID, P658
[10]  
Kupiec J., 1995, SIGIR FOR ACM SPEC I, P68, DOI DOI 10.1145/215206.215333