Can bibliographic pointers for known biological data be found automatically? Protein interactions as a case study

被引:35
作者
Blaschke, C [1 ]
Valencia, A [1 ]
机构
[1] CSIC, CNB, Natl Biotechnol Ctr, Prot Design Grp, E-28049 Madrid, Spain
来源
COMPARATIVE AND FUNCTIONAL GENOMICS | 2001年 / 2卷 / 04期
关键词
proteomics; information extraction; protein interactions; yeast two-hybrid; database;
D O I
10.1002/cfg.91
中图分类号
Q5 [生物化学]; Q7 [分子生物学];
学科分类号
071010 ; 081704 ;
摘要
The Dictionary of Interacting Proteins (DIP) (Xenarios et al, 2000) is a large repository of protein interactions- its March 2000 release included 2379 protein pairs whose interactions have been detected by experimental methods. Even if many of these correspond to poorly characterized proteins, the result of massive yeast two-hybrid screenings, as many as 851 correspond to interactions detected using direct biochemical methods. We used information retrieval technology to search automatically for sentences in Medline abstracts that support these 851 DIP interactions. Surprisingly, we found correspondence between DIP protein pairs and Medline sentences describing their interactions in only 30% of the cases. This low coverage has interesting consequences regarding the quality of annotations (references) introduced in the database and the limitations of the application of information extraction (IE) technology to Molecular Biology. It is clear that the limitation of analyzing abstracts rather than full papers and the lack of standard protein names are difficulties of considerably more importance than the limitations of the IE methodology employed. A positive finding is the capacity of the IE system to identify new relations between proteins, even in a set of proteins previously characterized by human experts. These identifications are made with a considerable degree of precision. This is, to our knowledge, the first large scale assessment of IE capacity to detect previously known interactions: we thus propose the use of the DIP data set as a biological reference to benchmark IE systems. Copyright (C) 2001 John Wiley & Sons, Ltd.
引用
收藏
页码:196 / 206
页数:11
相关论文
共 34 条
[1]   Automatic extraction of keywords from scientific text: application to the knowledge domain of protein families [J].
Andrade, MA ;
Valencia, A .
BIOINFORMATICS, 1998, 14 (07) :600-607
[2]  
[Anonymous], 1998, GENOME INFORM
[3]   BIND - The Biomolecular Interaction Network Database [J].
Bader, GD ;
Donaldson, I ;
Wolting, C ;
Ouellette, BFF ;
Pawson, T ;
Hogue, CWV .
NUCLEIC ACIDS RESEARCH, 2001, 29 (01) :242-245
[4]   The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000 [J].
Bairoch, A ;
Apweiler, R .
NUCLEIC ACIDS RESEARCH, 2000, 28 (01) :45-48
[5]  
BAKER WC, 2000, NUCLEIC ACIDS RES, V28, P45
[6]   GenBank [J].
Benson, DA ;
Boguski, MS ;
Lipman, DJ ;
Ostell, J ;
Ouellette, BFF .
NUCLEIC ACIDS RESEARCH, 1998, 26 (01) :1-7
[7]  
Blaschke C, 1999, Proc Int Conf Intell Syst Mol Biol, P60
[8]   Mining functional information associated with expression arrays [J].
Blaschke C. ;
Oliveros J.C. ;
Valencia A. .
Functional & Integrative Genomics, 2001, 1 (4) :256-268
[9]   THE 2-HYBRID SYSTEM - A METHOD TO IDENTIFY AND CLONE GENES FOR PROTEINS THAT INTERACT WITH A PROTEIN OF INTEREST [J].
CHIEN, CT ;
BARTEL, PL ;
STERNGLANZ, R ;
FIELDS, S .
PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES OF THE UNITED STATES OF AMERICA, 1991, 88 (21) :9578-9582
[10]  
Eilbeck K, 1999, Proc Int Conf Intell Syst Mol Biol, P87