A dictionary to identify small molecules and drugs in free text

被引:124
作者
Hettne, Kristina M. [1 ,2 ,3 ]
Stierum, Rob H. [3 ,4 ]
Schuemie, Martijn J. [2 ]
Hendriksen, Peter J. M. [5 ]
Schijvenaars, Bob J. A. [6 ]
van Mulligen, Erik M. [2 ]
Kleinjans, Jos [1 ,3 ]
Kors, Jan A. [2 ]
机构
[1] Maastricht Univ, Dept Hlth Risk Anal & Toxicol, Maastricht, Netherlands
[2] Erasmus Univ, Dept Med Informat, Med Ctr, NL-3000 DR Rotterdam, Netherlands
[3] Netherlands Toxicogenom Ctr, Dept Toxicoinformat, Maastricht, Netherlands
[4] TNO Qual Life, Business Unit Biosci, Physiol Genom, Zeist, Netherlands
[5] RIKILT Inst Food Safety, Wageningen, Netherlands
[6] Collexis Holdings Inc, Columbia, SC USA
关键词
CHEMISTRY; DATABASE; GENE; INFORMATION; SYSTEM; ABBREVIATIONS; KNOWLEDGEBASE; EXTRACTION; INTERNET; PROTEIN;
D O I
10.1093/bioinformatics/btp535
中图分类号
Q5 [生物化学];
学科分类号
070307 [化学生物学];
摘要
Motivation: From the scientific community, a lot of effort has been spent on the correct identification of gene and protein names in text, while less effort has been spent on the correct identification of chemical names. Dictionary-based term identification has the power to recognize the diverse representation of chemical information in the literature and map the chemicals to their database identifiers. Results: We developed a dictionary for the identification of small molecules and drugs in text, combining information from UMLS, MeSH, ChEBI, DrugBank, KEGG, HMDB and ChemIDplus. Rule-based term filtering, manual check of highly frequent terms and disambiguation rules were applied. We tested the combined dictionary and the dictionaries derived from the individual resources on an annotated corpus, and conclude the following: (i) each of the different processing steps increase precision with a minor loss of recall; (ii) the overall performance of the combined dictionary is acceptable (precision 0.67, recall 0.40 (0.80 for trivial names); (iii) the combined dictionary performed better than the dictionary in the chemical recognizer OSCAR3; (iv) the performance of a dictionary based on ChemIDplus alone is comparable to the performance of the combined dictionary.
引用
收藏
页码:2983 / 2991
页数:9
相关论文
共 56 条
[1]
Literature mining in support of drug discovery [J].
Agarwal, Pankaj ;
Searls, David B. .
BRIEFINGS IN BIOINFORMATICS, 2008, 9 (06) :479-492
[2]
Agirre E, 2006, TEXT SPEECH LANG TEC, V33, P1, DOI 10.1007/1-4020-4809-2_1
[3]
Biomedical word sense disambiguation with ontologies and metadata: automation meets accuracy [J].
Alexopoulou, Dimitra ;
Andreopoulos, Bill ;
Dietze, Heiko ;
Doms, Andreas ;
Gandon, Fabien ;
Hakenberg, Joerg ;
Khelif, Khaled ;
Schroeder, Michael ;
Waechter, Thomas .
BMC BIOINFORMATICS, 2009, 10
[4]
Mining chemical structural information from the drug literature [J].
Banville, DL .
DRUG DISCOVERY TODAY, 2006, 11 (1-2) :35-42
[5]
BINGJUN S, 2007, P 16 INT C WORLD WID
[6]
BINGJUN S, 2008, P 17 INT C WORLD WID
[7]
The Unified Medical Language System (UMLS): integrating biomedical terminology [J].
Bodenreider, O .
NUCLEIC ACIDS RESEARCH, 2004, 32 :D267-D270
[8]
ChemDB update - full-text search and virtual chemical space [J].
Chen, Jonathan H. ;
Linstead, Erik ;
Swamidass, S. Joshua ;
Wang, Dennis ;
Baldi, Pierre .
BIOINFORMATICS, 2007, 23 (17) :2348-2351
[9]
A survey of current work in biomedical text mining [J].
Cohen, AM ;
Hersh, WR .
BRIEFINGS IN BIOINFORMATICS, 2005, 6 (01) :57-71
[10]
Corbett P., 2007, P WORKSHOP BIONLP 20, P57