Standardizing formats of corporate source data

被引:25
作者
Galvez, Carmen [1 ]
Moya-Anegon, Felix [1 ]
机构
[1] Univ Granada, Dept Informat Sci, Scimago Res Grp, E-18071 Granada, Spain
关键词
D O I
10.1007/s11192-007-0101-0
中图分类号
TP39 [计算机的应用];
学科分类号
081203 ; 0835 ;
摘要
This paper describe an approach for improving the data quality of corporate sources when databases are used for bibliometric purposes. Research management relies on bibliographic databases and citation index systems as analytical tools, yet the raw resources for bibliometric studies are plagued by a lack of consistency in fied formatting for institution data. The present contribution puts forth a Natural Language Processing (NLP)-oriented method for the identification of the structures guiding corporate data and their mapping into a standardized format. The proposed unification process is based on the definition of address patterns and the ensuing application of Enhanced Finite-State Transducers (E-FST). Our procedure was tested on address formats downloaded from the INSPEC, MEDLINE and CAB Abstracts. The results demonstrate the helpfulness of the method as long as close control of errors is exercised as far as the formats to be unified. The computational efficacy of the model is noteworthy, due to the fact that it is firmly guided by the definition of data in the application domain.
引用
收藏
页码:3 / 26
页数:24
相关论文
共 57 条
[1]  
ABNEY S, 2000, P 40 ANN M ASS COMP
[2]  
ABNEY S, 1996, P ESSLLI 96 ROB PARS, P8
[3]  
ATRIN P, 2003, P 4 DUTCH BELG INF R, P16
[4]   Institutions and the map of science: matching university departments and fields of research [J].
Bourke, P ;
Butler, L .
RESEARCH POLICY, 1998, 26 (06) :711-718
[5]   Standards issues in a national bibliometric database: The Australian case [J].
Bourke, P ;
Butler, L .
SCIENTOMETRICS, 1996, 35 (02) :199-207
[6]   HYPHENATION OF DATABASES IN BUILDING SCIENTOMETRIC INDICATORS - PHYSICS BRIEFS - SCI BASED INDICATORS OF 13 EUROPEAN COUNTRIES, 1980-1989 [J].
BRAUN, T ;
BROCKEN, M ;
GLANZEL, W ;
RINIA, E ;
SCHUBERT, A .
SCIENTOMETRICS, 1995, 33 (02) :131-148
[7]   BIBLIOMETRIC PROFILES FOR BRITISH ACADEMIC-INSTITUTIONS - AN EXPERIMENT TO DEVELOP RESEARCH OUTPUT INDICATORS [J].
CARPENTER, MP ;
GIBB, F ;
HARRIS, M ;
IRVINE, J ;
MARTIN, BR ;
NARIN, F .
SCIENTOMETRICS, 1988, 14 (3-4) :213-233
[8]   Special issue on data quality in cooperative information systems - Editorial [J].
Catarci, T .
INFORMATION SYSTEMS, 2004, 29 (07) :529-530
[9]  
Chomsky N., 1957, SYNTACTIC STRUCTURES, DOI 10.1515/9783112316009
[10]  
Chomsky N., 1965, Aspects of the Theory of Syntax