Towards graphical models for text processing

被引:25
作者
Aggarwal, Charu C. [1 ]
Zhao, Peixiang [2 ]
机构
[1] IBM Corp, Thomas J Watson Res Ctr, Yorktown Hts, NY 10598 USA
[2] Florida State Univ, Tallahassee, FL 32306 USA
关键词
Text clustering; Text classification; Text representation; Text search;
D O I
10.1007/s10115-012-0552-3
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
The rapid proliferation of the World Wide Web has increased the importance and prevalence of text as a medium for dissemination of information. A variety of text mining and management algorithms have been developed in recent years such as clustering, classification, indexing, and similarity search. Almost all these applications use the well-known vector-space model for text representation and analysis. While the vector-space model has proven itself to be an effective and efficient representation for mining purposes, it does not preserve information about the ordering of the words in the representation. In this paper, we will introduce the concept of distance graph representations of text data. Such representations preserve information about the relative ordering and distance between the words in the graphs and provide a much richer representation in terms of sentence structure of the underlying data. Recent advances in graph mining and hardware capabilities of modern computers enable us to process more complex representations of text. We will see that such an approach has clear advantages from a qualitative perspective. This approach enables knowledge discovery from text which is not possible with the use of a pure vector-space representation, because it loses much less information about the ordering of the underlying words. Furthermore, this representation does not require the development of new mining and management techniques. This is because the technique can also be converted into a structural version of the vector-space representation, which allows the use of all existing tools for text. In addition, existing techniques for graph and XML data can be directly leveraged with this new representation. Thus, a much wider spectrum of algorithms is available for processing this representation. We will apply this technique to a variety of mining and management applications and show its advantages and richness in exploring the structure of the underlying text documents.
引用
收藏
页码:1 / 21
页数:21
相关论文
共 33 条
[1]  
Aggarwal C, 2012, EDBT C, P348
[2]  
Aggarwal CC, 2010, ADV DATABASE SYST, V40, P1, DOI 10.1007/978-1-4419-6045-0
[3]  
Aggarwal CC, 2010, SIGIR 2010: PROCEEDINGS OF THE 33RD ANNUAL INTERNATIONAL ACM SIGIR CONFERENCE ON RESEARCH DEVELOPMENT IN INFORMATION RETRIEVAL, P899
[4]   On clustering massive text and categorical data streams [J].
Aggarwal, Charu C. ;
Yu, Philip S. .
KNOWLEDGE AND INFORMATION SYSTEMS, 2010, 24 (02) :171-196
[5]  
Aggarwal CC, 2007, KDD-2007 PROCEEDINGS OF THE THIRTEENTH ACM SIGKDD INTERNATIONAL CONFERENCE ON KNOWLEDGE DISCOVERY AND DATA MINING, P46
[6]  
[Anonymous], 1992, COLING 1992, DOI DOI 10.3115/992133.992154
[7]  
[Anonymous], 2008, Proceedings of the 2008 ACM SIGMOD international conference on Management of data
[8]  
[Anonymous], P INT C MACH LEARN
[9]  
[Anonymous], 1996, Bow: A toolkit for statistical language modeling, text retrieval, classification and clustering
[10]  
[Anonymous], 2012, MINING TEXT DATA