Comparing intermittency and network measurements of words and their dependence on authorship

被引:38
作者
Amancio, Diego Raphael [2 ]
Altmann, Eduardo G. [1 ]
Oliveira, Osvaldo N., Jr. [2 ]
Costa, Luciano da Fontoura [2 ]
机构
[1] Max Planck Inst Phys Komplexer Syst, Dresden, Germany
[2] Univ Sao Paulo, Inst Phys Sao Carlos, BR-13560970 Sao Paulo, Brazil
来源
NEW JOURNAL OF PHYSICS | 2011年 / 13卷
基金
巴西圣保罗研究基金会;
关键词
COMPLEX NETWORKS; KEYWORD DETECTION; LEAST EFFORT; SMALL-WORLD; LANGUAGE; CLASSIFICATION; DISTRIBUTIONS;
D O I
10.1088/1367-2630/13/12/123024
中图分类号
O4 [物理学];
学科分类号
0702 ;
摘要
Many features of texts and languages can now be inferred from statistical analyses using concepts from complex networks and dynamical systems. In this paper, we quantify how topological properties of word co-occurrence networks and intermittency (or burstiness) in word distribution depend on the style of authors. Our database contains 40 books by eight authors who lived in the nineteenth and twentieth centuries, for which the following network measurements were obtained: the clustering coefficient, average shortest path lengths and betweenness. We found that the two factors with stronger dependence on authors were skewness in the distribution of word intermittency and the average shortest paths. Other factors such as betweenness and Zipf's law exponent show only weak dependence on authorship. Also assessed was the contribution from each measurement to authorship recognition using three machine learning methods. The best performance was about 65% accuracy upon combining complex networks and intermittency features with the nearest-neighbor algorithm of automatic authorship. From a detailed analysis of the interdependence of the various metrics, it is concluded that the methods used here are complementary for providing short- and long-scale perspectives on texts, which are useful for applications such as the identification of topical words and information retrieval.
引用
收藏
页数:17
相关论文
共 52 条
[41]  
Quinlan J.R., 1993, C4 5 PROGRAMS MACHIN
[42]  
RATNAPARKI A, 1996, P EMP METH NAT LANG
[43]   Prose and Poetry Classification and Boundary Detection Using Word Adjacency Network Analysis [J].
Roxas, Ranzivelle Marianne ;
Tapang, Giovanni .
INTERNATIONAL JOURNAL OF MODERN PHYSICS C, 2010, 21 (04) :503-512
[44]  
SHANNON CE, 1948, BELL SYST TECH J, V27, P379, DOI DOI 10.1002/J.1538-7305.1948.TB01338.X
[45]   Language Networks: Their Structure, Function, and Evolution [J].
Sole, Ricard V. ;
Corominas-Murtra, Bernat ;
Valverde, Sergi ;
Steels, Luc .
COMPLEXITY, 2010, 15 (06) :20-26
[46]  
Stevanak J T, 2010, ARXIV10073254
[47]  
TANKARD WJ, 2001, APPL COMPUTER CONTEN, pCH4
[48]  
UZUNER O, 2005, SIGIR WORKSH STYL AN
[49]   Can analysis of word frequency distinguish between writings of different authors? [J].
Vilensky, B .
PHYSICA A-STATISTICAL MECHANICS AND ITS APPLICATIONS, 1996, 231 (04) :705-711
[50]  
Witten IH, 2011, MOR KAUF D, P1