Identification of literary movements using complex networks to represent texts

被引:34
作者
Amancio, Diego Raphael [1 ]
Oliveira, Osvaldo N., Jr. [1 ]
Costa, Luciano da Fontoura [1 ]
机构
[1] Univ Sao Paulo, Inst Phys Sao Carlos, BR-13560970 Sao Carlos, SP, Brazil
基金
巴西圣保罗研究基金会;
关键词
Graph theory - Behavioral research;
D O I
10.1088/1367-2630/14/4/043029
中图分类号
O4 [物理学];
学科分类号
0702 ;
摘要
The use of statistical methods to analyze large databases of text has been useful in unveiling patterns of human behavior and establishing historical links between cultures and languages. In this study, we identified literary movements by treating books published from 1590 to 1922 as complex networks, whose metrics were analyzed with multivariate techniques to generate six clusters of books. The latter correspond to time periods coinciding with relevant literary movements over the last five centuries. The most important factor contributing to the distinctions between different literary styles was the average shortest path length, in particular the asymmetry of its distribution. Furthermore, over time there has emerged a trend toward larger average shortest path lengths, which is correlated with increased syntactic complexity, and a more uniform use of the words reflected in a smaller power-law coefficient for the distribution of word frequency. Changes in literary style were also found to be driven by opposition to earlier writing styles, as revealed by the analysis performed with geometrical concepts. The approaches adopted here are generic and may be extended to analyze a number of features of languages and cultures.
引用
收藏
页数:15
相关论文
共 32 条
[1]   Using metrics from complex networks to evaluate machine translation [J].
Amancio, D. R. ;
Nunes, M. G. V. ;
Oliveira, O. N., Jr. ;
Pardo, T. A. S. ;
Antiqueira, L. ;
Costa, L. da F. .
PHYSICA A-STATISTICAL MECHANICS AND ITS APPLICATIONS, 2011, 390 (01) :131-142
[2]   Complex networks analysis of manual and machine translations [J].
Amancio, Diego R. ;
Antiqueira, Lucas ;
Pardo, Thiago A. S. ;
Costa, Luciano da F. ;
Oliveira, Osvaldo N., Jr. ;
Nunes, Maria G. V. .
INTERNATIONAL JOURNAL OF MODERN PHYSICS C, 2008, 19 (04) :583-598
[3]   Comparing intermittency and network measurements of words and their dependence on authorship [J].
Amancio, Diego Raphael ;
Altmann, Eduardo G. ;
Oliveira, Osvaldo N., Jr. ;
Costa, Luciano da Fontoura .
NEW JOURNAL OF PHYSICS, 2011, 13
[4]  
[Anonymous], 1949, Human behaviour and the principle of least-effort
[5]  
[Anonymous], 2010, Networks: An Introduction, DOI 10.1162/artl_r_00062
[6]   Strong correlations between text quality and complex networks features [J].
Antiqueira, L. ;
Nunes, M. G. V. ;
Oliveira, O. N., Jr. ;
Costa, L. da F. .
PHYSICA A-STATISTICAL MECHANICS AND ITS APPLICATIONS, 2007, 373 :811-820
[7]   A complex network approach to text summarization [J].
Antiqueira, Lucas ;
Oliveira, Osvaldo N., Jr. ;
Costa, Luciano da Fontoura ;
Volpe Nunes, Maria das Gracas .
INFORMATION SCIENCES, 2009, 179 (05) :584-599
[8]   Scale-Free Networks: A Decade and Beyond [J].
Barabasi, Albert-Laszlo .
SCIENCE, 2009, 325 (5939) :412-413
[10]  
Boginski V L, 2005, THESIS U FLORIDA