Search arXivSearch

arXiv · 1802.07258

Rank dynamics of word usage at multiple scales

Abstract

The recent dramatic increase in online data availability has allowed researchers to explore human culture with unprecedented detail, such as the growth and diversification of language. In particular, it provides statistical tools to explore whether word use is similar across languages, and if so, whether these generic features appear at different scales of language structure. Here we use the Google Books $N$-grams dataset to analyze the temporal evolution of word usage in several languages. We apply measures proposed recently to study rank dynamics, such as the diversity of $N$-grams in a given rank, the probability that an $N$-gram changes rank between successive time intervals, the rank entropy, and the rank complexity. Using different methods, results show that there are generic properties for different languages at different scales, such as a core of words necessary to minimally understand a language. We also propose a null model to explore the relevance of linguistic structure across multiple scales, concluding that $N$-gram statistics cannot be reduced to word statistics. We expect our results to be useful in improving text prediction algorithms, as well as in shedding light on the large-scale features of language use, beyond linguistic and cultural differences across human populations.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

José A. Morales, Ewan Colman, Sergio Sánchez, Fernanda Sánchez-Puig, Carlos Pineda, Gerardo Iñiguez, Germinal Cocho, Jorge Flores, Carlos Gershenson. 2018-02-20. Rank dynamics of word usage at multiple scales. https://doi.org/10.3389/fphy.2018.00045

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Extending the Biswas--Chatterjee--Sen model with nonconformists and inflexibles

Originally, the Biswas--Chatterjee--Sen model was shown to exhibit an order/disorder phase transition for a sufficiently large number of negative interactions among actors. In this paper, the model is extended by the existence of anticonformists and inflexibles. Anticonformists are actors who define themselves in opposition to the group and may intentionally reject what most people accept, while inflexibles are those who do not change their opinions at all. Both discrete and continuous opinions are considered. With direct Monte Carlo simulations and mean-field calculations, we check the influence of fractions of anticonformists and inflexibles on the mean opinion in the system. With the mean-field calculations, we identify ranges of fractions of anticonformists where an ordered phase of the system is available. The results of the mean-field calculations perfectly match the results of the Monte Carlo simulations. We consider inflexibles adhered: (i) to extreme opinions; (ii) to specific opinions, and (iii) chosen independently of their initial opinion. For inflexibles adhered to specific and extreme opinions, they play a role of an effective bias suppressing the disordered phase in the system. The qualitative results of introducing anticonformists (inflexibles) in various ways (discrete/continuous opinions and annealed/quenched disorder) are roughly the same. However, for the model extended by inflexibles, we can observe a systematic shift of the mean order parameter to its higher values for quenched disorder compared with annealed disorder. On the other hand, for anticonformists modeled with a continuous space of opinions, we can observe a systematic shift of the mean order parameter to its higher values compared with the discrete space of opinions.

physics.soc-ph

Inferring Coupling Strength from the Kuramoto Order Parameter

Accurately estimating the coupling strength in oscillator networks from macroscopic observations is essential for predicting synchronization transitions. We consider the inverse problem of reconstructing the unknown coupling strength $K$ in the globally coupled Kuramoto model from scalar observations of the macroscopic order parameter $R(t)$, assuming that the natural frequencies and the initial phase configuration are known. After initialization, individual phase trajectories are treated as hidden, and only the scalar order parameter is observed. We employ an extended Kalman filter with an augmented state representation that recursively estimates the coupling strength from observations of $R(t)$. By exploiting the mean-field structure of the globally coupled Kuramoto model, the covariance prediction step can be computed efficiently, substantially reducing the computational cost. Numerical simulations demonstrate that the proposed estimator accurately reconstructs the coupling strength and remains stable even when $R(t)$ is small and strongly fluctuating.

physics.soc-ph

Blind directions of physical learning networks: where to measure and what to measure

A physical learning network is read at a few accessible nodes. Every task and learning rule that uses only the steady voltages and currents there, at a fixed operating point, acts through the boundary response map. Changes in the kernel of its Jacobian are blind to first order. A walk along one fiber ran to a 600-step cap, one edge then at 15.2 times its start. A decomposition theorem splits the response Jacobian over the hidden components. The blind dimension adds over components whose surviving slots are disjoint. A boundary-to-boundary edge adds one parameter and deletes one slot. Exposing a hidden node changes only its component. For one hidden node the contribution counts the bipartite components of the non-adjacency graph of its neighbors. For a pocket of h hidden nodes a factor-analysis bound caps what outside electrodes can expose. It is attained when every pocket node meets every neighbor and no edge joins two of those neighbors, so past a threshold further electrodes outside such a pocket expose nothing. An electrode inside such a pocket, when its hidden nodes form a clique, is worth h-1 directions where one outside is worth none. Read as vector displacements rather than potentials, the same nodes left no deficit beyond counting in all 90 spring networks tested, each a pocket fully joined to five or more accessible nodes, and nothing blind in 88. What limits the reading is the quantity measured as much as the number of contacts. Maximum-weight spanning forests give a proved upper bound on the blind dimension. The matching equality is proved for one hidden node and conjectured beyond. It holds in all 5,343 components whose maximum could be attained and certified, and the forest count matched the certified blind dimension in 1,448 of 1,500 held-out networks, where the maximum must be searched. A self-learning circuit shows the split on hardware.

physics.soc-ph