Quantifying the dynamics of topical fluctuations in language

June 02, 2018 · Entered Twilight · 🏛 Language Dynamics and Change

"Last commit was 6.0 years ago (≥5 year threshold)"

Evidence collected by the PWNC Scanner

Repo contents: .gitattributes, .gitignore, LICENSE, README.md, advection_functions.R, advection_paper.R

Authors Andres Karjus, Richard A. Blythe, Simon Kirby, Kenny Smith arXiv ID 1806.00699 Category cs.CL: Computation & Language Citations 20 Venue Language Dynamics and Change Repository https://github.com/andreskarjus/topical_cultural_advection_model ⭐ 6 Last Checked 1 month ago

Abstract

The availability of large diachronic corpora has provided the impetus for a growing body of quantitative research on language evolution and meaning change. The central quantities in this research are token frequencies of linguistic elements in texts, with changes in frequency taken to reflect the popularity or selective fitness of an element. However, corpus frequencies may change for a wide variety of reasons, including purely random sampling effects, or because corpora are composed of contemporary media and fiction texts within which the underlying topics ebb and flow with cultural and socio-political trends. In this work, we introduce a simple model for controlling for topical fluctuations in corpora - the topical-cultural advection model - and demonstrate how it provides a robust baseline of variability in word frequency changes over time. We validate the model on a diachronic corpus spanning two centuries, and a carefully-controlled artificial language change scenario, and then use it to correct for topical fluctuations in historical time series. Finally, we use the model to show that the emergence of new words typically corresponds with the rise of a trending topic. This suggests that some lexical innovations occur due to growing communicative need in a subspace of the lexicon, and that the topical-cultural advection model can be used to quantify this.