A data-based classification of Slavic languages: Indices of qualitative variation applied to grapheme frequencies
April 14, 2015 Β· Declared Dead Β· π Journal of Quantitative Linguistics
"No code URL or promise found in abstract"
Evidence collected by the PWNC Scanner
Authors
Michaela KoscovΓ‘, JΓ‘n Macutek, Emmerich Kelih
arXiv ID
1504.03608
Category
stat.AP
Cross-listed
cs.CL
Citations
6
Venue
Journal of Quantitative Linguistics
Last Checked
6 months ago
Abstract
The Ord's graph is a simple graphical method for displaying frequency distributions of data or theoretical distributions in the two-dimensional plane. Its coordinates are proportions of the first three moments, either empirical or theoretical ones. A modification of the Ord's graph based on proportions of indices of qualitative variation is presented. Such a modification makes the graph applicable also to data of categorical character. In addition, the indices are normalized with values between 0 and 1, which enables comparing data files divided into different numbers of categories. Both the original and the new graph are used to display grapheme frequencies in eleven Slavic languages. As the original Ord's graph requires an assignment of numbers to the categories, graphemes were ordered decreasingly according to their frequencies. Data were taken from parallel corpora, i.e., we work with grapheme frequencies from a Russian novel and its translations to ten other Slavic languages. Then, cluster analysis is applied to the graph coordinates. While the original graph yields results which are not linguistically interpretable, the modification reveals meaningful relations among the languages.
Community Contributions
Found the code? Know the venue? Think something is wrong? Let us know!
π Similar Papers
In the same crypt β stat.AP
R.I.P.
π»
Ghosted
R.I.P.
π»
Ghosted
Sequence-to-point learning with neural networks for nonintrusive load monitoring
R.I.P.
π»
Ghosted
Predictive Business Process Monitoring with LSTM Neural Networks
R.I.P.
π»
Ghosted
Forecasting: theory and practice
R.I.P.
π»
Ghosted
Accurate estimation of influenza epidemics using Google search data via ARGO
R.I.P.
π»
Ghosted
Survey of resampling techniques for improving classification performance in unbalanced datasets
Died the same way β π» Ghosted
R.I.P.
π»
Ghosted
Federated Learning: Strategies for Improving Communication Efficiency
R.I.P.
π»
Ghosted
In-Datacenter Performance Analysis of a Tensor Processing Unit
R.I.P.
π»
Ghosted
Deep Convolutional Neural Networks for Computer-Aided Detection: CNN Architectures, Dataset Characteristics and Transfer Learning
R.I.P.
π»
Ghosted