Lexical diversity as a lens into the classification of Slavic languages: A quantitative typology perspective.
This study proposes a linguistic classification method based on quantitative typology, which leverages a large-scale multilingual parallel corpus to obtain valid language classification result by excluding the influence of covariates such as text genre and semantic content in cross-language comparis...
| Publicado en: | Digital Scholarship in the Humanities Vol. 38; no. 3; pp. 1359 - 1372 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Sep2023
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=171389434&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 171389434 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Sep2023 vid: 38 iid: 3 pid: 622 pub: Oxford University Press / USA artinfo: ui: 171389434 10.1093/llc/fqad042 ppf: 1359 ppct: 13 formats: fmt: – @attributes: type: T – @attributes: type: P size: 753KB tig: atl: Lexical diversity as a lens into the classification of Slavic languages: A quantitative typology perspective. aug: au: Zhou, Chenliang Liu, Haitao affil: Department of Linguistics, Zhejiang University , Hangzhou, China Institute of Quantitative Linguistics, Beijing Language and Culture University , Beijing, China Center for Linguistics and Applied Linguistics, Guangdong University of Foreign Studies , Guangdong, Guangzhou, China su: Slavic languages Linguistic complexity Classification sug: subj: Slavic languages Linguistic complexity Classification ab: This study proposes a linguistic classification method based on quantitative typology, which leverages a large-scale multilingual parallel corpus to obtain valid language classification result by excluding the influence of covariates such as text genre and semantic content in cross-language comparison. To achieve this, we model the type–token relationships of each Slavic parallel text and calculate the lexical diversity to approximate the morphological complexity of the language. We perform automatic clustering of languages based on these lexical diversity metrics. Our findings show that (1) the lexical diversity metrics can well reflect that the language is located somewhere on the continuum of 'analytism-synthetism'; (2) the automatic clustering based on these metrics effectively reflects the genealogical classification of Slavic languages; and (3) the geographical distribution of lexical diversity in the region where Slavic languages are spoken shows a monotonic increasing trend from southwest to northeast, which is consistent with the pattern found by previous authors on a global scale. The methodological approach taken in this study is data-driven, with the benefit of being independent of theoretical assumptions and easy for computer processing. This approach can offer a better insight into corpus-based typology and may shed light on the understanding of language as a human-driven complex adaptive system. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2023 holdings: @attributes: islocal: N |
|---|