Automatic language identification: a case study of Pahari languages.
In an attempt to expand the inclusiveness of Natural Language Processing, this paper focuses on developing resources and building machine learning models to identify four languages of the Northern Indo-Aryan family, also known as Pahari languages—Nepali, Garhwali, Kumaoni, and Dogri. This is the fir...
| Publicado en: | Language Resources & Evaluation Vol. 57; no. 3; pp. 1361 - 1388 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Sep2023
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=170029266&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 170029266 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2023 vid: 57 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 170029266 10.1007/s10579-023-09651-6 ppf: 1361 ppct: 27 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.2MB tig: atl: Automatic language identification: a case study of Pahari languages. aug: au: Gusain, Rachana Dash, Satya Ranjan Parida, Shantipriya Jha, Girish Nath affil: Doon University, Dehradun, Uttarakhand, India KIIT University, Bhubaneswar, Odisha, India Silo AI, Helsinki, Finland Jawaharlal Nehru University, New Delhi, India su: Machine learning Natural language processing Automatic identification Support vector machines Native language sug: subj: Machine learning Natural language processing Automatic identification Support vector machines Native language keyword: Corpus development Dogri Garhwali Kumaoni Language identification Low-resource languages Nepali Northern Indo-Aryan Pahari Statistical analysis ab: In an attempt to expand the inclusiveness of Natural Language Processing, this paper focuses on developing resources and building machine learning models to identify four languages of the Northern Indo-Aryan family, also known as Pahari languages—Nepali, Garhwali, Kumaoni, and Dogri. This is the first attempt towards building identification models for Pahari languages and developing a plain text corpus for Garhwali and Kumaoni, both of which are lesser-known and under-resourced languages/mother tongues of India. The collected corpus, including data in Nepali and Dogri, is statistically analyzed at the word level. We also trained traditional machine learning models for Pahari language identification on this corpus and found that character n-grams based Linear Support Vector Machines performed best with 99.28% accuracy. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2023. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2023 holdings: @attributes: islocal: N |
|---|