Automatic language identification: a case study of Pahari languages.

In an attempt to expand the inclusiveness of Natural Language Processing, this paper focuses on developing resources and building machine learning models to identify four languages of the Northern Indo-Aryan family, also known as Pahari languages—Nepali, Garhwali, Kumaoni, and Dogri. This is the fir...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 57; no. 3; pp. 1361 - 1388
Autores principales: Gusain, Rachana, Dash, Satya Ranjan, Parida, Shantipriya, Jha, Girish Nath
Formato: Artículo
Publicado: Springer Nature Sep2023
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=170029266&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 170029266
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2023
      vid: 57
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        170029266
        10.1007/s10579-023-09651-6
      ppf: 1361
      ppct: 27
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.2MB
      tig:
        atl: Automatic language identification: a case study of Pahari languages.
      aug:
        au:
          Gusain, Rachana
          Dash, Satya Ranjan
          Parida, Shantipriya
          Jha, Girish Nath
        affil:
          Doon University, Dehradun, Uttarakhand, India
          KIIT University, Bhubaneswar, Odisha, India
          Silo AI, Helsinki, Finland
          Jawaharlal Nehru University, New Delhi, India
      su:
        Machine learning
        Natural language processing
        Automatic identification
        Support vector machines
        Native language
      sug:
        subj:
          Machine learning
          Natural language processing
          Automatic identification
          Support vector machines
          Native language
      keyword:
        Corpus development
        Dogri
        Garhwali
        Kumaoni
        Language identification
        Low-resource languages
        Nepali
        Northern Indo-Aryan
        Pahari
        Statistical analysis
      ab: In an attempt to expand the inclusiveness of Natural Language Processing, this paper focuses on developing resources and building machine learning models to identify four languages of the Northern Indo-Aryan family, also known as Pahari languages—Nepali, Garhwali, Kumaoni, and Dogri. This is the first attempt towards building identification models for Pahari languages and developing a plain text corpus for Garhwali and Kumaoni, both of which are lesser-known and under-resourced languages/mother tongues of India. The collected corpus, including data in Nepali and Dogri, is statistically analyzed at the word level. We also trained traditional machine learning models for Pahari language identification on this corpus and found that character n-grams based Linear Support Vector Machines performed best with 99.28% accuracy.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2023. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2023
    holdings:
      @attributes:
        islocal: N