Modelling multi-level prosody and spectral features using deep neural network for an automatic tonal and non-tonal pre-classification-based Indian language identification system.

In this paper an attempt has been made to prepare an automatic tonal and non-tonal pre-classification-based Indian language identification (LID) system using multi-level prosody and spectral features. Languages are first categorized into tonal and non-tonal groups, and then, from among the languages...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 55; no. 3; pp. 689 - 731
Autores principales: China Bhanja, Chuya, Laskar, Mohammad Azharuddin, Laskar, Rabul Hussain
Formato: Artículo
Publicado: Springer Nature Sep2021
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=151686300&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 151686300
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2021
      vid: 55
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        151686300
        10.1007/s10579-020-09527-z
      ppf: 689
      ppct: 42
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 483KB
      tig:
        atl: Modelling multi-level prosody and spectral features using deep neural network for an automatic tonal and non-tonal pre-classification-based Indian language identification system.
      aug:
        au:
          China Bhanja, Chuya
          Laskar, Mohammad Azharuddin
          Laskar, Rabul Hussain
        affil: Department of Electronics and Communication Engineering, National Institute of Technology Silchar, 788 010, Assam, India
      su:
        Multilevel models
        System identification
        Prosodic analysis (Linguistics)
        Gaussian mixture models
        Artificial neural networks
      sug:
        subj:
          Multilevel models
          System identification
          Prosodic analysis (Linguistics)
          Gaussian mixture models
          Artificial neural networks
      keyword:
        Classifiers
        Databases
        Multi-level analysis
        Prosody and spectral features
        Tonal and non-tonal languages
      ab: In this paper an attempt has been made to prepare an automatic tonal and non-tonal pre-classification-based Indian language identification (LID) system using multi-level prosody and spectral features. Languages are first categorized into tonal and non-tonal groups, and then, from among the languages of the respective groups, individual languages are identified. The system uses syllable, word (tri-syllable) and phrase level (multi-word) prosody (collectively called multi-level prosody) along with spectral features, namely Mel-frequency cepstral coefficients (MFCCs), Mean Hilbert envelope coefficients (MHEC), and shifted delta cepstral coefficients of MFCCs and MHECs for the pre-classification task. Multi-level analysis of spectral features has also been proposed and the complementarity of the syllable, word and phrase level (spectral + prosody) has been examined for pre-classification-based LID task. Four different models, particularly, Gaussian Mixture Model (GMM)-Universal Background Model (UBM), Artificial Neural Network (ANN), i-vector based support vector machine (SVM) and Deep Neural Network (DNN) have been developed to identify the languages. Experiments have been carried out on National Institute of Technology Silchar language database (NITS-LD) and OGI Multi-language Telephone Speech corpus (OGI-MLTS). The experiments confirm that both prosody and (spectral + prosody) obtained from syllable-, word- and phrase-level carry complementary information for pre-classification-based LID task. At the pre-classification stage, DNN models based on multi-level (prosody + MFCC) features, coupled with score combination technique results in the lowest EER value of 9.6% for NITS-LD. For OGI-MLTS database, the lowest EER value of 10.2% is observed for multi-level (prosody + MHEC). The pre-classification module helps to improve the performance of baseline single-stage LID system by 3.2% and 4.2% for NITS-LD and OGI-MLTS database respectively.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2021. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2021
    holdings:
      @attributes:
        islocal: N