The development of a labelled te reo Māori–English bilingual database for language technology: The development of a labelled te reo Māori–English bilingual corpus: J. James et al.

Te reo Māori (referred to as Māori), New Zealand's indigenous language, is under-resourced in language technology. Māori speakers are bilingual, where Māori is code-switched with English. Unfortunately, there are minimal resources available for Māori language technology, language detection and code-...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 59; no. 1; pp. 1 - 27
Autores principales: James, Jesin, Shields, Isabella, Yogarajan, Vithya, Keegan, Peter J., Watson, Catherine I., Jones, Peter-Lucas, Mahelona, Keoni
Formato: Artículo
Publicado: Springer Nature Mar2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=183750653&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 183750653
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Mar2025
      vid: 59
      iid: 1
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        183750653
        10.1007/s10579-023-09680-1
      ppf: 1
      ppct: 26
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 2.3MB
      tig:
        atl: The development of a labelled te reo Māori–English bilingual database for language technology: The development of a labelled te reo Māori–English bilingual corpus: J. James et al.
      aug:
        au:
          James, Jesin
          Shields, Isabella
          Yogarajan, Vithya
          Keegan, Peter J.
          Watson, Catherine I.
          Jones, Peter-Lucas
          Mahelona, Keoni
        affil:
          https://ror.org/03b94tp07 Department of Electrical, Computer, and Software Engineering, The University of Auckland, Auckland, New Zealand
          https://ror.org/03b94tp07 Strong AI Lab, The University of Auckland, Auckland, New Zealand
          https://ror.org/03b94tp07 Te Puna Wānanga, The University of Auckland, Auckland, New Zealand
          Te Hiku Media, Kaitaia, New Zealand
      su:
        Linguistics
        Language & languages
        Language acquisition
        Speech
        Cognitive psychology
      sug:
        subj:
          Linguistics
          Language & languages
          Language acquisition
          Speech
          Cognitive psychology
      keyword:
        Code-switching
        Communication and Culture Linguistics Psychology and Cognitive Sciences Cognitive Sciences
        Language
        Language identification
        Language technology
        Linguistic rules
        Low-resourced languages
        Māori
      ab: Te reo Māori (referred to as Māori), New Zealand's indigenous language, is under-resourced in language technology. Māori speakers are bilingual, where Māori is code-switched with English. Unfortunately, there are minimal resources available for Māori language technology, language detection and code-switch detection between Māori–English pair. Both English and Māori use Roman-derived orthography making rule-based systems for detecting language and code-switching restrictive. Most Māori language detection is done manually by language experts. This research builds a Māori–English bilingual database of 66,016,807 words with word-level language annotation. The New Zealand Parliament Hansard debates reports were used to build the database. The language labels are assigned automatically using language-specific rules and expert manual annotations. Words with the same spelling, but different meanings, exist for Māori and English. These words could not be categorised as Māori or English based on word-level language rules. Hence, manual annotations were necessary. An analysis reporting the various aspects of the database such as metadata, year-wise analysis, frequently occurring words, sentence length and N-grams is also reported. The database developed here is a valuable tool for future language and speech technology development for Aotearoa New Zealand. The methodology followed to label the database can also be followed by other low-resourced language pairs.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N