A comparative study of dictionaries and corpora as methods for language resource addition.

In this paper, we investigate the relative effect of two strategies for language resource addition for Japanese morphological analysis, a joint task of word segmentation and part-of-speech tagging. The first strategy is adding entries to the dictionary and the second is adding annotated sentences to...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 50; no. 2; pp. 245 - 262
Autores principales: Mori, Shinsuke, Neubig, Graham
Formato: Artículo
Publicado: Springer Nature Jun2016
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=116036662&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 116036662
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Jun2016
      vid: 50
      iid: 2
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        116036662
        10.1007/s10579-016-9354-7
      ppf: 245
      ppct: 17
      formats:
        fmt:
          @attributes:
            type: P
            size: 2.3MB
      tig:
        atl: A comparative study of dictionaries and corpora as methods for language resource addition.
      aug:
        au:
          Mori, Shinsuke
          Neubig, Graham
        affil:
          Academic Center for Computing and Media Studies, Kyoto University, Yoshidahonmachi, Sakyo-ku Kyoto Japan
          Nara Institute of Science and Technology, 8916-5 Takayamacho Ikoma Japan
      su:
        Annotations
        Encyclopedias & dictionaries
        Corpora
        Word recognition
        Japanese language
      sug:
        subj:
          Annotations
          Encyclopedias & dictionaries
          Corpora
          Word recognition
          Japanese language
      keyword:
        Dictionary
        Domain adaptation
        Non-maleficence of language resources
        Partial annotation
        POS tagging
        Word segmentation
      ab: In this paper, we investigate the relative effect of two strategies for language resource addition for Japanese morphological analysis, a joint task of word segmentation and part-of-speech tagging. The first strategy is adding entries to the dictionary and the second is adding annotated sentences to the training corpus. The experimental results showed that addition of annotated sentences to the training corpus is better than the addition of entries to the dictionary. In particular, adding annotated sentences is especially efficient when we add new words with contexts of several real occurrences as partially annotated sentences, i.e. sentences in which only some words are annotated with word boundary information. According to this knowledge, we performed real annotation experiments on invention disclosure texts and observed word segmentation accuracy. Finally we investigated various language resource addition cases and introduced the notion of non-maleficence, asymmetricity, and additivity of language resources for a task. In the WS case, we found that language resource addition is non-maleficent (adding new resources causes no harm in other domains) and sometimes additive (adding new resources helps other domains). We conclude that it is reasonable for us, NLP tool providers, to distribute only one general-domain model trained from all the language resources we have.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2016. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2016
    holdings:
      @attributes:
        islocal: N