Compilation of an idiom example database for supervised idiom identification.

Some phrases can be interpreted in their context either idiomatically (figuratively) or literally. The precise identification of idioms is essential in order to achieve full-fledged natural language processing. Because of this, the authors of this paper have created an idiom corpus for Japanese. Thi...

Full description

Bibliographic Details
Published in:Language Resources & Evaluation Vol. 43; no. 4; pp. 355 - 385
Main Authors: Hashimoto, Chikara, Kawahara, Daisuke
Format: Article
Published: Springer Nature Dec2009
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=45284386&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 45284386
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Dec2009
      vid: 43
      iid: 4
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        45284386
        10.1007/s10579-009-9104-1
      ppf: 355
      ppct: 30
      formats:
        fmt:
          @attributes:
            type: P
            size: 882KB
      tig:
        atl: Compilation of an idiom example database for supervised idiom identification.
      aug:
        au:
          Hashimoto, Chikara
          Kawahara, Daisuke
        affil: National Institute of Information and Communications Technology, Kyoto, Japan.
      su:
        Idioms
        Artificial intelligence
        Computational linguistics
        Electronic data processing
        Human-computer interaction
      sug:
        subj:
          Idioms
          Artificial intelligence
          Computational linguistics
          Electronic data processing
          Human-computer interaction
      keyword:
        Corpus
        Idiom identification
        Japanese idiom
        Language resources
      ab: Some phrases can be interpreted in their context either idiomatically (figuratively) or literally. The precise identification of idioms is essential in order to achieve full-fledged natural language processing. Because of this, the authors of this paper have created an idiom corpus for Japanese. This paper reports on the corpus itself and the results of an idiom identification experiment conducted using the corpus. The corpus targeted 146 ambiguous idioms, and consists of 102,856 examples, each of which is annotated with a literal/idiomatic label. All sentences were collected from the World Wide Web. For idiom identification, 90 out of the 146 idioms were targeted and a word sense disambiguation (WSD) method was adopted using both common WSD features and idiom-specific features. The corpus and the experiment are both, as far as can be determined, the largest of their kinds. It was discovered that a standard supervised WSD method works well for idiom identification and it achieved accuracy levels of 89.25 and 88.86%, with and without idiom-specific features, respectively. It was also found that the most effective idiom-specific feature is the one that involves the adjacency of idiom constituents.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2009. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2009
    holdings:
      @attributes:
        islocal: N