Vietnamese treebank construction and entropy-based error detection.

Treebanks, especially the Penn treebank for natural language processing (NLP) in English, play an essential role in both research into and the application of NLP. However, many languages still lack treebanks and building a treebank can be very complicated and difficult. This work has a twofold objec...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 49; no. 3; pp. 487 - 520
Autores principales: Nguyen, Phuong-Thai, Le, Anh-Cuong, Ho, Tu-Bao, Nguyen, Van-Hiep
Formato: Artículo
Publicado: Springer Nature Sep2015
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=108465722&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 108465722
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2015
      vid: 49
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        108465722
        10.1007/s10579-015-9308-5
      ppf: 487
      ppct: 33
      formats:
        fmt:
          @attributes:
            type: P
            size: 3.5MB
      tig:
        atl: Vietnamese treebank construction and entropy-based error detection.
      aug:
        au:
          Nguyen, Phuong-Thai
          Le, Anh-Cuong
          Ho, Tu-Bao
          Nguyen, Van-Hiep
        affil:
          University of Engineering and Technology, Vietnam National University, Hanoi Vietnam
          Japan Advanced Institute of Science and Technology, Nomi Japan
          Institute of Linguistics, Vietnam Academy of Social Sciences, Hanoi Vietnam
      su:
        Vietnamese language
        Syntax (Grammar)
        Comparative grammar
        Foreign language education
        Ethnology
        Education
        Vietnam
      sug:
        subj:
          Vietnam
          Vietnamese language
          Syntax (Grammar)
          Comparative grammar
          Foreign language education
          Ethnology
          Education
      keyword:
        Entropy
        Error detection
        Treebank
      ab: Treebanks, especially the Penn treebank for natural language processing (NLP) in English, play an essential role in both research into and the application of NLP. However, many languages still lack treebanks and building a treebank can be very complicated and difficult. This work has a twofold objective. Firstly, to share our results in constructing a large Vietnamese treebank (VTB) with three levels of annotation including word segmentation, part-of-speech tagging, and syntactic analysis. Major steps in the treebank construction process are described with particular regard to specific Vietnamese properties such as lack of word delimiter and isolation. Those properties make sentences highly syntactically ambiguous, and therefore it is difficult to ensure a high level of agreement among annotators. Various studies of Vietnamese syntax were employed not only to define annotations but also to systematically deal with ambiguities. Annotators were supported by automatic labelling tools, which are based on statistical machine learning methods, for sentence pre-processing and a tree editor for supporting manual annotation. As a result, an annotation agreement of around 90 % was achieved. Our second objective is to present our method for automatically finding errors and inconsistencies in treebank corpora and its application to the construction of the VTB. This method employs the Shannon entropy measure in a manner that the more reduced entropy the more corrected errors in a treebank. The method ranks error candidates by using a scoring function based on conditional entropy. Our experiments showed that this method detected high-error-density subsets of original error candidate sets, and that the corpus entropy was significantly reduced after error correction. The size of these subsets was only about one third of the whole set, while these subsets contained 80-90 % of the total errors. This method can also be applied to languages similar to Vietnamese.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2015. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2015
    holdings:
      @attributes:
        islocal: N