A feature-based approach to better automatic treebank conversion.

In the field of constituency parsing, there exist multiple human-labeled treebanks which are built on non-overlapping text samples and follow different annotation standards. Due to the extreme cost of annotating parse trees by human, it is desirable to automatically convert one treebank (called ) to...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 47; no. 4; pp. 1213 - 1232
Autores principales: Zhu, Muhua, Zhu, Jingbo, Wang, Huizhen
Formato: Artículo
Publicado: Springer Nature Dec2013
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=92719447&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 92719447
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Dec2013
      vid: 47
      iid: 4
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        92719447
        10.1007/s10579-013-9234-3
      ppf: 1213
      ppct: 19
      formats:
        fmt:
          @attributes:
            type: P
            size: 1.8MB
      tig:
        atl: A feature-based approach to better automatic treebank conversion.
      aug:
        au:
          Zhu, Muhua
          Zhu, Jingbo
          Wang, Huizhen
        affil: Natural Language Processing Laboratory, Northeastern University, Wenhua Road 3-11, Heping District, Shenyang, Liaoning Province, China
      su:
        Parts of speech
        Parametric modeling
        Annotations
        Syntax (Grammar)
        Chinese language
      sug:
        subj:
          Parts of speech
          Parametric modeling
          Annotations
          Syntax (Grammar)
          Chinese language
      keyword:
        Automatic treebank conversion
        Constituency syntactic structure
        Feature-based approach
        Part of speech
      ab: In the field of constituency parsing, there exist multiple human-labeled treebanks which are built on non-overlapping text samples and follow different annotation standards. Due to the extreme cost of annotating parse trees by human, it is desirable to automatically convert one treebank (called ) to the standard of another treebank (called ) which we are interested in. Conversion results can be manually corrected to obtain higher-quality annotations or can be directly used as additional training data for building syntactic parsers. To perform automatic treebank conversion, we divide constituency parses into two separate levels: the part-of-speech (POS) and syntactic structure (bracketing structures and constituent labels), and conduct conversion on these two levels respectively with a feature-based approach. The basic idea of the approach is to encode original annotations in a source treebank as guide features during the conversion process. Experiments on two Chinese treebanks show that our approach can convert POS tags and syntactic structures with the accuracy of 96.6 and 84.8 %, respectively, which are the best reported results on this task.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2013. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2013
    holdings:
      @attributes:
        islocal: N