A feature-based approach to better automatic treebank conversion.
In the field of constituency parsing, there exist multiple human-labeled treebanks which are built on non-overlapping text samples and follow different annotation standards. Due to the extreme cost of annotating parse trees by human, it is desirable to automatically convert one treebank (called ) to...
| Publicado en: | Language Resources & Evaluation Vol. 47; no. 4; pp. 1213 - 1232 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Dec2013
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=92719447&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 92719447 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Dec2013 vid: 47 iid: 4 pid: 237 pub: Springer Nature artinfo: ui: 92719447 10.1007/s10579-013-9234-3 ppf: 1213 ppct: 19 formats: fmt: @attributes: type: P size: 1.8MB tig: atl: A feature-based approach to better automatic treebank conversion. aug: au: Zhu, Muhua Zhu, Jingbo Wang, Huizhen affil: Natural Language Processing Laboratory, Northeastern University, Wenhua Road 3-11, Heping District, Shenyang, Liaoning Province, China su: Parts of speech Parametric modeling Annotations Syntax (Grammar) Chinese language sug: subj: Parts of speech Parametric modeling Annotations Syntax (Grammar) Chinese language keyword: Automatic treebank conversion Constituency syntactic structure Feature-based approach Part of speech ab: In the field of constituency parsing, there exist multiple human-labeled treebanks which are built on non-overlapping text samples and follow different annotation standards. Due to the extreme cost of annotating parse trees by human, it is desirable to automatically convert one treebank (called ) to the standard of another treebank (called ) which we are interested in. Conversion results can be manually corrected to obtain higher-quality annotations or can be directly used as additional training data for building syntactic parsers. To perform automatic treebank conversion, we divide constituency parses into two separate levels: the part-of-speech (POS) and syntactic structure (bracketing structures and constituent labels), and conduct conversion on these two levels respectively with a feature-based approach. The basic idea of the approach is to encode original annotations in a source treebank as guide features during the conversion process. Experiments on two Chinese treebanks show that our approach can convert POS tags and syntactic structures with the accuracy of 96.6 and 84.8 %, respectively, which are the best reported results on this task. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2013. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2013 holdings: @attributes: islocal: N |
|---|