A feature-based approach to better automatic treebank conversion.

In the field of constituency parsing, there exist multiple human-labeled treebanks which are built on non-overlapping text samples and follow different annotation standards. Due to the extreme cost of annotating parse trees by human, it is desirable to automatically convert one treebank (called ) to...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 47; no. 4; pp. 1213 - 1232
Autores principales: Zhu, Muhua, Zhu, Jingbo, Wang, Huizhen
Formato: Artículo
Publicado: Springer Nature Dec2013
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:In the field of constituency parsing, there exist multiple human-labeled treebanks which are built on non-overlapping text samples and follow different annotation standards. Due to the extreme cost of annotating parse trees by human, it is desirable to automatically convert one treebank (called ) to the standard of another treebank (called ) which we are interested in. Conversion results can be manually corrected to obtain higher-quality annotations or can be directly used as additional training data for building syntactic parsers. To perform automatic treebank conversion, we divide constituency parses into two separate levels: the part-of-speech (POS) and syntactic structure (bracketing structures and constituent labels), and conduct conversion on these two levels respectively with a feature-based approach. The basic idea of the approach is to encode original annotations in a source treebank as guide features during the conversion process. Experiments on two Chinese treebanks show that our approach can convert POS tags and syntactic structures with the accuracy of 96.6 and 84.8 %, respectively, which are the best reported results on this task.