Morphosyntactic annotation of CHILDES transcripts.
Corpora of child language are essential for research in child language acquisition and psycholinguistics. Linguistic annotation of the corpora provides researchers with better means for exploring the development of grammatical constructions and their usage. We describe a project whose goal is to ann...
| Publicado en: | Journal of Child Language Vol. 37; no. 3; pp. 705 - 730 |
|---|---|
| Autores principales: | , , , , |
| Formato: | equations & formulas research tables/charts Journal Article |
| Publicado: |
Cambridge University Press
Jun2010
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=105020670&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 105020670 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 03050009 JCL jtl: Journal of Child Language issn: 03050009 maglogo: N pubinfo: dt: Jun2010 vid: 37 iid: 3 pid: 15979 pub: Cambridge University Press artinfo: ui: 105020670 52872452 2010676543 10.1017/S0305000909990407 NLM20334720 105020670 ppf: 705 ppct: 25 formats: tig: atl: Morphosyntactic annotation of CHILDES transcripts. aug: au: Sagae K Davis E Lavie A MacWhinney B Wintner S affil: Institute for Creative Technologies, University of Southern California, CA 90292, USA. sagae@usc.edu sug: subj: Databases Utilization Grammar Evaluation Language Development In Infancy and Childhood Learning Methods In Infancy and Childhood Speech Sample In Infancy and Childhood Adult Algorithms Utilization Child Funding Source Grammar Standards Human Multilingualism Research Methodology Speech Sample In Adulthood Validation Studies Adult: 19-44 years Child: 6-12 years ab: Corpora of child language are essential for research in child language acquisition and psycholinguistics. Linguistic annotation of the corpora provides researchers with better means for exploring the development of grammatical constructions and their usage. We describe a project whose goal is to annotate the English section of the CHILDES database with grammatical relations in the form of labeled dependency structures. We have produced a corpus of over 18,800 utterances (approximately 65,000 words) with manually curated gold-standard grammatical relation annotations. Using this corpus, we have developed a highly accurate data-driven parser for the English CHILDES data, which we used to automatically annotate the remainder of the English section of CHILDES. We have also extended the parser to Spanish, and are currently working on supporting more languages. The parser and the manually and automatically annotated data are freely available for research purposes. pubtype: Academic Journal doctype: equations & formulas research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|