Evaluating and automating the annotation of a learner corpus.
The paper describes a corpus of texts produced by non-native speakers of Czech. We discuss its annotation scheme, consisting of three interlinked tiers, designed to handle a wide range of error types present in the input. Each tier corrects different types of errors; links between the tiers allow ca...
| Publicado en: | Language Resources & Evaluation Vol. 48; no. 1; pp. 65 - 93 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Mar2014
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=95006720&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 95006720 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Mar2014 vid: 48 iid: 1 pid: 237 pub: Springer Nature artinfo: ui: 95006720 10.1007/s10579-013-9226-3 ppf: 65 ppct: 28 formats: fmt: @attributes: type: P size: 627KB tig: atl: Evaluating and automating the annotation of a learner corpus. aug: au: Rosen, Alexandr Hana, Jirka Štindlová, Barbora Feldman, Anna affil: Charles University, Prague Czech Republic Technical University, Liberec Czech Republic Montclair State University, Montclair USA su: Annotations Second language acquisition Grammar checkers (Computer software) Spell checkers (Computer programs) Czech language sug: subj: Annotations Second language acquisition Grammar checkers (Computer software) Spell checkers (Computer programs) Czech language keyword: Czech Error annotation Learner corpus ab: The paper describes a corpus of texts produced by non-native speakers of Czech. We discuss its annotation scheme, consisting of three interlinked tiers, designed to handle a wide range of error types present in the input. Each tier corrects different types of errors; links between the tiers allow capturing errors in word order and complex discontinuous expressions. Errors are not only corrected, but also classified. The annotation scheme is tested on a data set including approx. 175,000 words with fair inter-annotator agreement results. We also explore the possibility of applying automated linguistic annotation tools (taggers, spell checkers and grammar checkers) to the learner text to support or even substitute manual annotation. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2014. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2014 holdings: @attributes: islocal: N |
|---|