Building a learner corpus.

The need for data about the acquisition of Czech by non-native learners prompted the compilation of the first learner corpus of Czech. After introducing its basic design and parameters, including a multi-tier manual annotation scheme and error taxonomy, we focus on the more technical aspects: the tr...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 48; no. 4; pp. 741 - 753
Autores principales: Hana, Jirka, Rosen, Alexandr, Štindlová, Barbora, Štěpánek, Jan
Formato: Artículo
Publicado: Springer Nature Dec2014
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=99708562&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 99708562
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Dec2014
      vid: 48
      iid: 4
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        99708562
        10.1007/s10579-014-9278-z
      ppf: 741
      ppct: 12
      formats:
        fmt:
          @attributes:
            type: P
            size: 966KB
      tig:
        atl: Building a learner corpus.
      aug:
        au:
          Hana, Jirka
          Rosen, Alexandr
          Štindlová, Barbora
          Štěpánek, Jan
        affil:
          Faculty of Mathematics and Physics, Charles University, Prague Czech Republic
          Faculty of Arts, Charles University, Prague Czech Republic
          Technical University, Liberec Czech Republic
      su:
        Czech language
        Corpora
        Annotations
        Spell checkers (Computer programs)
        Native language
        Grammar
      sug:
        subj:
          Czech language
          Corpora
          Annotations
          Spell checkers (Computer programs)
          Native language
          Grammar
      keyword:
        Czech
        Error annotation
        Learner corpus
      ab: The need for data about the acquisition of Czech by non-native learners prompted the compilation of the first learner corpus of Czech. After introducing its basic design and parameters, including a multi-tier manual annotation scheme and error taxonomy, we focus on the more technical aspects: the transcription of hand-written source texts, process of annotation, and options for exploiting the result, together with tools used for these tasks and decisions behind the choices. To support or even substitute manual annotation we assign some error tags automatically and use automatic annotation tools (tagger, spell checker).
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2014. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2014
    holdings:
      @attributes:
        islocal: N