Creation of an annotated corpus of Old and Middle Hungarian court records and private correspondence.

The paper introduces a novel annotated corpus of Old and Middle Hungarian (16–18 century), the texts of which were selected in order to approximate the vernacular of the given historical periods as closely as possible. The corpus consists of testimonies of witnesses in trials and samples of private...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 52; no. 1; pp. 1 - 29
Autores principales: Novák, Attila, Gugán, Katalin, Varga, Mónika, Dömötör, Adrienne
Formato: Artículo
Publicado: Springer Nature Mar2018
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=127930800&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 127930800
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Mar2018
      vid: 52
      iid: 1
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        127930800
        10.1007/s10579-017-9393-8
      ppf: 1
      ppct: 28
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.2MB
      tig:
        atl: Creation of an annotated corpus of Old and Middle Hungarian court records and private correspondence.
      aug:
        au:
          Novák, Attila
          Gugán, Katalin
          Varga, Mónika
          Dömötör, Adrienne
        affil:
          MTA-PPKE Hungarian Language Technology Research Group, Práter u. 50/a, Budapest, Hungary
          Pázmány Péter Catholic University, Faculty of Information Technology and Bionics, Práter u. 50/a, 1083, Budapest, Hungary
          Research Institute for Linguistics of the Hungarian Academy of Sciences, Benczúr u. 33, 1068, Budapest, Hungary
      su:
        Native language
        Metadata
        Sociolinguistic research
        Annotations
        Court records
      sug:
        subj:
          Native language
          Metadata
          Sociolinguistic research
          Annotations
          Court records
      keyword:
        Corpus annotation
        Corpus query tool
        Historical corpus
        Middle Hungarian
        Morphological analysis
        Old Hungarian
        PoS tagging
      ab: The paper introduces a novel annotated corpus of Old and Middle Hungarian (16–18 century), the texts of which were selected in order to approximate the vernacular of the given historical periods as closely as possible. The corpus consists of testimonies of witnesses in trials and samples of private correspondence. The texts are not only analyzed morphologically, but each file contains metadata that would also facilitate sociolinguistic research. The texts were segmented into clauses, manually normalized and morphosyntactically annotated using an annotation system consisting of the PurePos PoS tagger and the Hungarian morphological analyzer HuMor originally developed for Modern Hungarian but adapted to analyze Old and Middle Hungarian morphological constructions. The automatically disambiguated morphological annotation was manually checked and corrected using an easy-to-use web-based manual disambiguation interface. The normalization process and the manual validation of the annotation required extensive teamwork and provided continuous feedback for the refinement of the computational morphology and iterative retraining of the statistical models of the tagger. The paper discusses some of the typical problems that occurred during the normalization procedure and their tentative solutions. Besides, we also describe the automatic annotation tools, the process of semi-automatic disambiguation, and the query interface, a special function of which also makes correction of the annotation possible. Displaying the original, the normalized and the parsed versions of the selected texts, the beta version of the first fully normalized and annotated historical corpus of Hungarian is freely accessible at the address <ext-link>http://tmk.nytud.hu/</ext-link>.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2018. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2018
    holdings:
      @attributes:
        islocal: N