Syntactic annotation for Portuguese corpora: standards, parsers, and search interfaces.

In the last two decades, four Portuguese syntactically annotated corpora were built along the lines initially defined for the Penn Parsed Historical Corpora (Santorini, 2016). They cover the old, the middle, the classical and the modern periods of European Portuguese, as well as the nineteenth and t...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 58; no. 1; pp. 301 - 347
Autores principales: Faria, Pablo, Galves, Charlotte, Magro, Catarina
Formato: Artículo
Publicado: Springer Nature Mar2024
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=176079993&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 176079993
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Mar2024
      vid: 58
      iid: 1
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        176079993
        10.1007/s10579-023-09699-4
      ppf: 301
      ppct: 46
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 2.3MB
      tig:
        atl: Syntactic annotation for Portuguese corpora: standards, parsers, and search interfaces.
      aug:
        au:
          Faria, Pablo
          Galves, Charlotte
          Magro, Catarina
        affil:
          https://ror.org/04wffgt70 Linguistics Department, Universidade Estadual de Campinas, Campinas, Brazil
          https://ror.org/01c27hj86 Faculdade de Letras, Centro de Linguística, Universidade de Lisboa, Lisbon, Portugal
      su:
        Corpora
        Annotations
        Nineteenth century
        Twentieth century
        Standards
      sug:
        subj:
          Corpora
          Annotations
          Nineteenth century
          Twentieth century
          Standards
      keyword:
        Corpus search
        Portuguese parsed corpora
        Probabilistic parsers
        Rule-based parsers
      ab: In the last two decades, four Portuguese syntactically annotated corpora were built along the lines initially defined for the Penn Parsed Historical Corpora (Santorini, 2016). They cover the old, the middle, the classical and the modern periods of European Portuguese, as well as the nineteenth and twentieth century Brazilian Portuguese, and include different textual genres and oral discourse excerpts. Together they provide a fundamental resource for the study of variation and change in Portuguese. In the last years, an effort was made to maximally unify the annotation scheme applied to those corpora, in such a way that the searches done on one corpus could be done in exactly the same manner on the others. This effort resulted in the Portuguese Syntactic Annotation Manual (Magro & Galves, 2019). In this paper, we present the syntactic annotation for the Portuguese Corpora. We describe the functioning of ParsPort, a rule-based parser which makes use of the revision mode of the query language Corpus Search (Randall, 2005–2015). We argue that ParsPort is more efficient to our annotation efforts than the probabilistic parser developed by Bikel (2004), previously used for the syntactic annotation of the Portuguese Corpora. Finally we mention recent advances towards more user-friendly tools for syntactic searches.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2024. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2024
    holdings:
      @attributes:
        islocal: N