Enriching Portuguese Medieval Texts with Named Entity Recognition.

Historical data poses unique challenges to natural language processing (NLP) and information retrieval (IR) tools, including digitization errors, lack of annotated data, and diachronic-specific issues. However, the increasing recognition of the value in historical documents has promoted efforts to s...

Descripción completa

Detalles Bibliográficos
Publicado en:International Journal of Humanities & Arts Computing: A Journal of Digital Humanities Vol. 18; no. 1; pp. 109 - 125
Autores principales: Inês Bico, Maria, Baptista, Jorge, Batista, Fernando, Cardeira, Esperança
Formato: Artículo
Publicado: Edinburgh University Press Mar2024
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=176431506&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 176431506
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        17538548
        2QD7
      jtl: International Journal of Humanities & Arts Computing: A Journal of Digital Humanities
      issn: 17538548
      maglogo: N
    pubinfo:
      dt: Mar2024
      vid: 18
      iid: 1
      pid: 2327
      pub: Edinburgh University Press
    artinfo:
      ui:
        176431506
        10.3366/ijhac.2024.0324
      ppf: 109
      ppct: 16
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 284KB
      tig:
        atl: Enriching Portuguese Medieval Texts with Named Entity Recognition.
      aug:
        au:
          Inês Bico, Maria
          Baptista, Jorge
          Batista, Fernando
          Cardeira, Esperança
      su:
        Information retrieval
        Natural language processing
        Portuguese language
        Corpora
        Historical source material
        Knowledge base
      sug:
        subj:
          Information retrieval
          Natural language processing
          Portuguese language
          Corpora
          Historical source material
          Knowledge base
      keyword:
        corpus analysis
        information retrieval
        named entity disambiguation
        named entity linking
        natural language processing
        Portuguese medieval texts
      ab: Historical data poses unique challenges to natural language processing (NLP) and information retrieval (IR) tools, including digitization errors, lack of annotated data, and diachronic-specific issues. However, the increasing recognition of the value in historical documents has promoted efforts to semantically enrich and optimize their analysis. This article contributes to this endeavour by enriching the Corpus de Textos Antigos through NLP tools and techniques to enhance its usability and support research. The corpus undergoes linguistic annotation, including part-of-speech tagging, lemma annotation and named entity recognition (NER). Subsequently, the article delves into the tasks of entity disambiguation and entity linking, which involve identifying and disambiguating named entities by referring to a knowledge base (KB). Addressing the challenges posed by factors such as text state, epoch and the chosen KB, the article presents insights into related work, annotation results and the linguistic interest of a medieval annotated corpus for named entities. It concludes by discussing the challenges and providing avenues for future research in this domain.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Copyright of International Journal of Humanities & Arts Computing: A Journal of Digital Humanities is the property of Edinburgh University Press and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use.
      item: International Journal of Humanities & Arts Computing: A Journal of Digital Humanities
      holder: Edinburgh University Press
      dt:
        @attributes:
          year: 2024
    holdings:
      @attributes:
        islocal: N