Enriching Portuguese Medieval Texts with Named Entity Recognition.
Historical data poses unique challenges to natural language processing (NLP) and information retrieval (IR) tools, including digitization errors, lack of annotated data, and diachronic-specific issues. However, the increasing recognition of the value in historical documents has promoted efforts to s...
| Publicado en: | International Journal of Humanities & Arts Computing: A Journal of Digital Humanities Vol. 18; no. 1; pp. 109 - 125 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Edinburgh University Press
Mar2024
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=176431506&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 176431506 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 17538548 2QD7 jtl: International Journal of Humanities & Arts Computing: A Journal of Digital Humanities issn: 17538548 maglogo: N pubinfo: dt: Mar2024 vid: 18 iid: 1 pid: 2327 pub: Edinburgh University Press artinfo: ui: 176431506 10.3366/ijhac.2024.0324 ppf: 109 ppct: 16 formats: fmt: – @attributes: type: T – @attributes: type: P size: 284KB tig: atl: Enriching Portuguese Medieval Texts with Named Entity Recognition. aug: au: Inês Bico, Maria Baptista, Jorge Batista, Fernando Cardeira, Esperança su: Information retrieval Natural language processing Portuguese language Corpora Historical source material Knowledge base sug: subj: Information retrieval Natural language processing Portuguese language Corpora Historical source material Knowledge base keyword: corpus analysis information retrieval named entity disambiguation named entity linking natural language processing Portuguese medieval texts ab: Historical data poses unique challenges to natural language processing (NLP) and information retrieval (IR) tools, including digitization errors, lack of annotated data, and diachronic-specific issues. However, the increasing recognition of the value in historical documents has promoted efforts to semantically enrich and optimize their analysis. This article contributes to this endeavour by enriching the Corpus de Textos Antigos through NLP tools and techniques to enhance its usability and support research. The corpus undergoes linguistic annotation, including part-of-speech tagging, lemma annotation and named entity recognition (NER). Subsequently, the article delves into the tasks of entity disambiguation and entity linking, which involve identifying and disambiguating named entities by referring to a knowledge base (KB). Addressing the challenges posed by factors such as text state, epoch and the chosen KB, the article presents insights into related work, annotation results and the linguistic interest of a medieval annotated corpus for named entities. It concludes by discussing the challenges and providing avenues for future research in this domain. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Copyright of International Journal of Humanities & Arts Computing: A Journal of Digital Humanities is the property of Edinburgh University Press and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. item: International Journal of Humanities & Arts Computing: A Journal of Digital Humanities holder: Edinburgh University Press dt: @attributes: year: 2024 holdings: @attributes: islocal: N |
|---|