NEAT—Named Entities in Archaeological Texts: A semantic approach to term extraction and classification.

The lack of annotated datasets affects the development of Natural Language Processing applications and heavily impacts the access to textual data, in particular for specific domains and specific languages. In this paper, we propose a methodology to annotate texts concerning domain-specific knowledge...

Full description

Bibliographic Details
Published in:Digital Scholarship in the Humanities Vol. 38; no. 3; pp. 997 - 1014
Main Authors: Buono, Maria Pia di, Nolano, Gennaro, Monti, Johanna
Format: Article
Published: Oxford University Press / USA Sep2023
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=171389427&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 171389427
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Sep2023
      vid: 38
      iid: 3
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        171389427
        10.1093/llc/fqad017
      ppf: 997
      ppct: 17
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 673KB
      tig:
        atl: NEAT—Named Entities in Archaeological Texts: A semantic approach to term extraction and classification.
      aug:
        au:
          Buono, Maria Pia di
          Nolano, Gennaro
          Monti, Johanna
        affil: UniOR NLP Research Group, Department of Literary, Linguistics and Comparative Studies , University of Naples "L'Orientale", Italy
      su:
        Classification
        Conceptual models
      sug:
        subj:
          Classification
          Conceptual models
      ab: The lack of annotated datasets affects the development of Natural Language Processing applications and heavily impacts the access to textual data, in particular for specific domains and specific languages. In this paper, we propose a methodology to annotate texts concerning domain-specific knowledge, to provide a reliable source of data for the task of Named Entity Recognition (NER) in the domain of archaeology for the Italian laguage. This method integrates syntactic and semantic information from several structured sources to annotate entities' mentions in unstructured texts. Furthermore, we make use of an ontology to label entities with the specific type they refer to. By using a corpus made up of item descriptions from Europeana's Archaeology Collection, we first test our proposed methodology on a mock dataset composed of 1,000 texts. After several steps of improvements, we use the final process to create a complete dataset composed of 5,000 descriptions. The resulting dataset, Named Entities in Archaeological Texts has a total of 41,002 spans of texts annotated with their domain-specific entity classification according to the CIDOC Conceptual Reference Model.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2023
    holdings:
      @attributes:
        islocal: N