NEAT—Named Entities in Archaeological Texts: A semantic approach to term extraction and classification.
The lack of annotated datasets affects the development of Natural Language Processing applications and heavily impacts the access to textual data, in particular for specific domains and specific languages. In this paper, we propose a methodology to annotate texts concerning domain-specific knowledge...
| Published in: | Digital Scholarship in the Humanities Vol. 38; no. 3; pp. 997 - 1014 |
|---|---|
| Main Authors: | , , |
| Format: | Article |
| Published: |
Oxford University Press / USA
Sep2023
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=171389427&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 171389427 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Sep2023 vid: 38 iid: 3 pid: 622 pub: Oxford University Press / USA artinfo: ui: 171389427 10.1093/llc/fqad017 ppf: 997 ppct: 17 formats: fmt: – @attributes: type: T – @attributes: type: P size: 673KB tig: atl: NEAT—Named Entities in Archaeological Texts: A semantic approach to term extraction and classification. aug: au: Buono, Maria Pia di Nolano, Gennaro Monti, Johanna affil: UniOR NLP Research Group, Department of Literary, Linguistics and Comparative Studies , University of Naples "L'Orientale", Italy su: Classification Conceptual models sug: subj: Classification Conceptual models ab: The lack of annotated datasets affects the development of Natural Language Processing applications and heavily impacts the access to textual data, in particular for specific domains and specific languages. In this paper, we propose a methodology to annotate texts concerning domain-specific knowledge, to provide a reliable source of data for the task of Named Entity Recognition (NER) in the domain of archaeology for the Italian laguage. This method integrates syntactic and semantic information from several structured sources to annotate entities' mentions in unstructured texts. Furthermore, we make use of an ontology to label entities with the specific type they refer to. By using a corpus made up of item descriptions from Europeana's Archaeology Collection, we first test our proposed methodology on a mock dataset composed of 1,000 texts. After several steps of improvements, we use the final process to create a complete dataset composed of 5,000 descriptions. The resulting dataset, Named Entities in Archaeological Texts has a total of 41,002 spans of texts annotated with their domain-specific entity classification according to the CIDOC Conceptual Reference Model. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2023 holdings: @attributes: islocal: N |
|---|