Human–machine interaction in building an English reference dataset for natural language processing tasks.
Rich in information and annotated instances, a reference annotated dataset is essential for the training and evaluation of Natural Language Processing (NLP) tools. However, the creation of such linguistic resources is a tedious and time-consuming task involving lexical, syntactic, and semantic annot...
| Publicado en: | Language Resources & Evaluation Vol. 59; no. 3; pp. 2781 - 2810 |
|---|---|
| Autores principales: | , , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Sep2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909091&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 186909091 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2025 vid: 59 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 186909091 10.1007/s10579-025-09835-2 ppf: 2781 ppct: 29 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.1MB tig: atl: Human–machine interaction in building an English reference dataset for natural language processing tasks. aug: au: Žitko, Branko Gašpar, Angelina Bročić, Lucija Vasić, Daniel Grubišić, Ani affil: https://ror.org/00m31ft63 Faculty of Science, University of Split, Ruđera Boškovića 33, 21000, Split, Croatia https://ror.org/00m31ft63 Catholic Faculty of Theology, University of Split, Zrinsko Frankopanska 19, 21000, Split, Croatia https://ror.org/00v89p354 Faculty of Science and Education, University of Mostar, Matice hrvatske b.b., 88000, Mostar, Bosnia and Herzegovina su: Natural language processing Annotations Human-machine systems Data extraction sug: subj: Natural language processing Annotations Human-machine systems Data extraction keyword: Artificial intelligence CCS concepts Communication and Culture Linguistics Information and Computing Sciences Artificial Intelligence and Image Processing Psychology and Cognitive Sciences Cognitive Sciences Computing methodologies Language ab: Rich in information and annotated instances, a reference annotated dataset is essential for the training and evaluation of Natural Language Processing (NLP) tools. However, the creation of such linguistic resources is a tedious and time-consuming task involving lexical, syntactic, and semantic annotations, typically at the sentence level. Assuming we could speed up the human annotation process, we employed pre-trained models (spaCy, AllenNLP, EWISER) to automatically annotate a dataset of 664 sentences (6853 tokens, including 1598 predicates) taken from grammar books. A multi-layered annotation task encompassed Lemmatization (LEM), Part-of-Speech Tagging (UPOS, XPOS), Named Entity Recognition (NER), Dependency Parsing (DEP, HEAD), Coreference Resolution (COREF), Semantic Role Labelling (SRL), Predicate Sense Disambiguation (PSD) and Word Sense Disambiguation (WSD). Three annotators post-edited the noisy automatic annotations, and their average Inter-Annotator Agreement (IAA) for all annotation tasks at the token level was 0.91 and at the sentence level 0.74. Evaluation metrics including Accuracy, Precision, Recall, and F1 revealed disparities between machine and human annotations, along with correlations between machine annotations at both token and sentence levels. Manual error analysis identified instances where NLP tools failed to generate accurate annotations. A comparison of time spent per layer revealed that refining a pre-annotated subset of sentences required significantly less time than annotating them manually from scratch. This process resulted in an English reference dataset, tailored for the development of a hypergraph-based knowledge extraction model, known as the Natural Language 2 Semantic Hyper-graph Dataset (NL2SH) 1.0), which is accessible through CLARIN. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|