Human–machine interaction in building an English reference dataset for natural language processing tasks.

Rich in information and annotated instances, a reference annotated dataset is essential for the training and evaluation of Natural Language Processing (NLP) tools. However, the creation of such linguistic resources is a tedious and time-consuming task involving lexical, syntactic, and semantic annot...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 59; no. 3; pp. 2781 - 2810
Autores principales: Žitko, Branko, Gašpar, Angelina, Bročić, Lucija, Vasić, Daniel, Grubišić, Ani
Formato: Artículo
Publicado: Springer Nature Sep2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909091&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 186909091
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2025
      vid: 59
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        186909091
        10.1007/s10579-025-09835-2
      ppf: 2781
      ppct: 29
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.1MB
      tig:
        atl: Human–machine interaction in building an English reference dataset for natural language processing tasks.
      aug:
        au:
          Žitko, Branko
          Gašpar, Angelina
          Bročić, Lucija
          Vasić, Daniel
          Grubišić, Ani
        affil:
          https://ror.org/00m31ft63 Faculty of Science, University of Split, Ruđera Boškovića 33, 21000, Split, Croatia
          https://ror.org/00m31ft63 Catholic Faculty of Theology, University of Split, Zrinsko Frankopanska 19, 21000, Split, Croatia
          https://ror.org/00v89p354 Faculty of Science and Education, University of Mostar, Matice hrvatske b.b., 88000, Mostar, Bosnia and Herzegovina
      su:
        Natural language processing
        Annotations
        Human-machine systems
        Data extraction
      sug:
        subj:
          Natural language processing
          Annotations
          Human-machine systems
          Data extraction
      keyword:
        Artificial intelligence
        CCS concepts
        Communication and Culture Linguistics Information and Computing Sciences Artificial Intelligence and Image Processing Psychology and Cognitive Sciences Cognitive Sciences
        Computing methodologies
        Language
      ab: Rich in information and annotated instances, a reference annotated dataset is essential for the training and evaluation of Natural Language Processing (NLP) tools. However, the creation of such linguistic resources is a tedious and time-consuming task involving lexical, syntactic, and semantic annotations, typically at the sentence level. Assuming we could speed up the human annotation process, we employed pre-trained models (spaCy, AllenNLP, EWISER) to automatically annotate a dataset of 664 sentences (6853 tokens, including 1598 predicates) taken from grammar books. A multi-layered annotation task encompassed Lemmatization (LEM), Part-of-Speech Tagging (UPOS, XPOS), Named Entity Recognition (NER), Dependency Parsing (DEP, HEAD), Coreference Resolution (COREF), Semantic Role Labelling (SRL), Predicate Sense Disambiguation (PSD) and Word Sense Disambiguation (WSD). Three annotators post-edited the noisy automatic annotations, and their average Inter-Annotator Agreement (IAA) for all annotation tasks at the token level was 0.91 and at the sentence level 0.74. Evaluation metrics including Accuracy, Precision, Recall, and F1 revealed disparities between machine and human annotations, along with correlations between machine annotations at both token and sentence levels. Manual error analysis identified instances where NLP tools failed to generate accurate annotations. A comparison of time spent per layer revealed that refining a pre-annotated subset of sentences required significantly less time than annotating them manually from scratch. This process resulted in an English reference dataset, tailored for the development of a hypergraph-based knowledge extraction model, known as the Natural Language 2 Semantic Hyper-graph Dataset (NL2SH) 1.0), which is accessible through CLARIN.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N