Making manual scoring of typed transcripts a thing of the past: a commentary on Herrmann (2025).

Coding the accuracy of typed transcripts from experiments testing speech intelligibility is an arduous endeavour. A recent study in this journal [Herrmann, B. 2025. Leveraging natural language processing models to automate speech-intelligibility scoring. Speech, Language and Hearing, 28(1)] presents...

Descripción completa

Detalles Bibliográficos
Publicado en:Speech, Language & Hearing Vol. 28; no. 1; pp. 1 - 4
Autor principal: Bosker, Hans Rutger
Formato: commentary Journal Article
Publicado: Taylor & Francis Ltd Dec2025
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=190352204&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 190352204
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        2050571X
        FKDR
      jtl: Speech, Language & Hearing
      issn: 2050571X
      maglogo: N
    pubinfo:
      dt: Dec2025
      vid: 28
      iid: 1
      pid: 377
      pub: Taylor & Francis Ltd
      place: Philadelphia, Pennsylvania
    artinfo:
      ui:
        190352204
        190352204
        190352204
        10.1080/2050571X.2025.2514395
        190352204
      ppf: 1
      ppct: 3
      formats:
      tig:
        atl: Making manual scoring of typed transcripts a thing of the past: a commentary on Herrmann (2025).
      aug:
        au: Bosker, Hans Rutger
        affil: Donders Institute for Brain, Cognition and Behaviour, Radboud University, Nijmegen, the Netherlands
      sug:
        subj:
          Speech Intelligibility
          Natural Language Processing
          Automation
      ab: Coding the accuracy of typed transcripts from experiments testing speech intelligibility is an arduous endeavour. A recent study in this journal [Herrmann, B. 2025. Leveraging natural language processing models to automate speech-intelligibility scoring. Speech, Language and Hearing, 28(1)] presents a novel approach for automating the scoring of such listener transcripts, leveraging Natural Language Processing (NLP) models. It involves the calculation of the semantic similarity between transcripts and target sentences using high-dimensional vectors, generated by such NLP models as ADA2, GPT2, BERT, and USE. This approach demonstrates exceptional accuracy, with negligible underestimation of intelligibility scores (by about 2-4%), numerically outperforming simpler computational tools like Autoscore and TSR. The method uniquely relies on semantic representations generated by large language models. At the same time, these models also form the Achilles heel of the technique: the transparency, accessibility, data security, ethical framework, and cost of the selected model directly impact the suitability of the NLP-based scoring method. Hence, working with such models can raise serious risks regarding the reproducibility of scientific findings. This in turn emphasises the need for fair, ethical, and evidence-based open source models. With such models, Herrmann's new tool represents a valuable addition to the speech scientist's toolbox.
      pubtype: Academic Journal
      doctype:
        commentary
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N