Automating Intended Target Identification for Paraphasias in Discourse Using a Large Language Model.

Purpose: To date, there are no automated tools for the identification and finegrained classification of paraphasias within discourse, the production of which is the hallmark characteristic of most people with aphasia (PWA). In this work, we fine-tune a large language model (LLM) to automatically pre...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Speech, Language & Hearing Research Vol. 66; no. 12; pp. 4949 - 4967
Autores principales: Salem, Alexandra C., Gale, Robert C., Fleegle, Mikala, Fergadiotis, Gerasimos, Bedrick, Steven
Formato: Artículo
Publicado: American Speech-Language-Hearing Association Dec2023
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=174209929&site=ehost-live
header:
  @attributes:
    shortDbName: ssf
    uiTerm: 174209929
    longDbName: Social Sciences Full Text (H.W. Wilson)
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        10924388
        1SM
      jtl: Journal of Speech, Language & Hearing Research
      issn: 10924388
      maglogo: N
    pubinfo:
      dt: Dec2023
      vid: 66
      iid: 12
      pid: 42
      pub: American Speech-Language-Hearing Association
    artinfo:
      ui:
        174209929
        10.1044/2023_JSLHR-23-00121
      ppf: 4949
      ppct: 18
      formats:
        fmt:
          @attributes:
            type: P
            size: 1MB
      tig:
        atl: Automating Intended Target Identification for Paraphasias in Discourse Using a Large Language Model.
      aug:
        au:
          Salem, Alexandra C.
          Gale, Robert C.
          Fleegle, Mikala
          Fergadiotis, Gerasimos
          Bedrick, Steven
        affil:
          Department of Medical Informatics and Clinical Epidemiology, Oregon Health & Science University, Portland
          Department of Speech & Hearing Sciences, Portland State University, OR
      su:
        Linguistics
        Language & languages
        Storytelling
        Diagnosis of aphasia
        Machine learning
        Comparative studies
        Automation
        Descriptive statistics
        Research funding
      sug:
        subj:
          Linguistics
          Language & languages
          Storytelling
          Diagnosis of aphasia
          Machine learning
          Comparative studies
          Automation
          Descriptive statistics
          Research funding
      ab: Purpose: To date, there are no automated tools for the identification and finegrained classification of paraphasias within discourse, the production of which is the hallmark characteristic of most people with aphasia (PWA). In this work, we fine-tune a large language model (LLM) to automatically predict paraphasia targets in Cinderella story retellings. Method: Data consisted of 332 Cinderella story retellings containing 2,489 paraphasias from PWA, for which research assistants identified their intended targets. We supplemented these training data with 256 sessions from control participants, to which we added 2,415 synthetic paraphasias. We conducted four experiments using different training data configurations to fine-tune the LLM to automatically "fill in the blank" of the paraphasia with a predicted target, given the context of the rest of the story retelling. We tested the experiments' predictions against our human-identified targets and stratified our results by ambiguity of the targets and clinical factors. Results: The model trained on controls and PWA achieved 50.7% accuracy at exactly matching the human-identified target. Fine-tuning on PWA data, with or without controls, led to comparable performance. The model performed better on targets with less human ambiguity and on paraphasias from participants with fluent or less severe aphasia. Conclusions: We were able to automatically identify the intended target of paraphasias in discourse using just the surrounding language about half of the time. These findings take us a step closer to automatic aphasic discourse analysis. In future work, we will incorporate phonological information from the paraphasia to further improve predictive utility. Supplemental Material: https://doi.org/10.23641/asha.24463543
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: N
    holdings:
      @attributes:
        islocal: N