Using Machine Learning to Explain Paraphasias in Narratives of People With Aphasia.

Purpose: This study examines how personal, clinical, and word-level features explain paraphasias when using machine learning--based error analysis on the narratives of people with aphasia (PWA). Method: We used AphasiaBank as the source of narrative transcript data for 236 PWA. We tested machine lea...

Descripción completa

Detalles Bibliográficos
Publicado en:Perspectives of the ASHA Special Interest Groups Vol. 10; no. 2; pp. 451 - 463
Autores principales: Zavaleta, Rosa, Brue, Jacob, Sen, Sandip, Wilson, Laura
Formato: research tables/charts Journal Article
Publicado: American Speech-Language-Hearing Association Apr2025
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=184219911&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 184219911
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        2381473X
        KTSD
      jtl: Perspectives of the ASHA Special Interest Groups
      issn: 2381473X
      maglogo: N
    pubinfo:
      dt: Apr2025
      vid: 10
      iid: 2
      pid: 42
      pub: American Speech-Language-Hearing Association
      place: Rockville, Maryland
    artinfo:
      ui:
        184219911
        184219911
        184219911
        10.1044/2024_PERSP-23-00291
        184219911
      ppf: 451
      ppct: 12
      formats:
        fmt:
          @attributes:
            type: P
      tig:
        atl: Using Machine Learning to Explain Paraphasias in Narratives of People With Aphasia.
      aug:
        au:
          Zavaleta, Rosa
          Brue, Jacob
          Sen, Sandip
          Wilson, Laura
        affil: Department of Communication Sciences & Disorders, The University of Tulsa, OK
      sug:
        subj:
          Aphasia Therapy
          Machine Learning Utilization
          Speech Production Measurement Methods
          Language Disorders Therapy
          Speech Disorders Therapy
          Narratives
          Human
          Funding Source
          Male
          Female
          Middle Age
          Models, Theoretical
          Natural Language Processing
          Algorithms
          Random Forest
          Apraxia
          Communication
          Speech Physiology
          Sensitivity and Specificity
          Descriptive Statistics
          Comparative Studies
          ROC Curve
          Questionnaires
          Middle Aged: 45-64 years
          Male
          Female
      ab: Purpose: This study examines how personal, clinical, and word-level features explain paraphasias when using machine learning--based error analysis on the narratives of people with aphasia (PWA). Method: We used AphasiaBank as the source of narrative transcript data for 236 PWA. We tested machine learning classification algorithms including decision trees and random forests on the utterances of PWA, including nonparaphasic words and intended words when paraphasias were produced. We classified target words as paraphasic or nonparaphasic based on PWA's age; aphasia severity, duration, and type; presence of apraxia or dysarthria; and word-level features including part of speech, word frequency, imageability, syllable count, and location in the utterance. We measured the models' predictive accuracy across classification thresholds on held-out test sets, and we used feature analysis to compare feature importance. Results: At the word level, our random forest model achieved an area under curve (AUC) of 0.896. We found a sensitivity of 0.821 for semantic paraphasias, 0.764 for phonemic paraphasias, and 0.872 for neologistic paraphasias. The most salient features, in order of importance, were word frequency, imageability, part of speech, age, severity, and syllable count, followed by aphasia duration, location of word, presence of apraxia, type of aphasia (e.g., fluent), and presence of dysarthria. Our random forest model that included information about surrounding words achieved AUC scores ranging from 0.881 to 0.899. Additionally, we developed a model that was trained on surrounding words and their respective features, but not given the actual error word. The best model achieved an AUC of 0.745. Conclusions: Machine learning can aid in the explanation of paraphasias. In this study, we analyzed word- and person-level features and highlighted the nonrandom nature of paraphasic productions. Furthermore, this lays the groundwork for developing machine learning models with clinical applications at the various stages of treatment of PWA.
      pubtype: Academic Journal
      doctype:
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N