Preclinical Dialogue Simulation: Evaluating Response Accessibility in Conversational Artificial Intelligence for Aphasia Therapy.

Purpose: Large language model (LLM)-driven conversational agents are increasingly considered for use in clinical contexts, yet systematic approaches for evaluating their behavior in impairment-rich, speech-based therapeutic interactions remain limited. This study extends the Agent-Based Conversation...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Speech, Language & Hearing Research Vol. 69; no. 8; pp. 3781 - 3797
Autores principales: Imaezue, Gerald C., Maram, Krishna V., Ajayi, David, Alohali, Isra, Butta, Rajesh K.
Formato: Artículo
Publicado: American Speech-Language-Hearing Association Aug2026
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=196185582&site=ehost-live
header:
  @attributes:
    shortDbName: ssf
    uiTerm: 196185582
    longDbName: Social Sciences Full Text (H.W. Wilson)
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        10924388
        1SM
      jtl: Journal of Speech, Language & Hearing Research
      issn: 10924388
      maglogo: N
    pubinfo:
      dt: Aug2026
      vid: 69
      iid: 8
      pid: 42
      pub: American Speech-Language-Hearing Association
    artinfo:
      ui:
        196185582
        10.1044/2026_JSLHR-26-00031
      ppf: 3781
      ppct: 16
      formats:
        fmt:
          @attributes:
            type: P
            size: 4.2MB
      tig:
        atl: Preclinical Dialogue Simulation: Evaluating Response Accessibility in Conversational Artificial Intelligence for Aphasia Therapy.
      aug:
        au:
          Imaezue, Gerald C.
          Maram, Krishna V.
          Ajayi, David
          Alohali, Isra
          Butta, Rajesh K.
        affil:
          Department of Communication Sciences and Disorders, University of South Florida, Tampa
          Department of Computer Science and Engineering, University of South Florida, Tampa
      su:
        Artificial intelligence
        Readability (Literary style)
        Communicative disorders
        Communication devices for people with disabilities
        Statistical models
        Rehabilitation of aphasic persons
        Benchmarking (Management)
        Research evaluation
        Natural language processing
        Speech-language pathology
        Simulation methods in education
        Data analysis software
        Speech therapy
        Regression analysis
        Sensitivity & specificity (Statistics)
      sug:
        subj:
          Artificial intelligence
          Readability (Literary style)
          Communicative disorders
          Communication devices for people with disabilities
          Statistical models
          Rehabilitation of aphasic persons
          Benchmarking (Management)
          Research evaluation
          Natural language processing
          Speech-language pathology
          Simulation methods in education
          Data analysis software
          Speech therapy
          Regression analysis
          Sensitivity & specificity (Statistics)
      ab: Purpose: Large language model (LLM)-driven conversational agents are increasingly considered for use in clinical contexts, yet systematic approaches for evaluating their behavior in impairment-rich, speech-based therapeutic interactions remain limited. This study extends the Agent-Based Conversational Dialogue (ABCD) simulation method as a preclinical testbed to evaluate how LLMs generate accessible clinician language when responding to characteristic aphasic speech during Response Elaboration Training. Method: ABCD was used to simulate multi-turn spoken therapeutic dialogues between an LLM-driven clinician and an artificial intelligence (AI)-simulated aphasic patient, enabling controlled manipulation of impairment profiles, prompting strategies, and reasoning modes without human participants. Three LLM families (Claude, GPT, and Gemini) were benchmarked under zero-shot and few-shot prompting and standard versus advanced reasoning. Response accessibility was quantified using established readability metrics (Flesch Reading Ease, Dale-Chall) and a composite score derived from 16 standardized readability measures. Results: Distinct accessibility signatures emerged across model architectures and configurations. Few-shot prompting and advanced reasoning generally yielded more accessible clinician responses, whereas Gemini demonstrated superior accessibility under zero-shot, standard reasoning. Conclusions: LLMs differ systematically in their capacity to adapt clinician language to impaired speech. ABCD provides a scalable, preclinical dialogue simulation framework for benchmarking conversational AI in clinically oriented, impairment-rich dialogue across multidimensional constraints. It offers guidance for model selection and configuration prior to clinical translation in communication rehabilitation.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: N
    holdings:
      @attributes:
        islocal: N