Patient Experiences in the Cochlear Implant Reddit Community: Comparing Human and Large Language Model Categorization.

Purpose: Although some work has leveraged automated analyses of online communities to gain cochlear implant (CI) patient insights, there remains a gap in comparing human versus automated analysis of the nuanced, real-world experiences patients share outside clinical settings. This study characterize...

Full description

Bibliographic Details
Published in:American Journal of Audiology Vol. 35; no. 2; pp. 487 - 497
Main Authors: Habib, Daniel R. S., Depala, Kiran, Lin, Jack, Le, Samuel, McFall, Natalie, Dewan, Shiv S., Huang, Justin, Habib, Michael W. S., Bishay, Anthony E., Siebor, Konrad, Babaoglu, Gizem, Chowdhury, Naweed I., Moberly, Aaron C.
Format: research tables/charts Journal Article
Published: American Speech-Language-Hearing Association Jun2026
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=194359724&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 194359724
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        10590889
        5KS
      jtl: American Journal of Audiology
      issn: 10590889
      maglogo: N
    pubinfo:
      dt: Jun2026
      vid: 35
      iid: 2
      pid: 42
      pub: American Speech-Language-Hearing Association
      place: Rockville, Maryland
    artinfo:
      ui:
        194359724
        194359724
        194359724
        10.1044/2025_AJA-25-00216
        194359724
      ppf: 487
      ppct: 10
      formats:
        fmt:
          @attributes:
            type: P
      tig:
        atl: Patient Experiences in the Cochlear Implant Reddit Community: Comparing Human and Large Language Model Categorization.
      aug:
        au:
          Habib, Daniel R. S.
          Depala, Kiran
          Lin, Jack
          Le, Samuel
          McFall, Natalie
          Dewan, Shiv S.
          Huang, Justin
          Habib, Michael W. S.
          Bishay, Anthony E.
          Siebor, Konrad
          Babaoglu, Gizem
          Chowdhury, Naweed I.
          Moberly, Aaron C.
        affil: School of Medicine, Vanderbilt University, Nashville, TN
      sug:
        subj:
          Social Media
          Cochlear Implant
          Health Information
          Patient Attitudes Evaluation
          Natural Language Processing
          Sensitivity and Specificity
          Human
          Qualitative Studies
          Comparative Studies
          Thematic Analysis
          Coding
          kappa Statistic
          Predictive Value of Tests
          Support, Social
          Descriptive Statistics
          Interrater Reliability
          Coefficient alpha
          Data Analysis Software
          Automation
          Artificial Intelligence, Generative
          Community Role
      ab: Purpose: Although some work has leveraged automated analyses of online communities to gain cochlear implant (CI) patient insights, there remains a gap in comparing human versus automated analysis of the nuanced, real-world experiences patients share outside clinical settings. This study characterizes experiences within the r/Cochlearimplants Reddit community and compares human to large language model (LLM) performance in annotating posts. Method: Using reflexive thematic analysis, 996 publicly available r/Cochlearimplants posts (October 2024--June 2025) were manually coded and consolidated into themes. Three LLMs--OpenAI o3, Gemini 2.5 Pro, and Claude Sonnet 4--were prompted with the posts and human-generated codebook to perform post categorization. Model performance was evaluated against human coding using Cohen's kappa, percent agreement, sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and time. Results: Five themes emerged. Community engagement and support (N = 944, 94.8%) frequently involved eliciting advice (N = 721, 72.4%), seeking shared experiences (N = 249, 25.0%), and sharing negative experiences (N = 247, 24.8%). Other themes included the medical/surgical journey (N = 463, 46.5%), device/ technical issues (N = 343, 34.4%), daily life/adjustments (N = 236, 23.7%), and media/outreach (7.2%, N = 72). OpenAI o3 and Gemini 2.5 Pro achieved the highest interrater reliability with human annotators (Κ = .35 and Κ = .34, respectively). OpenAI o3 had higher sensitivity (46.7%) but lower specificity (90.4%) than Gemini 2.5 Pro, which had the highest specificity (93.4%) but lower sensitivity (38.0%). Claude Sonnet 4 showed the lowest agreement (Κ = .25) and PPV (30.9%). Compared to human annotation requiring 52 hr across all annotators, each LLM required less than 20 min. Conclusions: Reddit posts revealed rich discourse across CI topics. LLMs demonstrated fair agreement with human coders and can quickly aid in large-scale qualitative analysis. Although careful model selection and human expertise remain essential for accurate interpretation, LLM annotation shows potential for real-time monitoring of patient concerns to inform counseling, rehabilitation strategies, and iterative device design.
      pubtype: Academic Journal
      doctype:
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N