Reliability and Readability Assessment of Atrial Fibrillation Patient Information Delivered by Artificial Intelligence-Based Language Models (ChatGPT, YouChat, Gemini, and Perplexity AI) in English and Spanish.

Background: Atrial fibrillation (AF) is the most prevalent arrhythmia and a significant cause of morbidity. Artificial intelligence (AI)-based language models represent a novel tool for searching for medical information; however, there is still uncertainty regarding their reliability and readability...

Descripción completa

Detalles Bibliográficos
Publicado en:Clinical Medicine Insights: Cardiology Vol. 19; pp. 1 - 11
Autores principales: Juan-Guardela, Emilio Jose Juan, Beltrán-España, Jesús Andrés, Ravagli-Baquero, María Paula, Porras-Bueno, Cristian Orlando, Cáceres-Méndez, Edward, Ávila, Daniel Fernandez, Muñoz-Velandia, Oscar, García-Peña, Ángel Alberto
Formato: research tables/charts Journal Article
Publicado: Sage Publications Inc. 10/31/2025
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=189212351&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 189212351
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        11795468
        B3KL
      jtl: Clinical Medicine Insights: Cardiology
      issn: 11795468
      maglogo: Y
    pubinfo:
      dt: 10/31/2025
      vid: 19
      pid: 344
      pub: Sage Publications Inc.
      place: Thousand Oaks, California
    artinfo:
      ui:
        189212351
        189212351
        189212351
        10.1177/11795468251383666
        189212351
      ppf: 1
      ppct: 10
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
      tig:
        atl: Reliability and Readability Assessment of Atrial Fibrillation Patient Information Delivered by Artificial Intelligence-Based Language Models (ChatGPT, YouChat, Gemini, and Perplexity AI) in English and Spanish.
      aug:
        au:
          Juan-Guardela, Emilio Jose Juan
          Beltrán-España, Jesús Andrés
          Ravagli-Baquero, María Paula
          Porras-Bueno, Cristian Orlando
          Cáceres-Méndez, Edward
          Ávila, Daniel Fernandez
          Muñoz-Velandia, Oscar
          García-Peña, Ángel Alberto
        affil: Department of Internal Medicine, Pontificia Universidad Javeriana, Bogota, Colombia
      sug:
        subj:
          Artificial Intelligence, Generative
          Atrial Fibrillation
          Consumer Health Information Evaluation
          Reliability Evaluation
          Readability Evaluation
          Human
          Cross Sectional Studies
          Analytic Research
          Nonexperimental Studies
          Spanish Language
          English Language
          Interrater Reliability
          Descriptive Statistics
          Fisher's Exact Test
          Mann-Whitney U Test
          Kruskal-Wallis Test
          Data Analysis Software
          Scales
      ab: Background: Atrial fibrillation (AF) is the most prevalent arrhythmia and a significant cause of morbidity. Artificial intelligence (AI)-based language models represent a novel tool for searching for medical information; however, there is still uncertainty regarding their reliability and readability in different languages. Objective: To assess the reliability and readability of information provided by AI-based models for patients with AF. Methods: A cross-sectional study was conducted to assess the reliability and readability of the responses generated by ChatGPT, YouChat, Gemini and Perplexity on AF in English and Spanish. Thirty standardised questions were posed in both languages. The quality of the responses was then assessed by 2 independent reviewers via a standardised tool. Readability was assessed via the Flesch–Szigrist formula. The results were then compared by tool and language. Results: ChatGPT demonstrated the highest interrater agreement (PA = 0.73 in Spanish, 0.80 in English), followed by Gemini in English (PA = 0.66). In Spanish, ChatGPT generated the highest percentage of complete responses (80%), followed by Perplexity (73%) and Gemini (47%). In English, Perplexity demonstrated the strongest performance, with a score of 93%, followed by ChatGPT, with 73%, and Gemini, with 53%. A readability analysis revealed significant differences between the models (P <.01). The ChatGPT demonstrated the highest performance, although its content was moderately challenging in Spanish and highly challenging in English. Conclusion: ChatGPT and Perplexity emerged as the most reliable models, although readability remains a concern. There is a clear need for improvements to optimise the accuracy and accessibility of AI-generated medical information.
      pubtype: Academic Journal
      doctype:
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N