Evaluation of Information Quality and Readability of Artificial Intelligence–Powered Chatbots in Systemic Isotretinoin Use.

Aim: This study aimed to evaluate and compare the readability and quality of information in responses generated by artificial intelligence (AI) models to patients' frequently asked questions about systemic isotretinoin, a medication commonly prescribed in dermatology. Materials and Methods: Thirty-f...

Full description

Bibliographic Details
Published in:Turkish Journal of Dermatology / Türk Dermatoloji Dergisi Vol. 20; no. 2; pp. 58 - 64
Main Authors: Koç, Huriye Aybüke, Özenir, Elif, Güney, Cansu Altınöz
Format: research tables/charts Journal Article
Published: Galenos Yayinevi Tic. LTD. STI 2026
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=194399982&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 194399982
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        13077635
        87WM
      jtl: Turkish Journal of Dermatology / Türk Dermatoloji Dergisi
      issn: 13077635
      maglogo: N
    pubinfo:
      dt: 2026
      vid: 20
      iid: 2
      pid: 28155
      pub: Galenos Yayinevi Tic. LTD. STI
    artinfo:
      ui:
        194399982
        194399982
        194399982
        10.4274/tjd.galenos.2026.53244
        194399982
      ppf: 58
      ppct: 6
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
      tig:
        atl: Evaluation of Information Quality and Readability of Artificial Intelligence–Powered Chatbots in Systemic Isotretinoin Use.
      aug:
        au:
          Koç, Huriye Aybüke
          Özenir, Elif
          Güney, Cansu Altınöz
        affil: Department of Dermatology and Venereology, Giresun University Faculty of Medicine, Giresun, Türkiye
      sug:
        subj:
          Data Quality Evaluation
          Isotretinoin Therapeutic Use
          Chatbot Utilization
          Readability Evaluation
          Consumer Health Information
          User-Computer Interface
          Human
          Artificial Intelligence
          Turkiye
          Professional-Patient Relations
          Clinical Assessment Tools
          Dermatologists
          Post Hoc Analysis
          Comparative Studies
          Data Analysis Software
          Descriptive Statistics
          Kruskal-Wallis Test
          Chi Square Test
          Dermatology
          Drug Information Services
          Patient Education
      ab: Aim: This study aimed to evaluate and compare the readability and quality of information in responses generated by artificial intelligence (AI) models to patients' frequently asked questions about systemic isotretinoin, a medication commonly prescribed in dermatology. Materials and Methods: Thirty-four frequently asked questions from patients using isotretinoin were prepared by a team of dermatology specialists. These questions were posed to three AI-based text-generation tools (ChatGPT, Gemini 2.0, and Copilot), and the responses were analyzed. The resulting texts were compared in terms of readability levels [Flesch Reading Ease score (FRES), Flesch-Kincaid Grade Level (FKGL), Simple Measure of Gobbledygook (SMOG), Gunning Fog index (GFOG), Coleman-Liau index (CLI), and automated readability index (ARI)], sentence lengths, and content quality, which was evaluated by dermatologists. Results: None of the AI models achieved the optimal readability threshold (FRES ≥ 60). Readability metrics differed significantly among models. Gemini produced responses that were significantly less readable and more complex than those produced by ChatGPT and Copilot across all readability indices, including FRES, FKGL, SMOG, GFOG, CLI, and ARI; post-hoc analyses confirmed differences between Gemini and the other models. Sentence counts also differed significantly, with Gemini generating longer responses than Copilot. In contrast, Likert-based quality scores and response appropriateness were comparable across models, with no statistically significant differences observed. Conclusion: This study demonstrates that AI models produce academic responses that are difficult for those unfamiliar with medical terminology to understand, and can generate outputs with variable readability in health-related content. These findings highlight the need for careful evaluation of AI-based content for use in healthcare.
      pubtype: Academic Journal
      doctype:
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N