Validity evidence supporting clinical skills assessment by artificial intelligence compared with trained clinician raters.

Background: Artificial intelligence (AI) is becoming increasingly used in medical education, but our understanding of the validity of AI‐based assessments (AIBA) as compared with traditional clinical expert‐based assessments (EBA) is limited. In this study, the authors aimed to compare and contrast...

Descripción completa

Detalles Bibliográficos
Publicado en:Medical Education Vol. 58; no. 1; pp. 105 - 118
Autores principales: Johnsson, Vilma, Søndergaard, Morten Bo, Kulasegaram, Kulamakan, Sundberg, Karin, Tiblad, Eleonor, Herling, Lotta, Petersen, Olav Bjørn, Tolsgaard, Martin G.
Formato: research tables/charts Journal Article
Publicado: Wiley-Blackwell Jan2024
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=174576527&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 174576527
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        03080110
        ESF
      jtl: Medical Education
      issn: 03080110
      maglogo: Y
    pubinfo:
      dt: Jan2024
      vid: 58
      iid: 1
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        174576527
        170072994
        174576527
        174576527
        10.1111/medu.15190
        174576527
      ppf: 105
      ppct: 13
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: C
          – @attributes:
              type: P
      tig:
        atl: Validity evidence supporting clinical skills assessment by artificial intelligence compared with trained clinician raters.
      aug:
        au:
          Johnsson, Vilma
          Søndergaard, Morten Bo
          Kulasegaram, Kulamakan
          Sundberg, Karin
          Tiblad, Eleonor
          Herling, Lotta
          Petersen, Olav Bjørn
          Tolsgaard, Martin G.
        affil: Center for Fetal Medicine, Department of Obstetrics, Copenhagen University Hospital, Rigshospitalet, Copenhagen, Denmark
      sug:
        subj:
          Competency Assessment
          Clinical Competence Evaluation
          Artificial Intelligence Utilization
          Sensitivity and Specificity
          Expert Clinicians
          Human
          Chorionic Villi Sampling
          Simulations
          Conceptual Framework
          Videorecording
          Neural Networks (Computer)
          Motion Capture
          Eye Movements
          Novice Clinicians
          Funding Source
      ab: Background: Artificial intelligence (AI) is becoming increasingly used in medical education, but our understanding of the validity of AI‐based assessments (AIBA) as compared with traditional clinical expert‐based assessments (EBA) is limited. In this study, the authors aimed to compare and contrast the validity evidence for the assessment of a complex clinical skill based on scores generated from an AI and trained clinical experts, respectively. Methods: The study was conducted between September 2020 to October 2022. The authors used Kane's validity framework to prioritise and organise their evidence according to the four inferences: scoring, generalisation, extrapolation and implications. The context of the study was chorionic villus sampling performed within the simulated setting. AIBA and EBA were used to evaluate performances of experts, intermediates and novice based on video recordings. The clinical experts used a scoring instrument developed in a previous international consensus study. The AI used convolutional neural networks for capturing features on video recordings, motion tracking and eye movements to arrive at a final composite score. Results: A total of 45 individuals participated in the study (22 novices, 12 intermediates and 11 experts). The authors demonstrated validity evidence for scoring, generalisation, extrapolation and implications for both EBA and AIBA. The plausibility of assumptions related to scoring, evidence of reproducibility and relation to different training levels was examined. Issues relating to construct underrepresentation, lack of explainability, and threats to robustness were identified as potential weak links in the AIBA validity argument compared with the EBA validity argument. Conclusion: There were weak links in the use of AIBA compared with EBA, mainly in their representation of the underlying construct but also regarding their explainability and ability to transfer to other datasets. However, combining AI and clinical expert‐based assessments may offer complementary benefits, which is a promising subject for future research. Is artificial intelligence (AI) ready to replace clinicians in clinical skills assessment? In this paper, Johnsson et al. use Kane's framework to compare validity evidence between the two.
      pubtype: Academic Journal
      doctype:
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N