Developing a video‐based method to compare and adjust examiner effects in fully nested OSCEs.

Background: Although averaging across multiple examiners' judgements reduces unwanted overall score variability in objective structured clinical examinations (OSCE), designs involving several parallel circuits of the OSCE require that different examiner cohorts collectively judge performances to the...

Descripción completa

Detalles Bibliográficos
Publicado en:Medical Education Vol. 53; no. 3; pp. 250 - 264
Autores principales: Yeates, Peter, Cope, Natalie, Hawarden, Ashley, Bradshaw, Hannah, McCray, Gareth, Homer, Matt
Formato: research tables/charts Journal Article
Publicado: Wiley-Blackwell Mar2019
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=134737302&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 134737302
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        03080110
        ESF
      jtl: Medical Education
      issn: 03080110
      maglogo: Y
    pubinfo:
      dt: Mar2019
      vid: 53
      iid: 3
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        134737302
        134737302
        134737302
        10.1111/medu.13783
        134737302
      ppf: 250
      ppct: 14
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: C
          – @attributes:
              type: P
      tig:
        atl: Developing a video‐based method to compare and adjust examiner effects in fully nested OSCEs.
      aug:
        au:
          Yeates, Peter
          Cope, Natalie
          Hawarden, Ashley
          Bradshaw, Hannah
          McCray, Gareth
          Homer, Matt
        affil: Medical School Education Research Group (MERG), Keele University School of Medicine, Keele UK
      sug:
        subj:
          Competency Assessment
          Videorecording
          Students, Medical Evaluation
          Human
          Rasch Analysis
          Error of Leniency
          Descriptive Statistics
          Aptitude
          Reliability
          Prospective Studies
          Sampling Methods
      ab: Background: Although averaging across multiple examiners' judgements reduces unwanted overall score variability in objective structured clinical examinations (OSCE), designs involving several parallel circuits of the OSCE require that different examiner cohorts collectively judge performances to the same standard in order to avoid bias. Prior research suggests the potential for important examiner‐cohort effects in distributed or national examinations that could compromise fairness or patient safety, but despite their importance, these effects are rarely investigated because fully nested assessment designs make them very difficult to study. We describe initial use of a new method to measure and adjust for examiner‐cohort effects on students' scores. Methods: We developed video‐based examiner score comparison and adjustment (VESCA): volunteer students were filmed 'live' on 10 out of 12 OSCE stations. Following the examination, examiners additionally scored station‐specific common‐comparator videos, producing partial crossing between examiner cohorts. Many‐facet Rasch modelling and linear mixed modelling were used to estimate and adjust for examiner‐cohort effects on students' scores. Results: After accounting for students' ability, examiner cohorts differed substantially in their stringency or leniency (maximal global score difference of 0.47 out of 7.0 [Cohen's d = 0.96]; maximal total percentage score difference of 5.7% [Cohen's d = 1.06] for the same student ability by different examiner cohorts). Corresponding adjustment of students' global and total percentage scores altered the theoretical classification of 6.0% of students for both measures (either pass to fail or fail to pass), whereas 8.6–9.5% students' scores were altered by at least 0.5 standard deviations of student ability. Conclusions: Despite typical reliability, the examiner cohort that students encountered had a potentially important influence on their score, emphasising the need for adequate sampling and examiner training. Development and validation of VESCA may offer a means to measure and adjust for potential systematic differences in scoring patterns that could exist between locations in distributed or national OSCE examinations, thereby ensuring equivalence and fairness. Finding that scores by different groups of examiners can differ by a whole standard deviation of student ability, the authors offer a video‐based method to address examiner‐cohort effects in OSCEs.
      pubtype: Academic Journal
      doctype:
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N