Supervised Machine Learning Algorithms Can Classify Open-Text Feedback of Doctor Performance With Human-Level Accuracy.

Background: Machine learning techniques may be an effective and efficient way to classify open-text reports on doctor's activity for the purposes of quality assurance, safety, and continuing professional development.Objective: The objective of the study was to evaluate the accuracy of machine learni...

Full description

Bibliographic Details
Published in:Journal of Medical Internet Research Vol. 19; no. 3; pp. 1 - 2
Main Authors: Gibbons, Chris, Richards, Suzanne, Valderas, Jose Maria, Campbell, John
Format: research Journal Article
Published: JMIR Publications Inc. Mar2017
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=122323774&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 122323774
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        14394456
        DNC
      jtl: Journal of Medical Internet Research
      issn: 14394456
      maglogo: N
    pubinfo:
      dt: Mar2017
      vid: 19
      iid: 3
      pid: 21567
      pub: JMIR Publications Inc.
      place: Toronto, Ontario
    artinfo:
      ui:
        122323774
        122323774
        NLM28298265
        122323774
        10.2196/jmir.6533
        NLM28298265
        122323774
      ppf: 1
      ppct: 1
      formats:
      tig:
        atl: Supervised Machine Learning Algorithms Can Classify Open-Text Feedback of Doctor Performance With Human-Level Accuracy.
      aug:
        au:
          Gibbons, Chris
          Richards, Suzanne
          Valderas, Jose Maria
          Campbell, John
        affil: Centre for Health Services Research, University of Cambridge, Cambridge, United Kingdom
      sug:
        subj:
          Clinical Competence
          Algorithms
          Physicians Standards
          Feedback
          Human
          Questionnaires
      ab: Background: Machine learning techniques may be an effective and efficient way to classify open-text reports on doctor's activity for the purposes of quality assurance, safety, and continuing professional development.Objective: The objective of the study was to evaluate the accuracy of machine learning algorithms trained to classify open-text reports of doctor performance and to assess the potential for classifications to identify significant differences in doctors' professional performance in the United Kingdom.Methods: We used 1636 open-text comments (34,283 words) relating to the performance of 548 doctors collected from a survey of clinicians' colleagues using the General Medical Council Colleague Questionnaire (GMC-CQ). We coded 77.75% (1272/1636) of the comments into 5 global themes (innovation, interpersonal skills, popularity, professionalism, and respect) using a qualitative framework. We trained 8 machine learning algorithms to classify comments and assessed their performance using several training samples. We evaluated doctor performance using the GMC-CQ and compared scores between doctors with different classifications using t tests.Results: Individual algorithm performance was high (range F score=.68 to .83). Interrater agreement between the algorithms and the human coder was highest for codes relating to "popular" (recall=.97), "innovator" (recall=.98), and "respected" (recall=.87) codes and was lower for the "interpersonal" (recall=.80) and "professional" (recall=.82) codes. A 10-fold cross-validation demonstrated similar performance in each analysis. When combined together into an ensemble of multiple algorithms, mean human-computer interrater agreement was .88. Comments that were classified as "respected," "professional," and "interpersonal" related to higher doctor scores on the GMC-CQ compared with comments that were not classified (P<.05). Scores did not vary between doctors who were rated as popular or innovative and those who were not rated at all (P>.05).Conclusions: Machine learning algorithms can classify open-text feedback of doctor performance into multiple themes derived by human raters with high performance. Colleague open-text comments that signal respect, professionalism, and being interpersonal may be key indicators of doctor's performance.
      pubtype: Academic Journal
      doctype:
        research
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N