Active learning for clinical text classification: is it better than random sampling?

Objective: This study explores active learning algorithms as a way to reduce the requirements for large training sets in medical text classification tasks.Design: Three existing active learning algorithms (distance-based (DIST), diversity-based (DIV), and a combination of both (CMB)) were used to cl...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of the American Medical Informatics Association Vol. 19; no. 5; pp. 809 - 817
Autores principales: Figueroa, Rosa L, Zeng-Treitler, Qing, Ngo, Long H, Goryachev, Sergey, Wiechmann, Eduardo P
Formato: research Journal Article
Publicado: Oxford University Press / USA Sep2012
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=104360933&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 104360933
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        10675027
        FZ9
      jtl: Journal of the American Medical Informatics Association
      issn: 10675027
      maglogo: N
    pubinfo:
      dt: Sep2012
      vid: 19
      iid: 5
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        104360933
        NLM22707743
        2011646702
        10.1136/amiajnl-2011-000648
        NLM22707743
        PMC3422824
        104360933
      ppf: 809
      ppct: 8
      formats:
      tig:
        atl: Active learning for clinical text classification: is it better than random sampling?
      aug:
        au:
          Figueroa, Rosa L
          Zeng-Treitler, Qing
          Ngo, Long H
          Goryachev, Sergey
          Wiechmann, Eduardo P
        affil: Departamento de Ingeniería Eléctrica, Facultad de Ingeniería, Universidad de Concepción, Concepción, Chile.
      sug:
        subj:
          Data Mining Methods
          Natural Language Processing
          Algorithms
          Artificial Intelligence
          Human
          ROC Curve
      ab: Objective: This study explores active learning algorithms as a way to reduce the requirements for large training sets in medical text classification tasks.Design: Three existing active learning algorithms (distance-based (DIST), diversity-based (DIV), and a combination of both (CMB)) were used to classify text from five datasets. The performance of these algorithms was compared to that of passive learning on the five datasets. We then conducted a novel investigation of the interaction between dataset characteristics and the performance results.Measurements: Classification accuracy and area under receiver operating characteristics (ROC) curves for each algorithm at different sample sizes were generated. The performance of active learning algorithms was compared with that of passive learning using a weighted mean of paired differences. To determine why the performance varies on different datasets, we measured the diversity and uncertainty of each dataset using relative entropy and correlated the results with the performance differences.Results: The DIST and CMB algorithms performed better than passive learning. With a statistical significance level set at 0.05, DIST outperformed passive learning in all five datasets, while CMB was found to be better than passive learning in four datasets. We found strong correlations between the dataset diversity and the DIV performance, as well as the dataset uncertainty and the performance of the DIST algorithm.Conclusion: For medical text classification, appropriate active learning algorithms can yield performance comparable to that of passive learning with considerably smaller training sets. In particular, our results suggest that DIV performs better on data with higher diversity and DIST on data with lower uncertainty.
      pubtype: Academic Journal
      doctype:
        research
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N