Identifying medical terms in patient-authored text: a crowdsourcing-based approach.

Background and Objective: As people increasingly engage in online health-seeking behavior and contribute to health-oriented websites, the volume of medical text authored by patients and other medical novices grows rapidly. However, we lack an effective method for automatically identifying medical te...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of the American Medical Informatics Association Vol. 20; no. 6; pp. 1120 - 1128
Autores principales: Maclean, Diana Lynn, Heer, Jeffrey
Formato: research Journal Article
Publicado: Oxford University Press / USA Nov2013
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=104103799&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 104103799
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        10675027
        FZ9
      jtl: Journal of the American Medical Informatics Association
      issn: 10675027
      maglogo: N
    pubinfo:
      dt: Nov2013
      vid: 20
      iid: 6
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        104103799
        NLM23645553
        2012338842
        10.1136/amiajnl-2012-001110
        NLM23645553
        PMC3822103
        104103799
      ppf: 1120
      ppct: 8
      formats:
      tig:
        atl: Identifying medical terms in patient-authored text: a crowdsourcing-based approach.
      aug:
        au:
          Maclean, Diana Lynn
          Heer, Jeffrey
        affil: Department of Computer Science, Stanford University, Stanford, California, USA.
      sug:
        subj:
          Consumer Health Information
          Crowdsourcing
          Nomenclature
          Resource Databases
          Human
          Information Retrieval
          Nurses
      ab: Background and Objective: As people increasingly engage in online health-seeking behavior and contribute to health-oriented websites, the volume of medical text authored by patients and other medical novices grows rapidly. However, we lack an effective method for automatically identifying medical terms in patient-authored text (PAT). We demonstrate that crowdsourcing PAT medical term identification tasks to non-experts is a viable method for creating large, accurately-labeled PAT datasets; moreover, such datasets can be used to train classifiers that outperform existing medical term identification tools.Materials and Methods: To evaluate the viability of using non-expert crowds to label PAT, we compare expert (registered nurses) and non-expert (Amazon Mechanical Turk workers; Turkers) responses to a PAT medical term identification task. Next, we build a crowd-labeled dataset comprising 10 000 sentences from MedHelp. We train two models on this dataset and evaluate their performance, as well as that of MetaMap, Open Biomedical Annotator (OBA), and NaCTeM's TerMINE, against two gold standard datasets: one from MedHelp and the other from CureTogether.Results: When aggregated according to a corroborative voting policy, Turker responses predict expert responses with an F1 score of 84%. A conditional random field (CRF) trained on 10 000 crowd-labeled MedHelp sentences achieves an F1 score of 78% against the CureTogether gold standard, widely outperforming OBA (47%), TerMINE (43%), and MetaMap (39%). A failure analysis of the CRF suggests that misclassified terms are likely to be either generic or rare.Conclusions: Our results show that combining statistical models sensitive to sentence-level context with crowd-labeled data is a scalable and effective technique for automatically identifying medical terms in PAT.
      pubtype: Academic Journal
      doctype:
        research
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N