Toward better public health reporting using existing off the shelf approaches: The value of medical dictionaries in automated cancer detection using plaintext medical data.
Objectives: Existing approaches to derive decision models from plaintext clinical data frequently depend on medical dictionaries as the sources of potential features. Prior research suggests that decision models developed using non-dictionary based feature sourcing approaches and "off the shelf" too...
| Publicado en: | Journal of Biomedical Informatics Vol. 69; pp. 160 - 177 |
|---|---|
| Autores principales: | , , , , , , |
| Formato: | research Journal Article |
| Publicado: |
Academic Press Inc.
May2017
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=122882468&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 122882468 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 15320464 OMB jtl: Journal of Biomedical Informatics issn: 15320464 maglogo: N pubinfo: dt: May2017 vid: 69 pid: 735 pub: Academic Press Inc. place: Burlington, Massachusetts artinfo: ui: 122882468 122882468 NLM28410983 122882468 10.1016/j.jbi.2017.04.008 NLM28410983 122882468 ppf: 160 ppct: 17 formats: tig: atl: Toward better public health reporting using existing off the shelf approaches: The value of medical dictionaries in automated cancer detection using plaintext medical data. aug: au: Kasthurirathne, Suranga N. Dixon, Brian E. Gichoya, Judy Xu, Huiping Xia, Yuni Mamlin, Burke Grannis, Shaun J. affil: Indiana University School of Informatics and Computing, Indianapolis, IN, USA sug: subj: Neoplasms Diagnosis Reference Books Algorithms Public Health ROC Curve Automation Human ab: Objectives: Existing approaches to derive decision models from plaintext clinical data frequently depend on medical dictionaries as the sources of potential features. Prior research suggests that decision models developed using non-dictionary based feature sourcing approaches and "off the shelf" tools could predict cancer with performance metrics between 80% and 90%. We sought to compare non-dictionary based models to models built using features derived from medical dictionaries.Materials and Methods: We evaluated the detection of cancer cases from free text pathology reports using decision models built with combinations of dictionary or non-dictionary based feature sourcing approaches, 4 feature subset sizes, and 5 classification algorithms. Each decision model was evaluated using the following performance metrics: sensitivity, specificity, accuracy, positive predictive value, and area under the receiver operating characteristics (ROC) curve.Results: Decision models parameterized using dictionary and non-dictionary feature sourcing approaches produced performance metrics between 70 and 90%. The source of features and feature subset size had no impact on the performance of a decision model.Conclusion: Our study suggests there is little value in leveraging medical dictionaries for extracting features for decision model building. Decision models built using features extracted from the plaintext reports themselves achieve comparable results to those built using medical dictionaries. Overall, this suggests that existing "off the shelf" approaches can be leveraged to perform accurate cancer detection using less complex Named Entity Recognition (NER) based feature extraction, automated feature selection and modeling approaches. pubtype: Academic Journal doctype: research Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|