Feature engineering combined with machine learning and rule-based methods for structured information extraction from narrative clinical discharge summaries.
Objective: A system that translates narrative text in the medical domain into structured representation is in great demand. The system performs three sub-tasks: concept extraction, assertion classification, and relation identification.Design: The overall system consists of five steps: (1) pre-proces...
| Publicado en: | Journal of the American Medical Informatics Association Vol. 19; no. 5; pp. 824 - 833 |
|---|---|
| Autores principales: | , , , |
| Formato: | research Journal Article |
| Publicado: |
Oxford University Press / USA
Sep2012
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=104360929&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 104360929 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 10675027 FZ9 jtl: Journal of the American Medical Informatics Association issn: 10675027 maglogo: N pubinfo: dt: Sep2012 vid: 19 iid: 5 pid: 622 pub: Oxford University Press / USA artinfo: ui: 104360929 NLM22586067 2011646698 10.1136/amiajnl-2011-000776 NLM22586067 PMC3422834 104360929 ppf: 824 ppct: 9 formats: tig: atl: Feature engineering combined with machine learning and rule-based methods for structured information extraction from narrative clinical discharge summaries. aug: au: Xu, Yan Hong, Kai Tsujii, Junichi Chang, Eric I-Chao affil: State Key Laboratory of Software Development Environment, Key Laboratory of Biomechanics and Mechanobiology of the Ministry of Education, Beihang University, Beijing, China. sug: subj: Data Mining Methods Electronic Health Records Natural Language Processing Patient Discharge Artificial Intelligence Human Vocabulary, Controlled ab: Objective: A system that translates narrative text in the medical domain into structured representation is in great demand. The system performs three sub-tasks: concept extraction, assertion classification, and relation identification.Design: The overall system consists of five steps: (1) pre-processing sentences, (2) marking noun phrases (NPs) and adjective phrases (APs), (3) extracting concepts that use a dosage-unit dictionary to dynamically switch two models based on Conditional Random Fields (CRF), (4) classifying assertions based on voting of five classifiers, and (5) identifying relations using normalized sentences with a set of effective discriminating features.Measurements: Macro-averaged and micro-averaged precision, recall and F-measure were used to evaluate results.Results: The performance is competitive with the state-of-the-art systems with micro-averaged F-measure of 0.8489 for concept extraction, 0.9392 for assertion classification and 0.7326 for relation identification.Conclusions: The system exploits an array of common features and achieves state-of-the-art performance. Prudent feature engineering sets the foundation of our systems. In concept extraction, we demonstrated that switching models, one of which is especially designed for telegraphic sentences, improved extraction of the treatment concept significantly. In assertion classification, a set of features derived from a rule-based classifier were proven to be effective for the classes such as conditional and possible. These classes would suffer from data scarcity in conventional machine-learning methods. In relation identification, we use two-staged architecture, the second of which applies pairwise classifiers to possible candidate classes. This architecture significantly improves performance. pubtype: Academic Journal doctype: research Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|