Improving textual medication extraction using combined conditional random fields and rule-based systems.
Objective: In the i2b2 Medication Extraction Challenge, medication names together with details of their administration were to be extracted from medical discharge summaries.Design: The task of the challenge was decomposed into three pipelined components: named entity identification, context-aware fi...
| Published in: | Journal of the American Medical Informatics Association Vol. 17; no. 5; pp. 540 - 545 |
|---|---|
| Main Authors: | , , , |
| Format: | research Journal Article |
| Published: |
Oxford University Press / USA
Sep2010
|
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=105091995&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 105091995 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 10675027 FZ9 jtl: Journal of the American Medical Informatics Association issn: 10675027 maglogo: N pubinfo: dt: Sep2010 vid: 17 iid: 5 pid: 622 pub: Oxford University Press / USA artinfo: ui: 105091995 NLM20819860 2010773952 10.1136/jamia.2010.004119 NLM20819860 PMC2995683 105091995 ppf: 540 ppct: 5 formats: tig: atl: Improving textual medication extraction using combined conditional random fields and rule-based systems. aug: au: Tikk D Solt I Tikk, Domonkos Solt, Illés affil: Department of Telecommunications and Media Informatics, Budapest University of Technology and Economics, Budapest, Hungary sug: subj: Electronic Health Records Information Retrieval Methods Natural Language Processing Drug Administration Human Software Design ab: Objective: In the i2b2 Medication Extraction Challenge, medication names together with details of their administration were to be extracted from medical discharge summaries.Design: The task of the challenge was decomposed into three pipelined components: named entity identification, context-aware filtering and relation extraction. For named entity identification, first a rule-based (RB) method that was used in our overall fifth place-ranked solution at the challenge was investigated. Second, a conditional random fields (CRF) approach is presented for named entity identification (NEI) developed after the completion of the challenge. The CRF models are trained on the 17 ground truth documents, the output of the rule-based NEI component on all documents, a larger but potentially inaccurate training dataset. For both NEI approaches their effect on relation extraction performance was investigated. The filtering and relation extraction components are both rule-based.Measurements: In addition to the official entry level evaluation of the challenge, entity level analysis is also provided.Results: On the test data an entry level F(1)-score of 80% was achieved for exact matching and 81% for inexact matching with the RB-NEI component. The CRF produces a significantly weaker result, but CRF outperforms the rule-based model with 81% exact and 82% inexact F(1)-score (p<0.02).Conclusion: This study shows that a simple rule-based method is on a par with more complicated machine learners; CRF models can benefit from the addition of the potentially inaccurate training data, when only very few training documents are available. Such training data could be generated using the outputs of rule-based methods. pubtype: Academic Journal doctype: research Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|