Named entity recognition from Chinese adverse drug event reports with lexical feature based BiLSTM-CRF and tri-training.
Background: The Adverse Drug Event Reports (ADERs) from the spontaneous reporting system are important data sources for studying Adverse Drug Reactions (ADRs) as well as post-marketing pharmacovigilance. Apart from the conventional ADR information contained in the structured section of ADERs, more d...
| Publicado en: | Journal of Biomedical Informatics Vol. 96 |
|---|---|
| Autores principales: | , , , , , , |
| Formato: | research tables/charts Journal Article |
| Publicado: |
Academic Press Inc.
Aug2019
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=138031266&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 138031266 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 15320464 OMB jtl: Journal of Biomedical Informatics issn: 15320464 maglogo: N pubinfo: dt: Aug2019 vid: 96 pid: 735 pub: Academic Press Inc. place: Burlington, Massachusetts artinfo: ui: 138031266 138031266 NLM31323311 138031266 10.1016/j.jbi.2019.103252 NLM31323311 138031266 ppct: 1 formats: tig: atl: Named entity recognition from Chinese adverse drug event reports with lexical feature based BiLSTM-CRF and tri-training. aug: au: Chen, Yao Zhou, Changjiang Li, Tianxin Wu, Hong Zhao, Xia Ye, Kai Liao, Jun affil: School of Science, China Pharmaceutical University, Nanjing, China sug: subj: Adverse Drug Event Equipment and Supplies Algorithms Bioinformatics Language Adverse Drug Event Pharmacovigilance Natural Language Processing Reproducibility of Results Data Collection China Human Hospitals Validation Studies Comparative Studies Evaluation Research Multicenter Studies ab: Background: The Adverse Drug Event Reports (ADERs) from the spontaneous reporting system are important data sources for studying Adverse Drug Reactions (ADRs) as well as post-marketing pharmacovigilance. Apart from the conventional ADR information contained in the structured section of ADERs, more detailed information such as pre- and post- ADR symptoms, multi-drug usages and ADR-relief treatments are described in the free-text section, which can be mined through Natural Language Processing (NLP) tools.Objective: The goal of this study was to extract ADR-related entities from free-text section of Chinese ADERs, which can act as supplements for the information contained in structured section, so as to further assist in ADR evaluation.Methods: Three models of Conditional Random Field (CRF), Bidirectional Long Short-Term Memory-CRF (BiLSTM-CRF) and Lexical Feature based BiLSTM-CRF (LF-BiLSTM-CRF) were constructed to conduct Named Entity Recognition (NER) tasks in free-text section of Chinese ADERs. A semi-supervised learning method of tri-training was applied on the basis of the three established models to give un-annotated raw data with reliable tags.Results: Among the three basic models, the LF-BiLSTM-CRF achieved the highest average F1 score of 94.35%. After the process of tri-training, almost half of the un-annotated cases were tagged with labels, and the performances of all the three models improved after iterative training.Conclusions: The LF-BiLSTM-CRF model that we constructed could achieve a comparatively high F1 score, and the fusion of CRF, while BiLSTM-CRF and LF-BiLSTM-CRF in tri-training might further strengthen the reliability of predicted tags. The results suggested the usefulness of our methods in developing the specialized NER tools for identifying ADR-related information from Chinese ADERs. pubtype: Academic Journal doctype: research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|