A comprehensive study of named entity recognition in Chinese clinical text.
Objective: Named entity recognition (NER) is one of the fundamental tasks in natural language processing. In the medical domain, there have been a number of studies on NER in English clinical notes; however, very limited NER research has been carried out on clinical notes written in Chinese. The goa...
| Published in: | Journal of the American Medical Informatics Association Vol. 21; no. 5; pp. 808 - 815 |
|---|---|
| Main Authors: | , , , , , |
| Format: | research Journal Article |
| Published: |
Oxford University Press / USA
Sep2014
|
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=103985559&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 103985559 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 10675027 FZ9 jtl: Journal of the American Medical Informatics Association issn: 10675027 maglogo: N pubinfo: dt: Sep2014 vid: 21 iid: 5 pid: 622 pub: Oxford University Press / USA artinfo: ui: 103985559 NLM24347408 2012683757 10.1136/amiajnl-2013-002381 NLM24347408 PMC4147609 103985559 ppf: 808 ppct: 7 formats: tig: atl: A comprehensive study of named entity recognition in Chinese clinical text. aug: au: Lei, Jianbo Tang, Buzhou Lu, Xueqin Gao, Kaihua Jiang, Min Xu, Hua affil: Center for Medical Informatics, Peking University, Beijing, China The University of Texas School of Biomedical Informatics at Houston, Houston, Texas, USA. sug: subj: Algorithms Electronic Health Records Natural Language Processing Artificial Intelligence China Human Patient Admission ab: Objective: Named entity recognition (NER) is one of the fundamental tasks in natural language processing. In the medical domain, there have been a number of studies on NER in English clinical notes; however, very limited NER research has been carried out on clinical notes written in Chinese. The goal of this study was to systematically investigate features and machine learning algorithms for NER in Chinese clinical text.Materials and Methods: We randomly selected 400 admission notes and 400 discharge summaries from Peking Union Medical College Hospital in China. For each note, four types of entity-clinical problems, procedures, laboratory test, and medications-were annotated according to a predefined guideline. Two-thirds of the 400 notes were used to train the NER systems and one-third for testing. We investigated the effects of different types of feature including bag-of-characters, word segmentation, part-of-speech, and section information, and different machine learning algorithms including conditional random fields (CRF), support vector machines (SVM), maximum entropy (ME), and structural SVM (SSVM) on the Chinese clinical NER task. All classifiers were trained on the training dataset and evaluated on the test set, and micro-averaged precision, recall, and F-measure were reported.Results: Our evaluation on the independent test set showed that most types of feature were beneficial to Chinese NER systems, although the improvements were limited. The system achieved the highest performance by combining word segmentation and section information, indicating that these two types of feature complement each other. When the same types of optimized feature were used, CRF and SSVM outperformed SVM and ME. More specifically, SSVM achieved the highest performance of the four algorithms, with F-measures of 93.51% and 90.01% for admission notes and discharge summaries, respectively. pubtype: Academic Journal doctype: research Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|