Annotating patient clinical records with syntactic chunks and named entities: the Harvey Corpus.
The free text notes typed by physicians during patient consultations contain valuable information for the study of disease and treatment. These notes are difficult to process by existing natural language analysis tools since they are highly telegraphic (omitting many words), and contain many spellin...
| Publicado en: | Language Resources & Evaluation Vol. 50; no. 3; pp. 523 - 549 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Sep2016
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=117418392&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 117418392 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2016 vid: 50 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 117418392 10.1007/s10579-015-9330-7 ppf: 523 ppct: 26 formats: fmt: @attributes: type: P size: 910KB tig: atl: Annotating patient clinical records with syntactic chunks and named entities: the Harvey Corpus. aug: au: Savkov, Aleksandar Carroll, John Koeling, Rob Cassell, Jackie affil: Department of Informatics , University of Sussex , Brighton BN1 9QJ UK Division of Primary Care and Public Health , Brighton and Sussex Medical School , Brighton BN1 9PH UK su: Annotations Medical records Corpora Spelling errors Language & languages Machine learning sug: subj: Annotations Medical records Corpora Spelling errors Language & languages Machine learning keyword: Annotation guidelines Chunking Clinical text Corpus annotation Named entities ab: The free text notes typed by physicians during patient consultations contain valuable information for the study of disease and treatment. These notes are difficult to process by existing natural language analysis tools since they are highly telegraphic (omitting many words), and contain many spelling mistakes, inconsistencies in punctuation, and non-standard word order. To support information extraction and classification tasks over such text, we describe a de-identified corpus of free text notes, a shallow syntactic and named entity annotation scheme for this kind of text, and an approach to training domain specialists with no linguistic background to annotate the text. Finally, we present a statistical chunking system for such clinical text with a stable learning rate and good accuracy, indicating that the manual annotation is consistent and that the annotation scheme is tractable for machine learning. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2016. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2016 holdings: @attributes: islocal: N |
|---|