Annotating patient clinical records with syntactic chunks and named entities: the Harvey Corpus.

The free text notes typed by physicians during patient consultations contain valuable information for the study of disease and treatment. These notes are difficult to process by existing natural language analysis tools since they are highly telegraphic (omitting many words), and contain many spellin...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 50; no. 3; pp. 523 - 549
Autores principales: Savkov, Aleksandar, Carroll, John, Koeling, Rob, Cassell, Jackie
Formato: Artículo
Publicado: Springer Nature Sep2016
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=117418392&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 117418392
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2016
      vid: 50
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        117418392
        10.1007/s10579-015-9330-7
      ppf: 523
      ppct: 26
      formats:
        fmt:
          @attributes:
            type: P
            size: 910KB
      tig:
        atl: Annotating patient clinical records with syntactic chunks and named entities: the Harvey Corpus.
      aug:
        au:
          Savkov, Aleksandar
          Carroll, John
          Koeling, Rob
          Cassell, Jackie
        affil:
          Department of Informatics , University of Sussex , Brighton BN1 9QJ UK
          Division of Primary Care and Public Health , Brighton and Sussex Medical School , Brighton BN1 9PH UK
      su:
        Annotations
        Medical records
        Corpora
        Spelling errors
        Language & languages
        Machine learning
      sug:
        subj:
          Annotations
          Medical records
          Corpora
          Spelling errors
          Language & languages
          Machine learning
      keyword:
        Annotation guidelines
        Chunking
        Clinical text
        Corpus annotation
        Named entities
      ab: The free text notes typed by physicians during patient consultations contain valuable information for the study of disease and treatment. These notes are difficult to process by existing natural language analysis tools since they are highly telegraphic (omitting many words), and contain many spelling mistakes, inconsistencies in punctuation, and non-standard word order. To support information extraction and classification tasks over such text, we describe a de-identified corpus of free text notes, a shallow syntactic and named entity annotation scheme for this kind of text, and an approach to training domain specialists with no linguistic background to annotate the text. Finally, we present a statistical chunking system for such clinical text with a stable learning rate and good accuracy, indicating that the manual annotation is consistent and that the annotation scheme is tractable for machine learning.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2016. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2016
    holdings:
      @attributes:
        islocal: N