A French clinical corpus with comprehensive semantic annotations: development of the Medical Entity and Relation LIMSI annOtated Text corpus (MERLOT).

Quality annotated resources are essential for Natural Language Processing. The objective of this work is to present a corpus of clinical narratives in French annotated for linguistic, semantic and structural information, aimed at clinical information extraction. Six annotators contributed to the cor...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 52; no. 2; pp. 571 - 602
Autores principales: Campillos, Leonardo, Deléger, Louise, Grouin, Cyril, Hamon, Thierry, Ligozat, Anne-Laure, Névéol, Aurélie
Formato: Artículo
Publicado: Springer Nature Jun2018
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=129593476&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 129593476
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Jun2018
      vid: 52
      iid: 2
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        129593476
        10.1007/s10579-017-9382-y
      ppf: 571
      ppct: 31
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.7MB
      tig:
        atl: A French clinical corpus with comprehensive semantic annotations: development of the Medical Entity and Relation LIMSI annOtated Text corpus (MERLOT).
      aug:
        au:
          Campillos, Leonardo
          Deléger, Louise
          Grouin, Cyril
          Hamon, Thierry
          Ligozat, Anne-Laure
          Névéol, Aurélie
        affil: LIMSI-CNRS, Université Paris Saclay, 91403, Orsay, France
      su:
        Natural language processing
        French language
        Corpora
        Semantics
        Medical records
      sug:
        subj:
          Natural language processing
          French language
          Corpora
          Semantics
          Medical records
      keyword:
        Clinical narrative
        Inter-annotator agreement
        Personal health information
        Semantic annotations
      ab: Quality annotated resources are essential for Natural Language Processing. The objective of this work is to present a corpus of clinical narratives in French annotated for linguistic, semantic and structural information, aimed at clinical information extraction. Six annotators contributed to the corpus annotation, using a comprehensive annotation scheme covering 21 entities, 11 attributes and 37 relations. All annotators trained on a small, common portion of the corpus before proceeding independently. An automatic tool was used to produce entity and attribute pre-annotations. About a tenth of the corpus was doubly annotated and annotation differences were resolved in consensus meetings. To ensure annotation consistency throughout the corpus, we devised harmonization tools to automatically identify annotation differences to be addressed to improve the overall corpus quality. The annotation project spanned over 24 months and resulted in a corpus comprising 500 documents (148,476 tokens) annotated with 44,740 entities and 26,478 relations. The average inter-annotator agreement is 0.793 F-measure for entities and 0.789 for relations. The performance of the pre-annotation tool for entities reached 0.814 F-measure when sufficient training data was available. The performance of our entity pre-annotation tool shows the value of the corpus to build and evaluate information extraction methods. In addition, we introduced harmonization methods that further improved the quality of annotations in the corpus.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2018. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2018
    holdings:
      @attributes:
        islocal: N