Automated classification of cancer morphology from Italian pathology reports using Natural Language Processing techniques: A rule-based approach.

Pathology reports represent a primary source of information for cancer registries. Hospitals routinely process high volumes of free-text reports, a valuable source of information regarding cancer diagnosis for improving clinical care and supporting research. Information extraction and coding of text...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Biomedical Informatics Vol. 116
Autores principales: Hammami, Linda, Paglialonga, Alessia, Pruneri, Giancarlo, Torresani, Michele, Sant, Milena, Bono, Carlo, Caiani, Enrico Gianluca, Baili, Paolo
Formato: research Journal Article
Publicado: Academic Press Inc. Apr2021
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=149784909&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 149784909
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        15320464
        OMB
      jtl: Journal of Biomedical Informatics
      issn: 15320464
      maglogo: N
    pubinfo:
      dt: Apr2021
      vid: 116
      pid: 735
      pub: Academic Press Inc.
      place: Burlington, Massachusetts
    artinfo:
      ui:
        149784909
        149784909
        NLM33609761
        149784909
        10.1016/j.jbi.2021.103712
        NLM33609761
        149784909
      ppct: 1
      formats:
      tig:
        atl: Automated classification of cancer morphology from Italian pathology reports using Natural Language Processing techniques: A rule-based approach.
      aug:
        au:
          Hammami, Linda
          Paglialonga, Alessia
          Pruneri, Giancarlo
          Torresani, Michele
          Sant, Milena
          Bono, Carlo
          Caiani, Enrico Gianluca
          Baili, Paolo
        affil: Analytical Epidemiology and Health Impact Unit, Fondazione IRCCS "Istituto Nazionale dei Tumori", Milan, Italy
      sug:
        subj:
          Neoplasms Diagnosis
          Natural Language Processing
          Language
          Italy
          Human
          Information Retrieval
          Comparative Studies
          Multicenter Studies
          Evaluation Research
          Validation Studies
          Scales
          Questionnaires
      ab: Pathology reports represent a primary source of information for cancer registries. Hospitals routinely process high volumes of free-text reports, a valuable source of information regarding cancer diagnosis for improving clinical care and supporting research. Information extraction and coding of textual unstructured data is typically a manual, labour-intensive process. There is a need to develop automated approaches to extract meaningful information from such texts in a reliable and accurate way. In this scenario, Natural Language Processing (NLP) algorithms offer a unique opportunity to automatically encode the unstructured reports into structured data, thus representing a potential powerful alternative to expensive manual processing. However, notwithstanding the increasing interest in this area, there is still limited availability of NLP approaches for pathology reports in languages other than English, including Italian, to date. The aim of our work was to develop an automated algorithm based on NLP techniques, able to identify and classify the morphological content of pathology reports in the Italian language with micro-averaged performance scores higher than 95%. Specifically, a novel, domain-specific classifier that uses linguistic rules was developed and tested on 27,239 pathology reports from a single Italian oncological centre, following the International Classification of Diseases for Oncology morphology classification standard (ICD-O-M). The proposed classification algorithm achieved successful results with a micro-F1 score of 98.14% on 9594 pathology reports in the test dataset. This algorithm relies on rules defined on data from a single hospital that is specifically dedicated to cancer, but it is based on general processing steps which can be applied to different datasets. Further research will be important to demonstrate the generalizability of the proposed approach on a larger corpus from different hospitals.
      pubtype: Academic Journal
      doctype:
        research
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N