Extracting comprehensive clinical information for breast cancer using deep learning methods.

Objective: Breast cancer is the most common malignant tumor among women. The diagnosis and treatment information of breast cancer patients is abundant in multiple types of clinical fields, including clinicopathological data, genotype and phenotype information, treatment information, and prognosis in...

Descripción completa

Detalles Bibliográficos
Publicado en:International Journal of Medical Informatics Vol. 132
Autores principales: Zhang, Xiaohui, Zhang, Yaoyun, Zhang, Qin, Ren, Yuankai, Qiu, Tinglin, Ma, Jianhui, Sun, Qiang
Formato: research Journal Article
Publicado: Elsevier B.V. Dec2019
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=139277842&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 139277842
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        13865056
        JR4
      jtl: International Journal of Medical Informatics
      issn: 13865056
      maglogo: N
    pubinfo:
      dt: Dec2019
      vid: 132
      pid: 467
      pub: Elsevier B.V.
      place: New York, New York
    artinfo:
      ui:
        139277842
        139277842
        NLM31627032
        139277842
        10.1016/j.ijmedinf.2019.103985
        NLM31627032
        139277842
      ppct: 1
      formats:
      tig:
        atl: Extracting comprehensive clinical information for breast cancer using deep learning methods.
      aug:
        au:
          Zhang, Xiaohui
          Zhang, Yaoyun
          Zhang, Qin
          Ren, Yuankai
          Qiu, Tinglin
          Ma, Jianhui
          Sun, Qiang
        affil: Peking Union Medical College Hospital, Peking Union Medical College & Chinese Academy of Medical Sciences, Beijing, China
      sug:
        subj:
          Breast Neoplasms Diagnosis
          Algorithms
          Natural Language Processing
          Breast Neoplasms Therapy
          Human
          Female
          Breast Neoplasms Epidemiology
          China
          Validation Studies
          Comparative Studies
          Evaluation Research
          Multicenter Studies
          Barthel Index
          Scales
          Short Portable Mental Status Questionnaire
          Female
      ab: Objective: Breast cancer is the most common malignant tumor among women. The diagnosis and treatment information of breast cancer patients is abundant in multiple types of clinical fields, including clinicopathological data, genotype and phenotype information, treatment information, and prognosis information. However, current studies are mainly focused on extracting information from one specific type of clinical field. This study defines a comprehensive information model to represent the whole-course clinical information of patients. Furthermore, deep learning approaches are used to extract the concepts and their attributes from clinical breast cancer documents by fine-tuning pretrained Bidirectional Encoder Representations from Transformers (BERT) language models.Materials and Methods: The clinical corpus that was used in this study was from one 3A cancer hospital in China, consisting of the encounter notes, operation records, pathology notes, radiology notes, progress notes and discharge summaries of 100 breast cancer patients. Our system consists of two components: a named entity recognition (NER) component and a relation recognition component. For each component, we implemented deep learning-based approaches by fine-tuning BERT, which outperformed other state-of-the-art methods on multiple natural language processing (NLP) tasks. A clinical language model is first pretrained using BERT on a large-scale unlabeled corpus of Chinese clinical text. For NER, the context embeddings that were pretrained using BERT were used as the input features of the Bi-LSTM-CRF (Bidirectional long-short-memory-conditional random fields) model and were fine-tuned using the annotated breast cancer notes. Furthermore, we proposed an approach to fine-tune BERT for relation extraction. It was considered to be a classification problem in which the two entities that were mentioned in the input sentence were replaced with their semantic types.Results: Our best-performing system achieved F1 scores of 93.53% for the NER and 96.73% for the relation extraction. Additional evaluations showed that the deep learning-based approaches that fine-tuned BERT did outperform the traditional Bi-LSTM-CRF and CRF machine learning algorithms in NER and the attention-Bi-LSTM and SVM (support vector machines) algorithms in relation recognition.Conclusion: In this study, we developed a deep learning approach that fine-tuned BERT to extract the breast cancer concepts and their attributes. It demonstrated its superior performance compared to traditional machine learning algorithms, thus supporting its uses in broader NER and relation extraction tasks in the medical domain.
      pubtype: Academic Journal
      doctype:
        research
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N