PASCAL: a pseudo cascade learning framework for breast cancer treatment entity normalization in Chinese clinical text.

Backgrounds: Knowledge discovery from breast cancer treatment records has promoted downstream clinical studies such as careflow mining and therapy analysis. However, the clinical treatment text from electronic health data might be recorded by different doctors under their hospital guidelines, making...

Descripción completa

Detalles Bibliográficos
Publicado en:BMC Medical Informatics & Decision Making Vol. 20; no. 1
Autores principales: An, Yang, Wang, Jianlin, Zhang, Liang, Zhao, Hanyu, Gao, Zhan, Huang, Haitao, Du, Zhenguang, Jiao, Zengtao, Yan, Jun, Wei, Xiaopeng, Jin, Bo
Formato: research Journal Article
Publicado: BioMed Central 8/28/2020
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=145372061&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 145372061
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        14726947
        1CI0
      jtl: BMC Medical Informatics & Decision Making
      issn: 14726947
      maglogo: N
    pubinfo:
      dt: 8/28/2020
      vid: 20
      iid: 1
      pid: 24147
      pub: BioMed Central
    artinfo:
      ui:
        145372061
        145372061
        NLM32859189
        145372061
        10.1186/s12911-020-01216-9
        NLM32859189
        145372061
      ppct: 1
      formats:
      tig:
        atl: PASCAL: a pseudo cascade learning framework for breast cancer treatment entity normalization in Chinese clinical text.
      aug:
        au:
          An, Yang
          Wang, Jianlin
          Zhang, Liang
          Zhao, Hanyu
          Gao, Zhan
          Huang, Haitao
          Du, Zhenguang
          Jiao, Zengtao
          Yan, Jun
          Wei, Xiaopeng
          Jin, Bo
        affil: School of Computer Science and Technology, Dalian University of Technology, No.2 Linggong Road, Ganjingzi District, 116024, Dalian, Liaoning, China
      sug:
        subj:
          Breast Neoplasms Drug Therapy
          Resource Databases
          Human
          Text Messaging
          Female
          Validation Studies
          Comparative Studies
          Evaluation Research
          Multicenter Studies
          Scales
          Female
      ab: Backgrounds: Knowledge discovery from breast cancer treatment records has promoted downstream clinical studies such as careflow mining and therapy analysis. However, the clinical treatment text from electronic health data might be recorded by different doctors under their hospital guidelines, making the final data rich in author- and domain-specific idiosyncrasies. Therefore, breast cancer treatment entity normalization becomes an essential task for the above downstream clinical studies. The latest studies have demonstrated the superiority of deep learning methods in named entity normalization tasks. Fundamentally, most existing approaches adopt pipeline implementations that treat it as an independent process after named entity recognition, which can propagate errors to later tasks. In addition, despite its importance in clinical and translational research, few studies directly deal with the normalization task in Chinese clinical text due to the complexity of composition forms.Methods: To address these issues, we propose PASCAL, an end-to-end and accurate framework for breast cancer treatment entity normalization (TEN). PASCAL leverages a gated convolutional neural network to obtain a representation vector that can capture contextual features and long-term dependencies. Additionally, it treats treatment entity recognition (TER) as an auxiliary task that can provide meaningful information to the primary TEN task and as a particular regularization to further optimize the shared parameters. Finally, by concatenating the context-aware vector and probabilistic distribution vector from TEN, we utilize the conditional random field layer (CRF) to model the normalization sequence and predict the TEN sequential results.Results: To evaluate the effectiveness of the proposed framework, we employ the three latest sequential models as baselines and build the model in single- and multitask on a real-world database. Experimental results show that our method achieves better accuracy and efficiency than state-of-the-art approaches.Conclusions: The effectiveness and efficiency of the presented pseudo cascade learning framework were validated for breast cancer treatment normalization in clinical text. We believe the predominant performance lies in its ability to extract valuable information from unstructured text data, which will significantly contribute to downstream tasks, such as treatment recommendations, breast cancer staging and careflow mining.
      pubtype: Academic Journal
      doctype:
        research
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N