A Multilevel Transfer Learning Technique and LSTM Framework for Generating Medical Captions for Limited CT and DBT Images.

Medical image captioning has been recently attracting the attention of the medical community. Also, generating captions for images involving multiple organs is an even more challenging task. Therefore, any attempt toward such medical image captioning becomes the need of the hour. In recent years, th...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Digital Imaging Vol. 35; no. 3; pp. 564 - 581
Autores principales: Aswiga, R. V., Shanthi, A. P.
Formato: equations & formulas pictorial research tables/charts Journal Article
Publicado: Springer Nature Jun2022
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=157184671&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 157184671
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        08971889
        DOQ
      jtl: Journal of Digital Imaging
      issn: 08971889
      maglogo: N
    pubinfo:
      dt: Jun2022
      vid: 35
      iid: 3
      pid: 237
      pub: Springer Nature
      place: New York, New York
    artinfo:
      ui:
        157184671
        155426397
        157184671
        157184671
        10.1007/s10278-021-00567-7
        157184671
      ppf: 564
      ppct: 17
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
      tig:
        atl: A Multilevel Transfer Learning Technique and LSTM Framework for Generating Medical Captions for Limited CT and DBT Images.
      aug:
        au:
          Aswiga, R. V.
          Shanthi, A. P.
        affil: Department of Computer Science & Engineering, Amrita School of Engineering, Amrita Vishwa Vidyapeetham, Tamil Nadu, 601103, Chennai, India
      sug:
        subj:
          Deep Learning
          Long Short-Term Memory
          Models, Psychological
          Human
          Experimental Studies
          Conceptual Framework
          Image Processing, Computer Assisted
          Word Processing
      ab: Medical image captioning has been recently attracting the attention of the medical community. Also, generating captions for images involving multiple organs is an even more challenging task. Therefore, any attempt toward such medical image captioning becomes the need of the hour. In recent years, the rapid developments in deep learning approaches have made them an effective option for the analysis of medical images and automatic report generation. But analyzing medical images that are scarce and limited is hard, and it is difficult even with machine learning approaches. The concept of transfer learning can be employed in such applications that suffer from insufficient training data. This paper presents an approach to develop a medical image captioning model based on a deep recurrent architecture that combines Multi Level Transfer Learning (MLTL) framework with a Long Short-Term-Memory (LSTM) model. A basic MLTL framework with three models is designed to detect and classify very limited datasets, using the knowledge acquired from easily available datasets. The first model for the source domain uses the abundantly available non-medical images and learns the generalized features. The acquired knowledge is then transferred to the second model for the intermediate and auxiliary domain, which is related to the target domain. This information is then used for the final target domain, which consists of medical datasets that are very limited in nature. Therefore, the knowledge learned from a non-medical source domain is transferred to improve the learning in the target domain that deals with medical images. Then, a novel LSTM model, which is used for sequence generation and machine translation, is proposed to generate captions for the given medical image from the MLTL framework. To improve the captioning of the target sentence further, an enhanced multi-input Convolutional Neural Network (CNN) model along with feature extraction techniques is proposed. This enhanced multi-input CNN model extracts the most important features of an image that help in generating a more precise and detailed caption of the medical image. Experimental results show that the proposed model performs well with an accuracy of 96.90%, with BLEU score of 76.9%, even with very limited datasets, when compared to the work reported in literature.
      pubtype: Academic Journal
      doctype:
        equations & formulas
        pictorial
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N