ConTEXTual Net: A Multimodal Vision-Language Model for Segmentation of Pneumothorax.

Radiology narrative reports often describe characteristics of a patient's disease, including its location, size, and shape. Motivated by the recent success of multimodal learning, we hypothesized that this descriptive text could guide medical image analysis algorithms. We proposed a novel vision-lan...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Digital Imaging Vol. 37; no. 4; pp. 1652 - 1664
Autores principales: Huemann, Zachary, Tie, Xin, Hu, Junjie, Bradshaw, Tyler J.
Formato: diagnostic images equations & formulas research tables/charts Journal Article
Publicado: Springer Nature Aug2024
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=179554139&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 179554139
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        08971889
        DOQ
      jtl: Journal of Digital Imaging
      issn: 08971889
      maglogo: N
    pubinfo:
      dt: Aug2024
      vid: 37
      iid: 4
      pid: 237
      pub: Springer Nature
      place: New York, New York
    artinfo:
      ui:
        179554139
        179554139
        179554139
        10.1007/s10278-024-01051-8
        179554139
      ppf: 1652
      ppct: 12
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
      tig:
        atl: ConTEXTual Net: A Multimodal Vision-Language Model for Segmentation of Pneumothorax.
      aug:
        au:
          Huemann, Zachary
          Tie, Xin
          Hu, Junjie
          Bradshaw, Tyler J.
        affil: https://ror.org/01y2jtd41 Department of Radiology, University of Wisconsin-Madison, 53705, Madison, WI, USA
      sug:
        subj:
          Pneumothorax Radiography
          Neural Networks (Computer)
          Radiography, Thoracic
          Human
          Natural Language Processing Methods
          Algorithms
          Radiography
          Descriptive Statistics
          Funding Source
      ab: Radiology narrative reports often describe characteristics of a patient's disease, including its location, size, and shape. Motivated by the recent success of multimodal learning, we hypothesized that this descriptive text could guide medical image analysis algorithms. We proposed a novel vision-language model, ConTEXTual Net, for the task of pneumothorax segmentation on chest radiographs. ConTEXTual Net extracts language features from physician-generated free-form radiology reports using a pre-trained language model. We then introduced cross-attention between the language features and the intermediate embeddings of an encoder-decoder convolutional neural network to enable language guidance for image analysis. ConTEXTual Net was trained on the CANDID-PTX dataset consisting of 3196 positive cases of pneumothorax with segmentation annotations from 6 different physicians as well as clinical radiology reports. Using cross-validation, ConTEXTual Net achieved a Dice score of 0.716±0.016, which was similar to the degree of inter-reader variability (0.712±0.044) computed on a subset of the data. It outperformed vision-only models (Swin UNETR: 0.670±0.015, ResNet50 U-Net: 0.677±0.015, GLoRIA: 0.686±0.014, and nnUNet 0.694±0.016) and a competing vision-language model (LAVT: 0.706±0.009). Ablation studies confirmed that it was the text information that led to the performance gains. Additionally, we show that certain augmentation methods degraded ConTEXTual Net's segmentation performance by breaking the image-text concordance. We also evaluated the effects of using different language models and activation functions in the cross-attention module, highlighting the efficacy of our chosen architectural design.
      pubtype: Academic Journal
      doctype:
        diagnostic images
        equations & formulas
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N