Improving Laryngoscopy Image Analysis Through Integration of Global Information and Local Features in VoFoCD Dataset.

The diagnosis and treatment of vocal fold disorders heavily rely on the use of laryngoscopy. A comprehensive vocal fold diagnosis requires accurate identification of crucial anatomical structures and potential lesions during laryngoscopy observation. However, existing approaches have yet to explore...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Digital Imaging Vol. 37; no. 6; pp. 2794 - 2810
Autores principales: Dao, Thao Thi Phuong, Huynh, Tuan-Luc, Pham, Minh-Khoi, Le, Trung-Nghia, Nguyen, Tan-Cong, Nguyen, Quang-Thuc, Tran, Bich Anh, Van, Boi Ngoc, Ha, Chanh Cong, Tran, Minh-Triet
Formato: pictorial research tables/charts Journal Article
Publicado: Springer Nature Dec2024
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=182283950&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 182283950
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        08971889
        DOQ
      jtl: Journal of Digital Imaging
      issn: 08971889
      maglogo: N
    pubinfo:
      dt: Dec2024
      vid: 37
      iid: 6
      pid: 237
      pub: Springer Nature
      place: New York, New York
    artinfo:
      ui:
        182283950
        182283950
        182283950
        10.1007/s10278-024-01068-z
        182283950
      ppf: 2794
      ppct: 16
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
      tig:
        atl: Improving Laryngoscopy Image Analysis Through Integration of Global Information and Local Features in VoFoCD Dataset.
      aug:
        au:
          Dao, Thao Thi Phuong
          Huynh, Tuan-Luc
          Pham, Minh-Khoi
          Le, Trung-Nghia
          Nguyen, Tan-Cong
          Nguyen, Quang-Thuc
          Tran, Bich Anh
          Van, Boi Ngoc
          Ha, Chanh Cong
          Tran, Minh-Triet
        affil: University of Science, Ho Chi Minh City, Vietnam
      sug:
        subj:
          Vocal Cords Pathology
          Laryngoscopy
          Diagnosis, Otorhinolaryngologic
          Detection Algorithms
          Image Interpretation, Computer Assisted
          Image Enhancement
          Decision Support Systems, Clinical
          Human
          Outpatients
          Inpatients
          Vietnam
          Funding Source
          Retrospective Design
          Record Review
          Vocal Cords Anatomy and Histology
          Classification Algorithms
          Cysts
          Polyps
          Papilloma
          Hyperplasia
          Keratosis
          Dislocations
          Neoplasms
          Carcinoma in Situ
          Carcinoma, Squamous Cell
          Convolutional Neural Networks
      ab: The diagnosis and treatment of vocal fold disorders heavily rely on the use of laryngoscopy. A comprehensive vocal fold diagnosis requires accurate identification of crucial anatomical structures and potential lesions during laryngoscopy observation. However, existing approaches have yet to explore the joint optimization of the decision-making process, including object detection and image classification tasks simultaneously. In this study, we provide a new dataset, VoFoCD, with 1724 laryngology images designed explicitly for object detection and image classification in laryngoscopy images. Images in the VoFoCD dataset are categorized into four classes and comprise six glottic object types. Moreover, we propose a novel Multitask Efficient trAnsformer network for Laryngoscopy (MEAL) to classify vocal fold images and detect glottic landmarks and lesions. To further facilitate interpretability for clinicians, MEAL provides attention maps to visualize important learned regions for explainable artificial intelligence results toward supporting clinical decision-making. We also analyze our model's effectiveness in simulated clinical scenarios where shaking of the laryngoscopy process occurs. The proposed model demonstrates outstanding performance on our VoFoCD dataset. The accuracy for image classification and mean average precision at an intersection over a union threshold of 0.5 (mAP50) for object detection are 0.951 and 0.874, respectively. Our MEAL method integrates global knowledge, encompassing general laryngoscopy image classification, into local features, which refer to distinct anatomical regions of the vocal fold, particularly abnormal regions, including benign and malignant lesions. Our contribution can effectively aid laryngologists in identifying benign or malignant lesions of vocal folds and classifying images in the laryngeal endoscopy process visually.
      pubtype: Academic Journal
      doctype:
        pictorial
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N