OralEpitheliumDB: A Dataset for Oral Epithelial Dysplasia Image Segmentation and Classification.

Early diagnosis of potentially malignant disorders, such as oral epithelial dysplasia, is the most reliable way to prevent oral cancer. Computational algorithms have been used as an auxiliary tool to aid specialists in this process. Usually, experiments are performed on private data, making it diffi...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Digital Imaging Vol. 37; no. 4; pp. 1691 - 1711
Autores principales: Silva, Adriano Barbosa, Martins, Alessandro Santana, Tosta, Thaína Aparecida Azevedo, Loyola, Adriano Mota, Cardoso, Sérgio Vitorino, Neves, Leandro Alves, de Faria, Paulo Rogério, do Nascimento, Marcelo Zanchetta
Formato: equations & formulas pictorial research tables/charts Journal Article
Publicado: Springer Nature Aug2024
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=179554132&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 179554132
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        08971889
        DOQ
      jtl: Journal of Digital Imaging
      issn: 08971889
      maglogo: N
    pubinfo:
      dt: Aug2024
      vid: 37
      iid: 4
      pid: 237
      pub: Springer Nature
      place: New York, New York
    artinfo:
      ui:
        179554132
        179554132
        179554132
        10.1007/s10278-024-01041-w
        179554132
      ppf: 1691
      ppct: 20
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
      tig:
        atl: OralEpitheliumDB: A Dataset for Oral Epithelial Dysplasia Image Segmentation and Classification.
      aug:
        au:
          Silva, Adriano Barbosa
          Martins, Alessandro Santana
          Tosta, Thaína Aparecida Azevedo
          Loyola, Adriano Mota
          Cardoso, Sérgio Vitorino
          Neves, Leandro Alves
          de Faria, Paulo Rogério
          do Nascimento, Marcelo Zanchetta
        affil: Faculty of Computer Science (FACOM) - Federal University of Uberlândia (UFU), Av. João Naves de Ávila 2121, BLB, 38400-902, Uberlândia, MG, Brazil
      sug:
        subj:
          Image Interpretation, Computer Assisted
          Data Management
          Epithelium Pathology
          Mouth Neoplasms Diagnosis
          Mouth Neoplasms Classification
          Precancerous Conditions Diagnosis
          Human
          Funding Source
          Neoplasm Grading
          Pathologists
          Neural Networks (Computer)
          Image Processing, Computer Assisted
          Machine Learning
          Algorithms
          Random Forest
          Validity
          Automation
      ab: Early diagnosis of potentially malignant disorders, such as oral epithelial dysplasia, is the most reliable way to prevent oral cancer. Computational algorithms have been used as an auxiliary tool to aid specialists in this process. Usually, experiments are performed on private data, making it difficult to reproduce the results. There are several public datasets of histological images, but studies focused on oral dysplasia images use inaccessible datasets. This prevents the improvement of algorithms aimed at this lesion. This study introduces an annotated public dataset of oral epithelial dysplasia tissue images. The dataset includes 456 images acquired from 30 mouse tongues. The images were categorized among the lesion grades, with nuclear structures manually marked by a trained specialist and validated by a pathologist. Also, experiments were carried out in order to illustrate the potential of the proposed dataset in classification and segmentation processes commonly explored in the literature. Convolutional neural network (CNN) models for semantic and instance segmentation were employed on the images, which were pre-processed with stain normalization methods. Then, the segmented and non-segmented images were classified with CNN architectures and machine learning algorithms. The data obtained through these processes is available in the dataset. The segmentation stage showed the F1-score value of 0.83, obtained with the U-Net model using the ResNet-50 as a backbone. At the classification stage, the most expressive result was achieved with the Random Forest method, with an accuracy value of 94.22%. The results show that the segmentation contributed to the classification results, but studies are needed for the improvement of these stages of automated diagnosis. The original, gold standard, normalized, and segmented images are publicly available and may be used for the improvement of clinical applications of CAD methods on oral epithelial dysplasia tissue images.
      pubtype: Academic Journal
      doctype:
        equations & formulas
        pictorial
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N