Finding a Suitable Class Distribution for Building Histological Images Datasets Used in Deep Model Training—The Case of Cancer Detection.

The class distribution of a training dataset is an important factor which influences the performance of a deep learning-based system. Understanding the optimal class distribution is therefore crucial when building a new training set which may be costly to annotate. This is the case for histological...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Digital Imaging Vol. 35; no. 5; pp. 1326 - 1350
Autores principales: Reshma, Ismat Ara, Franchet, Camille, Gaspard, Margot, Ionescu, Radu Tudor, Mothe, Josiane, Cussat-Blanc, Sylvain, Luga, Hervé, Brousset, Pierre
Formato: equations & formulas pictorial research tables/charts Journal Article
Publicado: Springer Nature Oct2022
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=159758922&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 159758922
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        08971889
        DOQ
      jtl: Journal of Digital Imaging
      issn: 08971889
      maglogo: N
    pubinfo:
      dt: Oct2022
      vid: 35
      iid: 5
      pid: 237
      pub: Springer Nature
      place: New York, New York
    artinfo:
      ui:
        159758922
        156417009
        159758922
        159758922
        10.1007/s10278-022-00618-7
        159758922
      ppf: 1326
      ppct: 24
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
      tig:
        atl: Finding a Suitable Class Distribution for Building Histological Images Datasets Used in Deep Model Training—The Case of Cancer Detection.
      aug:
        au:
          Reshma, Ismat Ara
          Franchet, Camille
          Gaspard, Margot
          Ionescu, Radu Tudor
          Mothe, Josiane
          Cussat-Blanc, Sylvain
          Luga, Hervé
          Brousset, Pierre
        affil: IRIT, UMR5505 CNRS, Université de Toulouse, Toulouse, France
      sug:
        subj:
          Neoplasms Diagnosis
          Staining and Labeling Methods
          Diagnosis, Computer Assisted Methods
          Deep Learning Methods
          Models, Statistical Evaluation
          Human
          Neoplasms Classification
          Artificial Intelligence
          Predictive Value of Tests
          Image Interpretation, Computer Assisted
          Medical Informatics
          Information Retrieval
          Histological Techniques
      ab: The class distribution of a training dataset is an important factor which influences the performance of a deep learning-based system. Understanding the optimal class distribution is therefore crucial when building a new training set which may be costly to annotate. This is the case for histological images used in cancer diagnosis where image annotation requires domain experts. In this paper, we tackle the problem of finding the optimal class distribution of a training set to be able to train an optimal model that detects cancer in histological images. We formulate several hypotheses which are then tested in scores of experiments with hundreds of trials. The experiments have been designed to account for both segmentation and classification frameworks with various class distributions in the training set, such as natural, balanced, over-represented cancer, and over-represented non-cancer. In the case of cancer detection, the experiments show several important results: (a) the natural class distribution produces more accurate results than the artificially generated balanced distribution; (b) the over-representation of non-cancer/negative classes (healthy tissue and/or background classes) compared to cancer/positive classes reduces the number of samples which are falsely predicted as cancer (false positive); (c) the least expensive to annotate non-ROI (non-region-of-interest) data can be useful in compensating for the performance loss in the system due to a shortage of expensive to annotate ROI data; (d) the multi-label examples are more useful than the single-label ones to train a segmentation model; and (e) when the classification model is tuned with a balanced validation set, it is less affected than the segmentation model by the class distribution of the training set.
      pubtype: Academic Journal
      doctype:
        equations & formulas
        pictorial
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N