Learning from Imbalanced Data: Integration of Advanced Resampling Techniques and Machine Learning Models for Enhanced Cancer Diagnosis and Prognosis.

Simple Summary: This research focuses on improving cancer diagnosis and prognosis by addressing a common problem in data analysis known as class imbalance, where some patient groups are underrepresented. The authors aim to evaluate different resampling methods that can balance the data and enhance t...

Descripción completa

Detalles Bibliográficos
Publicado en:Cancers Vol. 16; no. 19; pp. 3417 - 3436
Autores principales: Gurcan, Fatih, Soylu, Ahmet
Formato: research tables/charts Journal Article
Publicado: MDPI Oct2024
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=180274314&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 180274314
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        20726694
        B74B
      jtl: Cancers
      issn: 20726694
      maglogo: N
    pubinfo:
      dt: Oct2024
      vid: 16
      iid: 19
      pid: 97109
      pub: MDPI
    artinfo:
      ui:
        180274314
        180274314
        180274314
        10.3390/cancers16193417
        180274314
      ppf: 3417
      ppct: 19
      formats:
      tig:
        atl: Learning from Imbalanced Data: Integration of Advanced Resampling Techniques and Machine Learning Models for Enhanced Cancer Diagnosis and Prognosis.
      aug:
        au:
          Gurcan, Fatih
          Soylu, Ahmet
        affil: Department of Management Information Systems, Faculty of Economics and Administrative Sciences, Karadeniz Technical University, 61080 Trabzon, Turkey
      sug:
        subj:
          Neoplasms Diagnosis
          Neoplasms Prognosis
          Machine Learning
          Diagnosis, Computer Assisted
          Image Enhancement
          Sampling Methods
          Algorithms
          Prediction Models
          Oncology
          Reference Databases
          Human
          Comparative Studies
          Descriptive Statistics
          Random Forest
          Oncologists
          Early Detection of Cancer
          Cancer Screening
      ab: Simple Summary: This research focuses on improving cancer diagnosis and prognosis by addressing a common problem in data analysis known as class imbalance, where some patient groups are underrepresented. The authors aim to evaluate different resampling methods that can balance the data and enhance the performance of various classification algorithms used to predict cancer outcomes. By testing a wide range of techniques across multiple cancer datasets, this study identifies the best-performing classifier, Random Forest, along with the most effective resampling method, SMOTEENN. These findings provide valuable insights for researchers and healthcare professionals, enabling them to make more accurate predictions and ultimately improve patient care. This research could pave the way for the development of more reliable machine learning applications in the medical field. Background/Objectives: This study aims to evaluate the performance of various classification algorithms and resampling methods across multiple diagnostic and prognostic cancer datasets, addressing the challenges of class imbalance. Methods: A total of five datasets were analyzed, including three diagnostic datasets (Wisconsin Breast Cancer Database, Cancer Prediction Dataset, Lung Cancer Detection Dataset) and two prognostic datasets (Seer Breast Cancer Dataset, Differentiated Thyroid Cancer Recurrence Dataset). Nineteen resampling methods from three categories were employed, and ten classifiers from four distinct categories were utilized for comparison. Results: The results demonstrated that hybrid sampling methods, particularly SMOTEENN, achieved the highest mean performance at 98.19%, followed by IHT (97.20%) and RENN (96.48%). In terms of classifiers, Random Forest showed the best performance with a mean value of 94.69%, with Balanced Random Forest and XGBoost following closely. The baseline method (no resampling) yielded a significantly lower performance of 91.33%, highlighting the effectiveness of resampling techniques in improving model outcomes. Conclusions: This research underscores the importance of resampling methods in enhancing classification performance on imbalanced datasets, providing valuable insights for researchers and healthcare professionals. The findings serve as a foundation for future studies aimed at integrating machine learning techniques in cancer diagnosis and prognosis, with recommendations for further research on hybrid models and clinical applications.
      pubtype: Academic Journal
      doctype:
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N