Distance Metric Based Oversampling Method for Bioinformatics and Performance Evaluation.

An imbalanced classification means that a dataset has an unequal class distribution among its population. For any given dataset, regardless of any balancing issue, the predictions made by most classification methods are highly accurate for the majority class but significantly less accurate for the m...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Medical Systems Vol. 40; no. 7; pp. 1 - 10
Autores principales: Tsai, Meng-Fong, Yu, Shyr-Shen
Formato: equations & formulas research tables/charts Journal Article
Publicado: Springer Nature Jul2016
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=115925380&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 115925380
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        01485598
        4N0
      jtl: Journal of Medical Systems
      issn: 01485598
      maglogo: N
    pubinfo:
      dt: Jul2016
      vid: 40
      iid: 7
      pid: 237
      pub: Springer Nature
      place: New York, New York
    artinfo:
      ui:
        115925380
        115925380
        115925380
        10.1007/s10916-016-0516-3
        115925380
      ppf: 1
      ppct: 9
      formats:
        fmt:
          @attributes:
            type: P
      tig:
        atl: Distance Metric Based Oversampling Method for Bioinformatics and Performance Evaluation.
      aug:
        au:
          Tsai, Meng-Fong
          Yu, Shyr-Shen
        affil: Department of Computer Science and Engineering, National Chung Hsing University, Taichung 402 Taiwan
      sug:
        subj:
          Bioinformatics
          Algorithms
          Decision Support Techniques
          Human
          Descriptive Statistics
          Survival Analysis
          Breast Neoplasms Prognosis
          Female
          Female
      ab: An imbalanced classification means that a dataset has an unequal class distribution among its population. For any given dataset, regardless of any balancing issue, the predictions made by most classification methods are highly accurate for the majority class but significantly less accurate for the minority class. To overcome this problem, this study took several imbalanced datasets from the famed UCI datasets and designed and implemented an efficient algorithm which couples Top-N Reverse k-Nearest Neighbor (TR kNN) with the Synthetic Minority Oversampling TEchnique (SMOTE). The proposed algorithm was investigated by applying it to classification methods such as logistic regression (LR), C4.5, Support Vector Machine (SVM), and Back Propagation Neural Network (BPNN). This research also adopted different distance metrics to classify the same UCI datasets. The empirical results illustrate that the Euclidean and Manhattan distances are not only more accurate, but also show greater computational efficiency when compared to the Chebyshev and Cosine distances. Therefore, the proposed algorithm based on TR kNN and SMOTE can be widely used to handle imbalanced datasets. Our recommendations on choosing suitable distance metrics can also serve as a reference for future studies.
      pubtype: Academic Journal
      doctype:
        equations & formulas
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N