A Robust Supervised Variable Selection for Noisy High-Dimensional Data.

The Minimum Redundancy Maximum Relevance (MRMR) approach to supervised variable selection represents a successful methodology for dimensionality reduction, which is suitable for high-dimensional data observed in two or more different groups. Various available versions of the MRMR approach have been...

Descripción completa

Detalles Bibliográficos
Publicado en:BioMed Research International Vol. 2015; pp. 1 - 11
Autores principales: Kalina, Jan, Schlenker, Anna
Formato: equations & formulas research tables/charts Journal Article
Publicado: Wiley-Blackwell 6/2/2015
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=109274457&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 109274457
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        23146133
        FT2T
      jtl: BioMed Research International
      issn: 23146133
      maglogo: N
    pubinfo:
      dt: 6/2/2015
      vid: 2015
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        109274457
        109274457
        109274457
        10.1155/2015/320385
        109274457
      ppf: 1
      ppct: 10
      formats:
        fmt:
          @attributes:
            type: P
      tig:
        atl: A Robust Supervised Variable Selection for Noisy High-Dimensional Data.
      aug:
        au:
          Kalina, Jan
          Schlenker, Anna
        affil: Institute of Computer Science of the Czech Academy of Sciences, Pod Vodárenskou Vĕží 2, 182 07 Prague 8, Czech Republic
      sug:
        subj:
          Data Management Methods
          Data Analysis, Statistical
          Funding Source
      ab: The Minimum Redundancy Maximum Relevance (MRMR) approach to supervised variable selection represents a successful methodology for dimensionality reduction, which is suitable for high-dimensional data observed in two or more different groups. Various available versions of the MRMR approach have been designed to search for variables with the largest relevance for a classification task while controlling for redundancy of the selected set of variables. However, usual relevance and redundancy criteria have the disadvantages of being too sensitive to the presence of outlying measurements and/or being inefficient. We propose a novel approach called Minimum Regularized Redundancy Maximum Robust Relevance (MRRMRR), suitable for noisy high-dimensional data observed in two groups. It combines principles of regularization and robust statistics. Particularly, redundancy is measured by a new regularized version of the coefficient of multiple correlation and relevance is measured by a highly robust correlation coefficient based on the least weighted squares regression with data-adaptive weights. We compare various dimensionality reduction methods on three real data sets. To investigate the influence of noise or outliers on the data, we perform the computations also for data artificially contaminated by severe noise of various forms. The experimental results confirm the robustness of the method with respect to outliers.
      pubtype: Academic Journal
      doctype:
        equations & formulas
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N