Multiple Imputation based Clustering Validation (MIV) for Big Longitudinal Trial Data with Missing Values in eHealth.

Web-delivered trials are an important component in eHealth services. These trials, mostly behavior-based, generate big heterogeneous data that are longitudinal, high dimensional with missing values. Unsupervised learning methods have been widely applied in this area, however, validating the optimal...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Medical Systems Vol. 40; no. 6; pp. 1 - 10
Autores principales: Zhang, Zhaoyang, Fang, Hua, Wang, Honggang
Formato: algorithm equations & formulas tables/charts Journal Article
Publicado: Springer Nature Jun2016
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=115925350&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 115925350
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        01485598
        4N0
      jtl: Journal of Medical Systems
      issn: 01485598
      maglogo: N
    pubinfo:
      dt: Jun2016
      vid: 40
      iid: 6
      pid: 237
      pub: Springer Nature
      place: New York, New York
    artinfo:
      ui:
        115925350
        115925350
        115925350
        10.1007/s10916-016-0499-0
        115925350
      ppf: 1
      ppct: 9
      formats:
        fmt:
          @attributes:
            type: P
      tig:
        atl: Multiple Imputation based Clustering Validation (MIV) for Big Longitudinal Trial Data with Missing Values in eHealth.
      aug:
        au:
          Zhang, Zhaoyang
          Fang, Hua
          Wang, Honggang
        affil: Department of Quantitative Health Science, University of Massachusetts Medical School, Worcester 01655 USA
      sug:
        subj:
          Data Analytics
          Algorithms
          Clinical Trials
          Human
          Smoking Cessation
          Computer Simulation
          Funding Source
      ab: Web-delivered trials are an important component in eHealth services. These trials, mostly behavior-based, generate big heterogeneous data that are longitudinal, high dimensional with missing values. Unsupervised learning methods have been widely applied in this area, however, validating the optimal number of clusters has been challenging. Built upon our multiple imputation (MI) based fuzzy clustering, MIfuzzy, we proposed a new multiple imputation based validation (MIV) framework and corresponding MIV algorithms for clustering big longitudinal eHealth data with missing values, more generally for fuzzy-logic based clustering methods. Specifically, we detect the optimal number of clusters by auto-searching and -synthesizing a suite of MI-based validation methods and indices, including conventional (bootstrap or cross-validation based) and emerging (modularity-based) validation indices for general clustering methods as well as the specific one (Xie and Beni) for fuzzy clustering. The MIV performance was demonstrated on a big longitudinal dataset from a real web-delivered trial and using simulation. The results indicate MI-based Xie and Beni index for fuzzy-clustering are more appropriate for detecting the optimal number of clusters for such complex data. The MIV concept and algorithms could be easily adapted to different types of clustering that could process big incomplete longitudinal trial data in eHealth services.
      pubtype: Academic Journal
      doctype:
        algorithm
        equations & formulas
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N