Classifying Human Voices by Using Hybrid SFX Time-Series Preprocessing and Ensemble Feature Selection.

Voice biometrics is one kind of physiological characteristics whose voice is different for each individual person. Due to this uniqueness, voice classification has found useful applications in classifying speakers' gender, mother tongue or ethnicity (accent), emotion states, identity verification, v...

Descripción completa

Detalles Bibliográficos
Publicado en:BioMed Research International Vol. 2013; pp. 720834 - 720835
Autores principales: Fong, Simon, Lan, Kun, Wong, Raymond
Formato: Journal Article
Publicado: Wiley-Blackwell 2013
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=104120708&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 104120708
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        23146133
        FT2T
      jtl: BioMed Research International
      issn: 23146133
      maglogo: N
    pubinfo:
      dt: 2013
      vid: 2013
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        104120708
        2012401204
        NLM24288684
        PMC3830839
        104120708
      ppf: 720834
      ppct: 1
      formats:
        fmt:
          @attributes:
            type: P
      tig:
        atl: Classifying Human Voices by Using Hybrid SFX Time-Series Preprocessing and Ensemble Feature Selection.
      aug:
        au:
          Fong, Simon
          Lan, Kun
          Wong, Raymond
        affil: Department of Computer and Information Science, University of Macau, Macau.
      sug:
        subj:
          Algorithms
          Data Mining
          Reproduction
          Signal Processing, Computer Assisted
          Speech Acoustics
          Voice
          Female
          Male
          Female
          Male
      ab: Voice biometrics is one kind of physiological characteristics whose voice is different for each individual person. Due to this uniqueness, voice classification has found useful applications in classifying speakers' gender, mother tongue or ethnicity (accent), emotion states, identity verification, verbal command control, and so forth. In this paper, we adopt a new preprocessing method named Statistical Feature Extraction (SFX) for extracting important features in training a classification model, based on piecewise transformation treating an audio waveform as a time-series. Using SFX we can faithfully remodel statistical characteristics of the time-series; together with spectral analysis, a substantial amount of features are extracted in combination. An ensemble is utilized in selecting only the influential features to be used in classification model induction. We focus on the comparison of effects of various popular data mining algorithms on multiple datasets. Our experiment consists of classification tests over four typical categories of human voice data, namely, Female and Male, Emotional Speech, Speaker Identification, and Language Recognition. The experiments yield encouraging results supporting the fact that heuristically choosing significant features from both time and frequency domains indeed produces better performance in voice classification than traditional signal processing techniques alone, like wavelets and LPC-to-CC.
      pubtype: Academic Journal
      doctype: Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N