Classifying Human Voices by Using Hybrid SFX Time-Series Preprocessing and Ensemble Feature Selection.
Voice biometrics is one kind of physiological characteristics whose voice is different for each individual person. Due to this uniqueness, voice classification has found useful applications in classifying speakers' gender, mother tongue or ethnicity (accent), emotion states, identity verification, v...
| Publicado en: | BioMed Research International Vol. 2013; pp. 720834 - 720835 |
|---|---|
| Autores principales: | , , |
| Formato: | Journal Article |
| Publicado: |
Wiley-Blackwell
2013
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=104120708&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 104120708 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 23146133 FT2T jtl: BioMed Research International issn: 23146133 maglogo: N pubinfo: dt: 2013 vid: 2013 pid: 480 pub: Wiley-Blackwell place: Malden, Massachusetts artinfo: ui: 104120708 2012401204 NLM24288684 PMC3830839 104120708 ppf: 720834 ppct: 1 formats: fmt: @attributes: type: P tig: atl: Classifying Human Voices by Using Hybrid SFX Time-Series Preprocessing and Ensemble Feature Selection. aug: au: Fong, Simon Lan, Kun Wong, Raymond affil: Department of Computer and Information Science, University of Macau, Macau. sug: subj: Algorithms Data Mining Reproduction Signal Processing, Computer Assisted Speech Acoustics Voice Female Male Female Male ab: Voice biometrics is one kind of physiological characteristics whose voice is different for each individual person. Due to this uniqueness, voice classification has found useful applications in classifying speakers' gender, mother tongue or ethnicity (accent), emotion states, identity verification, verbal command control, and so forth. In this paper, we adopt a new preprocessing method named Statistical Feature Extraction (SFX) for extracting important features in training a classification model, based on piecewise transformation treating an audio waveform as a time-series. Using SFX we can faithfully remodel statistical characteristics of the time-series; together with spectral analysis, a substantial amount of features are extracted in combination. An ensemble is utilized in selecting only the influential features to be used in classification model induction. We focus on the comparison of effects of various popular data mining algorithms on multiple datasets. Our experiment consists of classification tests over four typical categories of human voice data, namely, Female and Male, Emotional Speech, Speaker Identification, and Language Recognition. The experiments yield encouraging results supporting the fact that heuristically choosing significant features from both time and frequency domains indeed produces better performance in voice classification than traditional signal processing techniques alone, like wavelets and LPC-to-CC. pubtype: Academic Journal doctype: Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|