Cochlea-inspired speech recognition interface.
Automatic speech recognition (ASR) technology provides a natural interface for human-machine interaction. Typical ASR systems can achieve high performance in quiet environments but, unlike humans, perform poorly in real-world situations. To better simulate the human auditory periphery and improve th...
| Publicado en: | Medical & Biological Engineering & Computing Vol. 57; no. 6; pp. 1393 - 1404 |
|---|---|
| Autores principales: | , , , |
| Formato: | Journal Article |
| Publicado: |
Springer Nature
Jun2019
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=136505534&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 136505534 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 01400118 PO0 jtl: Medical & Biological Engineering & Computing issn: 01400118 maglogo: N pubinfo: dt: Jun2019 vid: 57 iid: 6 pid: 237 pub: Springer Nature place: New York, New York artinfo: ui: 136505534 136505534 NLM30830542 10.1007/s11517-019-01963-6 NLM30830542 136505534 ppf: 1393 ppct: 11 formats: fmt: – @attributes: type: T – @attributes: type: P tig: atl: Cochlea-inspired speech recognition interface. aug: au: Russo, Mladen Stella, Maja Sikora, Marjan Šarić, Matko affil: Laboratory for Smart Environment Technologies, FESB - University of Split, Split, Croatia sug: subj: Speech Physiology Cochlea Physiology Probability Signal Processing, Computer Assisted Sound Spectrography Biophysics Models, Biological Ferrans and Powers Quality of Life Index ab: Automatic speech recognition (ASR) technology provides a natural interface for human-machine interaction. Typical ASR systems can achieve high performance in quiet environments but, unlike humans, perform poorly in real-world situations. To better simulate the human auditory periphery and improve the performance in realistic noisy scenarios, we propose two models of speech recognition front-ends based on a biophysical cochlear model. The first front-end is based on the method of signal reconstruction from a basilar membrane response. When applied to noisy speech, this method results in improved signal quality. This method can be used as a preprocessing step in a standard ASR system and can also be used as a noise reduction technique for other applications. The second front-end we propose is based on the construction of speech recognition coefficients directly from a basilar membrane response. Experimental results using a continuous-density hidden Markov model (HMM) recognizer demonstrate significant improvement in performance compared to standard Mel-frequency cepstral coefficients (MFCC) in various types of noisy conditions. Graphical Abstract Speech recognition model based on cochlear front-end. pubtype: Academic Journal doctype: Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|