Pathological Voice Detection and Classification Based on Multimodal Transmission Network.
Describing pronunciation features from multiple perspectives can help doctors accurately diagnose the pathological type of a patient's voice. According to the two modal information of sound signal and electroglottography (EGG) signal, this paper proposes a pathological voice detection and classifica...
| Publicado en: | Journal of Voice Vol. 39; no. 3; pp. 591 - 602 |
|---|---|
| Autores principales: | , , , , , |
| Formato: | research Journal Article |
| Publicado: |
Elsevier B.V.
May2025
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=184889992&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 184889992 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 08921997 H24 jtl: Journal of Voice issn: 08921997 maglogo: N pubinfo: dt: May2025 vid: 39 iid: 3 pid: 467 pub: Elsevier B.V. place: New York, New York artinfo: ui: 184889992 184889992 184889992 10.1016/j.jvoice.2022.11.018 184889992 ppf: 591 ppct: 11 formats: tig: atl: Pathological Voice Detection and Classification Based on Multimodal Transmission Network. aug: au: Geng, Lei Liang, Yan Shan, Hongfeng Xiao, Zhitao Wang, Wei Wei, Mei affil: School of Life Sciences, Tiangong University, Tianjin, China sug: subj: Voice Disorders Diagnosis Voice Disorders Classification Voice Disorders Physiopathology Speech Production Measurement Methods Electrodiagnosis Signal Processing, Computer Assisted Voice Quality Human Convolutional Neural Networks Acoustics Algorithms Models, Statistical Predictive Value of Tests Sound Spectrography Speech Acoustics ab: Describing pronunciation features from multiple perspectives can help doctors accurately diagnose the pathological type of a patient's voice. According to the two modal information of sound signal and electroglottography (EGG) signal, this paper proposes a pathological voice detection and classification algorithm based on multimodal transmission network. Firstly, we used the short-time Fourier transform (STFT) to map the features of the two signals, and designed the Mel filter to obtain the Mel spectogram. Then, the constructed multimodal transmission network extracted features from Mel spectogram and applied Multimodal Transfer Module (MMTM) module. Finally, the fusion layer can integrate multimodal information, and the full connection layer diagnoses and classifies voice pathology according to the fused features. The experiment was based on 1179 subjects in Saarbrücken voice database (SVD), and the average accuracy, recall, specificity and F1 score of pathological voice classification reached 98.02%, 98.23%, 97.82% and 97.95% respectively. Compared with other algorithms, the classification accuracy is significantly improved. The proposed model can integrate multiple modal information to obtain more comprehensive and stable voice features and improve the accuracy of pathological voice classification. Future research will further explore in reducing the time-consuming and complexity of the model. pubtype: Academic Journal doctype: research Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|