Pathological Voice Detection and Classification Based on Multimodal Transmission Network.

Describing pronunciation features from multiple perspectives can help doctors accurately diagnose the pathological type of a patient's voice. According to the two modal information of sound signal and electroglottography (EGG) signal, this paper proposes a pathological voice detection and classifica...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Voice Vol. 39; no. 3; pp. 591 - 602
Autores principales: Geng, Lei, Liang, Yan, Shan, Hongfeng, Xiao, Zhitao, Wang, Wei, Wei, Mei
Formato: research Journal Article
Publicado: Elsevier B.V. May2025
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=184889992&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 184889992
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        08921997
        H24
      jtl: Journal of Voice
      issn: 08921997
      maglogo: N
    pubinfo:
      dt: May2025
      vid: 39
      iid: 3
      pid: 467
      pub: Elsevier B.V.
      place: New York, New York
    artinfo:
      ui:
        184889992
        184889992
        184889992
        10.1016/j.jvoice.2022.11.018
        184889992
      ppf: 591
      ppct: 11
      formats:
      tig:
        atl: Pathological Voice Detection and Classification Based on Multimodal Transmission Network.
      aug:
        au:
          Geng, Lei
          Liang, Yan
          Shan, Hongfeng
          Xiao, Zhitao
          Wang, Wei
          Wei, Mei
        affil: School of Life Sciences, Tiangong University, Tianjin, China
      sug:
        subj:
          Voice Disorders Diagnosis
          Voice Disorders Classification
          Voice Disorders Physiopathology
          Speech Production Measurement Methods
          Electrodiagnosis
          Signal Processing, Computer Assisted
          Voice Quality
          Human
          Convolutional Neural Networks
          Acoustics
          Algorithms
          Models, Statistical
          Predictive Value of Tests
          Sound Spectrography
          Speech Acoustics
      ab: Describing pronunciation features from multiple perspectives can help doctors accurately diagnose the pathological type of a patient's voice. According to the two modal information of sound signal and electroglottography (EGG) signal, this paper proposes a pathological voice detection and classification algorithm based on multimodal transmission network. Firstly, we used the short-time Fourier transform (STFT) to map the features of the two signals, and designed the Mel filter to obtain the Mel spectogram. Then, the constructed multimodal transmission network extracted features from Mel spectogram and applied Multimodal Transfer Module (MMTM) module. Finally, the fusion layer can integrate multimodal information, and the full connection layer diagnoses and classifies voice pathology according to the fused features. The experiment was based on 1179 subjects in Saarbrücken voice database (SVD), and the average accuracy, recall, specificity and F1 score of pathological voice classification reached 98.02%, 98.23%, 97.82% and 97.95% respectively. Compared with other algorithms, the classification accuracy is significantly improved. The proposed model can integrate multiple modal information to obtain more comprehensive and stable voice features and improve the accuracy of pathological voice classification. Future research will further explore in reducing the time-consuming and complexity of the model.
      pubtype: Academic Journal
      doctype:
        research
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N