A Deep Learning Enhanced Novel Software Tool for Laryngeal Dynamics Analysis.

Purpose: High-speed videoendoscopy (HSV) is an emerging, but barely used, endoscopy technique in the clinic to assess and diagnose voice disorders because of the lack of dedicated software to analyze the data. HSV allows to quantify the vocal fold oscillations by segmenting the glottal area. This ch...

Full description

Bibliographic Details
Published in:Journal of Speech, Language & Hearing Research Vol. 64; no. 6; pp. 1889 - 1904
Main Authors: Kist, Andreas M., Gómez, Pablo, Dubrovskiy, Denis, Schlegel, Patrick, Kunduk, Melda, Echternach, Matthias, Patel, Rita, Semmler, Marion, Stephan Dürr, Christopher Bohr,e, Schützenberger, Anne, Döllinger, Michael
Format: Article
Published: American Speech-Language-Hearing Association Jun2021
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=150778854&site=ehost-live
header:
  @attributes:
    shortDbName: ssf
    uiTerm: 150778854
    longDbName: Social Sciences Full Text (H.W. Wilson)
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        10924388
        1SM
      jtl: Journal of Speech, Language & Hearing Research
      issn: 10924388
      maglogo: N
    pubinfo:
      dt: Jun2021
      vid: 64
      iid: 6
      pid: 42
      pub: American Speech-Language-Hearing Association
    artinfo:
      ui:
        150778854
        10.1044/2021_JSLHR-20-00498
      ppf: 1889
      ppct: 15
      formats:
        fmt:
          @attributes:
            type: P
            size: 1.3MB
      tig:
        atl: A Deep Learning Enhanced Novel Software Tool for Laryngeal Dynamics Analysis.
      aug:
        au:
          Kist, Andreas M.
          Gómez, Pablo
          Dubrovskiy, Denis
          Schlegel, Patrick
          Kunduk, Melda
          Echternach, Matthias
          Patel, Rita
          Semmler, Marion
          Stephan Dürr, Christopher Bohr,e
          Schützenberger, Anne
          Döllinger, Michael
        affil:
          Division of Phoniatrics and Pediatric Audiology, Department of Otorhinolaryngology—Head & Neck Surgery, University Hospital Erlangen, Germany.
          Department of Communication Sciences and Disorders, Louisiana State University, Baton Rouge.
          Division of Phoniatrics and Pediatric Audiology, Department of Otorhinolaryngology, Munich University Hospital (LMU), Germany.
          Department of Speech, Language and Hearing Sciences, College of Arts and Sciences, Indiana University, Bloomington.
      su:
        Deep learning
        Computer software
        Glottis
        Artificial neural networks
        Video recording
        Graphical user interfaces
        Algorithms
      sug:
        subj:
          Computer, computer peripheral and pre-packaged software merchant wholesalers
          Computer and Computer Peripheral Equipment and Software Merchant Wholesalers
          Computer and software stores
          Software publishers (except video game publishers)
          Deep learning
          Computer software
          Glottis
          Artificial neural networks
          Video recording
          Graphical user interfaces
          Algorithms
      ab: Purpose: High-speed videoendoscopy (HSV) is an emerging, but barely used, endoscopy technique in the clinic to assess and diagnose voice disorders because of the lack of dedicated software to analyze the data. HSV allows to quantify the vocal fold oscillations by segmenting the glottal area. This challenging task has been tackled by various studies; however, the proposed approaches are mostly limited and not suitable for daily clinical routine. Method: We developed a user-friendly software in C# that allows the editing, motion correction, segmentation, and quantitative analysis of HSV data. We further provide pretrained deep neural networks for fully automatic glottis segmentation. Results: We freely provide our software Glottis Analysis Tools (GAT). Using GAT, we provide a general thresholdbased region growing platform that enables the user to analyze data from various sources, such as in vivo recordings, ex vivo recordings, and high-speed footage of artificial vocal folds. Additionally, especially for in vivo recordings, we provide three robust neural networks at various speed and quality settings to allow a fully automatic glottis segmentation needed for application by untrained personnel. GAT further evaluates video and audio data in parallel and is able to extract various features from the video data, among others the glottal area waveform, that is, the changing glottal area over time. In total, GAT provides 79 unique quantitative analysis parameters for video- and audio-based signals. Many of these parameters have already been shown to reflect voice disorders, highlighting the clinical importance and usefulness of the GAT software. Conclusion: GAT is a unique tool to process HSV and audio data to determine quantitative, clinically relevant parameters for research, diagnosis, and treatment of laryngeal disorders.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: N
    holdings:
      @attributes:
        islocal: N