Automation of Language Sample Analysis.

Purpose: A major barrier to the wider use of language sample analysis (LSA) is the fact that transcription is very time intensive. Methods that can reduce the required time and effort could help in promoting the use of LSA for clinical prac- tice and research. Method: This article describes an autom...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Speech, Language & Hearing Research Vol. 66; no. 7; pp. 2421 - 2434
Autores principales: Houjun Liu, MacWhinney, Brian, Fromm, Davida, Lanzi, Alyssa
Formato: Artículo
Publicado: American Speech-Language-Hearing Association Jul2023
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=164881408&site=ehost-live
header:
  @attributes:
    shortDbName: ssf
    uiTerm: 164881408
    longDbName: Social Sciences Full Text (H.W. Wilson)
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        10924388
        1SM
      jtl: Journal of Speech, Language & Hearing Research
      issn: 10924388
      maglogo: N
    pubinfo:
      dt: Jul2023
      vid: 66
      iid: 7
      pid: 42
      pub: American Speech-Language-Hearing Association
    artinfo:
      ui:
        164881408
        10.1044/2023_JSLHR-22-00642
      ppf: 2421
      ppct: 13
      formats:
        fmt:
          @attributes:
            type: P
            size: 705KB
      tig:
        atl: Automation of Language Sample Analysis.
      aug:
        au:
          Houjun Liu
          MacWhinney, Brian
          Fromm, Davida
          Lanzi, Alyssa
        affil:
          The Nueva School, San Mateo, CA.
          Department of Psychology, Carnegie Mellon University, Pittsburgh, PA.
          Communication Sciences and Disorders Department, University of Delaware, Newark.
      su:
        Computer software
        Speech therapy
        Phonological awareness
        Speech evaluation
        Automatic speech recognition
        Acquisition of data
        Articulation disorders
        Automation
      sug:
        subj:
          Computer, computer peripheral and pre-packaged software merchant wholesalers
          Computer and Computer Peripheral Equipment and Software Merchant Wholesalers
          Computer and software stores
          Software publishers (except video game publishers)
          Computer software
          Speech therapy
          Phonological awareness
          Speech evaluation
          Automatic speech recognition
          Acquisition of data
          Articulation disorders
          Automation
      ab: Purpose: A major barrier to the wider use of language sample analysis (LSA) is the fact that transcription is very time intensive. Methods that can reduce the required time and effort could help in promoting the use of LSA for clinical prac- tice and research. Method: This article describes an automated pipeline, called Batchalign, that takes raw audio and creates full transcripts in Codes for the Human Analysis of Talk (CHAT) transcription format, complete with utterance- and word-level time alignments and morphosyntactic analysis. The pipeline only requires major human intervention for final checking. It combines a series of existing tools with additional novel reformatting processes. The steps in the pipeline are (a) auto- matic speech recognition, (b) utterance tokenization, (c) automatic corrections, (d) speaker ID assignment, (e) forced alignment, (f) user adjustments, and (g) automatic morphosyntactic and profiling analyses. Results: For work with recordings from adults with language disorders, six major results were obtained: (a) The word error rate was between 2.4% for con- trols and 3.4% for patients, (b) utterance tokenization accuracy was at the level reported for speakers without language disorders, (c) word-level diarization accuracy was at 93% for control participants and 83% for participants with lan- guage disorders, (d) utterance-level diarization accuracy based on word-level diarization was high, (e) adherence to CHAT format was fully accurate, and (f) human transcriber time was reduced by up to 75%. Conclusion: The pipeline dramatically shortens the time gap between data col- lection and data analysis and provides an output superior to that typically gen- erated by human transcribers.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: N
    holdings:
      @attributes:
        islocal: N