Speech recognition in edge environments: an exploration of support and impact of model compression.

Automatic Speech Recognition (ASR) is an essential component of Human Computer Interaction. With the surge in mobile and home assistant IoT devices, there is a rising need for edge computing solutions supporting ASR applications. Whisper is a state-of-art ASR model that exhibits superior performance...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 60; no. 1; pp. 1 - 21
Autores principales: Joy, Carolene, Martin, John Paul, Joseph, Christina Terese, Madhavan, Manu
Formato: Artículo
Publicado: Springer Nature Mar2026
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=192212482&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 192212482
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Mar2026
      vid: 60
      iid: 1
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        192212482
        10.1007/s10579-025-09885-6
      ppf: 1
      ppct: 20
      formats:
      tig:
        atl: Speech recognition in edge environments: an exploration of support and impact of model compression.
      aug:
        au:
          Joy, Carolene
          Martin, John Paul
          Joseph, Christina Terese
          Madhavan, Manu
        affil:
          https://ror.org/00gcgw028 Government College of Engineering Kannur, Kannur, India
          Indian Institute of Information Technology Kottayam, Kottayam, India
      su:
        Automatic speech recognition
        Edge computing
        Internet of things
        Integer approximations
        Computer memory management
      sug:
        subj:
          Automatic speech recognition
          Edge computing
          Internet of things
          Integer approximations
          Computer memory management
      keyword:
        Edge device
        Information and Computing Sciences Artificial Intelligence and Image Processing
        Quantization
        Speech recognition
      ab: Automatic Speech Recognition (ASR) is an essential component of Human Computer Interaction. With the surge in mobile and home assistant IoT devices, there is a rising need for edge computing solutions supporting ASR applications. Whisper is a state-of-art ASR model that exhibits superior performance. Being resource hungry, this model fails to deliver similar performance on resource constrained devices. Quantization is a thoroughly studied model compression technique in the premise of Deep Learning, which reduces the model size by using smaller integer representations for weights and activations. In this work, quantization is applied to the Whisper releases which are then deployed on Raspberry Pi and Jetson Orin Nano. To get a fair understanding, the models are also deployed on an AWS EC2 instance as well as on a personal computing device. This work is one of the initial attempts to study the performance of quantized Whisper models on edge devices. The quantized models achieved compression ratios of over 3.7 with only a marginal degradation in transcription accuracy. Real-time performance improved significantly, with up to 15% reduction in inference latency observed for the Base model on Raspberry Pi 5. Furthermore, quantized models demonstrated a notable reduction in memory usage, enabling deployment on devices with limited computational resources.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      dt:
        @attributes:
          year: 2026
    holdings:
      @attributes:
        islocal: N