Whisper-Small: Optimized for Qualcomm Devices

HuggingFace Whisper-Small ASR (Automatic Speech Recognition) model is a state-of-the-art system designed for transcribing spoken language into written text. This model is based on the transformer architecture and has been optimized for edge inference by replacing Multi-Head Attention (MHA) with Single-Head Attention (SHA) and linear layers with convolutional (conv) layers. It exhibits robust performance in realistic, noisy environments, making it highly reliable for real-world applications. Specifically, it excels in long-form transcription, capable of accurately transcribing audio clips up to 30 seconds long. Time to the first token is the encoder's latency, while time to each additional token is decoder's latency, where we assume a max decoded length specified below.

This is based on the implementation of Whisper-Small found here. This repository contains pre-exported model files optimized for Qualcomm® devices. You can use the Qualcomm® AI Hub Models library to export with custom configurations. More details on model performance across various devices, can be found here.

Qualcomm AI Hub Models uses Qualcomm AI Hub Workbench to compile, profile, and evaluate this model. Sign up to run these models on a hosted Qualcomm® device.

Deploying Whisper-Small on-device

This model is compatible with the Qualcomm Voice AI SDK. Download the SDK from the Qualcomm Package Manager to deploy this model on-device.

Getting Started

There are two ways to deploy this model on your device:

Option 1: Download Pre-Exported Models

Below are pre-exported model assets ready for deployment.

Runtime Precision Chipset SDK Versions Download
PRECOMPILED_QNN_ONNX float Snapdragon® 8 Elite Gen 5 For Galaxy Mobile QAIRT 2.50, ONNX Runtime 1.30.0 Download
PRECOMPILED_QNN_ONNX float Snapdragon® 8 Elite For Galaxy Mobile QAIRT 2.50, ONNX Runtime 1.30.0 Download
PRECOMPILED_QNN_ONNX float Snapdragon® X2 Elite QAIRT 2.50, ONNX Runtime 1.30.0 Download
PRECOMPILED_QNN_ONNX float Snapdragon® X Elite QAIRT 2.50, ONNX Runtime 1.30.0 Download
PRECOMPILED_QNN_ONNX float Snapdragon® 8 Gen 3 Mobile QAIRT 2.50, ONNX Runtime 1.30.0 Download
PRECOMPILED_QNN_ONNX float Snapdragon® 8 Gen 1 Mobile QAIRT 2.50, ONNX Runtime 1.30.0 Download
PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ IQ-8275 QAIRT 2.50, ONNX Runtime 1.30.0 Download
PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ QCS8550 (Proxy) QAIRT 2.50, ONNX Runtime 1.30.0 Download
PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ IQ-9075 QAIRT 2.50, ONNX Runtime 1.30.0 Download
QNN_CONTEXT_BINARY float Snapdragon® 8 Elite Gen 5 For Galaxy Mobile QAIRT 2.50 Download
QNN_CONTEXT_BINARY float Snapdragon® 8 Elite For Galaxy Mobile QAIRT 2.50 Download
QNN_CONTEXT_BINARY float Snapdragon® X2 Elite QAIRT 2.50 Download
QNN_CONTEXT_BINARY float Snapdragon® X Elite QAIRT 2.50 Download
QNN_CONTEXT_BINARY float Snapdragon® 8 Gen 3 Mobile QAIRT 2.50 Download
QNN_CONTEXT_BINARY float Snapdragon® 8 Gen 1 Mobile QAIRT 2.50 Download
QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ IQ-8275 QAIRT 2.50 Download
QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ QCS8550 (Proxy) QAIRT 2.50 Download
QNN_CONTEXT_BINARY float Qualcomm® SA8775P QAIRT 2.50 Download
QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ IQ-9075 QAIRT 2.50 Download
QNN_CONTEXT_BINARY float Qualcomm® SA7255P QAIRT 2.50 Download
QNN_CONTEXT_BINARY float Qualcomm® SA8295P QAIRT 2.50 Download
VOICE_AI float Snapdragon® 8 Elite Gen 5 For Galaxy Mobile QAIRT 2.50 Download
VOICE_AI float Snapdragon® 8 Elite For Galaxy Mobile QAIRT 2.50 Download
VOICE_AI float Snapdragon® X2 Elite QAIRT 2.50 Download
VOICE_AI float Snapdragon® X Elite QAIRT 2.50 Download
VOICE_AI float Snapdragon® 8 Gen 3 Mobile QAIRT 2.50 Download
VOICE_AI float Snapdragon® 8 Gen 1 Mobile QAIRT 2.50 Download
VOICE_AI float Qualcomm® Dragonwing™ IQ-8275 QAIRT 2.50 Download
VOICE_AI float Qualcomm® Dragonwing™ QCS8550 (Proxy) QAIRT 2.50 Download
VOICE_AI float Qualcomm® SA8775P QAIRT 2.50 Download
VOICE_AI float Qualcomm® Dragonwing™ IQ-9075 QAIRT 2.50 Download
VOICE_AI float Qualcomm® SA7255P QAIRT 2.50 Download
VOICE_AI float Qualcomm® SA8295P QAIRT 2.50 Download

For more device-specific assets and performance metrics, visit Whisper-Small on Qualcomm® AI Hub.

Option 2: Export with Custom Configurations

Use the Qualcomm® AI Hub Models Python library to compile and export the model with your own:

  • Custom weights (e.g., fine-tuned checkpoints)
  • Custom input shapes
  • Target device and runtime configurations

This option is ideal if you need to customize the model beyond the default configuration provided here.

See our repository for Whisper-Small on GitHub for usage instructions.

Model Details

Model Type: Model_use_case.speech_recognition

Model Stats:

  • Input resolution: 80x3000 (30 seconds audio)
  • Max decoded sequence length: 200 tokens
  • Model checkpoint: openai/whisper-small
  • Model size (decoder) (float): 533 MB
  • Model size (encoder) (float): 391 MB
  • Number of parameters (decoder): 139M
  • Number of parameters (encoder): 102M

Performance Summary

Model Runtime Precision Chipset Inference Time (ms) Peak Memory Range (MB) Primary Compute Unit
decoder PRECOMPILED_QNN_ONNX float Snapdragon® 8 Elite Gen 5 For Galaxy Mobile 7.356 ms 43 - 54 MB NPU
decoder PRECOMPILED_QNN_ONNX float Snapdragon® 8 Elite For Galaxy Mobile 8.407 ms 57 - 69 MB NPU
decoder PRECOMPILED_QNN_ONNX float Snapdragon® X2 Elite 6.116 ms 60 - 60 MB NPU
decoder PRECOMPILED_QNN_ONNX float Snapdragon® X Elite 10.449 ms 287 - 287 MB NPU
decoder PRECOMPILED_QNN_ONNX float Snapdragon® 8 Gen 3 Mobile 10.121 ms 75 - 87 MB NPU
decoder PRECOMPILED_QNN_ONNX float Snapdragon® 8 Gen 1 Mobile 14.284 ms 75 - 90 MB NPU
decoder PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ IQ-8275 14.781 ms 60 - 124 MB NPU
decoder PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ QCS8550 (Proxy) 12.419 ms 0 - 318 MB NPU
decoder PRECOMPILED_QNN_ONNX float Qualcomm® QCS8450 14.284 ms 75 - 90 MB NPU
decoder PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ IQ-9075 13.709 ms 60 - 123 MB NPU
decoder PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ IQ-X7181 10.449 ms 287 - 287 MB NPU
decoder PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ Q-8750 8.407 ms 57 - 69 MB NPU
decoder QNN_CONTEXT_BINARY float Snapdragon® 8 Elite Gen 5 For Galaxy Mobile 7.338 ms 1 - 10 MB NPU
decoder QNN_CONTEXT_BINARY float Snapdragon® 8 Elite For Galaxy Mobile 8.331 ms 0 - 8 MB NPU
decoder QNN_CONTEXT_BINARY float Snapdragon® X2 Elite 6.665 ms 60 - 60 MB NPU
decoder QNN_CONTEXT_BINARY float Snapdragon® X Elite 11.241 ms 60 - 60 MB NPU
decoder QNN_CONTEXT_BINARY float Snapdragon® 8 Gen 3 Mobile 9.835 ms 60 - 68 MB NPU
decoder QNN_CONTEXT_BINARY float Snapdragon® 8 Gen 1 Mobile 14.509 ms 60 - 75 MB NPU
decoder QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ IQ-8275 15.0 ms 60 - 130 MB NPU
decoder QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ QCS8550 (Proxy) 12.16 ms 39 - 40 MB NPU
decoder QNN_CONTEXT_BINARY float Qualcomm® SA8775P 13.816 ms 50 - 61 MB NPU
decoder QNN_CONTEXT_BINARY float Qualcomm® SA8650P 13.816 ms 50 - 61 MB NPU
decoder QNN_CONTEXT_BINARY float Qualcomm® SA8255P 13.816 ms 50 - 61 MB NPU
decoder QNN_CONTEXT_BINARY float Qualcomm® QCS8450 14.509 ms 60 - 75 MB NPU
decoder QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ IQ-9075 13.544 ms 60 - 129 MB NPU
decoder QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ IQ-X7181 11.241 ms 60 - 60 MB NPU
decoder QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ Q-8750 8.331 ms 0 - 8 MB NPU
decoder QNN_CONTEXT_BINARY float Qualcomm® SA7255P 19.127 ms 49 - 58 MB NPU
decoder QNN_CONTEXT_BINARY float Qualcomm® SA8295P 14.92 ms 60 - 66 MB NPU
decoder VOICE_AI float Snapdragon® 8 Elite Gen 5 For Galaxy Mobile 7.334 ms 1 - 9 MB NPU
decoder VOICE_AI float Snapdragon® 8 Elite For Galaxy Mobile 8.278 ms 0 - 8 MB NPU
decoder VOICE_AI float Snapdragon® X2 Elite 6.754 ms 60 - 60 MB NPU
decoder VOICE_AI float Snapdragon® X Elite 11.295 ms 60 - 60 MB NPU
decoder VOICE_AI float Snapdragon® 8 Gen 3 Mobile 9.881 ms 60 - 68 MB NPU
decoder VOICE_AI float Snapdragon® 8 Gen 1 Mobile 13.755 ms 60 - 74 MB NPU
decoder VOICE_AI float Qualcomm® Dragonwing™ IQ-8275 14.603 ms 60 - 130 MB NPU
decoder VOICE_AI float Qualcomm® Dragonwing™ QCS8550 (Proxy) 12.23 ms 39 - 40 MB NPU
decoder VOICE_AI float Qualcomm® SA8775P 13.775 ms 32 - 42 MB NPU
decoder VOICE_AI float Qualcomm® SA8650P 13.775 ms 32 - 42 MB NPU
decoder VOICE_AI float Qualcomm® SA8255P 13.775 ms 32 - 42 MB NPU
decoder VOICE_AI float Qualcomm® QCS8450 13.755 ms 60 - 74 MB NPU
decoder VOICE_AI float Qualcomm® Dragonwing™ IQ-9075 13.469 ms 60 - 129 MB NPU
decoder VOICE_AI float Qualcomm® Dragonwing™ IQ-X7181 11.295 ms 60 - 60 MB NPU
decoder VOICE_AI float Qualcomm® Dragonwing™ Q-8750 8.278 ms 0 - 8 MB NPU
decoder VOICE_AI float Qualcomm® SA8295P 14.987 ms 58 - 63 MB NPU
encoder PRECOMPILED_QNN_ONNX float Snapdragon® 8 Elite Gen 5 For Galaxy Mobile 50.626 ms 125 - 135 MB NPU
encoder PRECOMPILED_QNN_ONNX float Snapdragon® 8 Elite For Galaxy Mobile 66.052 ms 131 - 137 MB NPU
encoder PRECOMPILED_QNN_ONNX float Snapdragon® X2 Elite 52.93 ms 132 - 132 MB NPU
encoder PRECOMPILED_QNN_ONNX float Snapdragon® X Elite 118.025 ms 254 - 254 MB NPU
encoder PRECOMPILED_QNN_ONNX float Snapdragon® 8 Gen 3 Mobile 87.214 ms 128 - 140 MB NPU
encoder PRECOMPILED_QNN_ONNX float Snapdragon® 8 Gen 1 Mobile 179.204 ms 126 - 140 MB NPU
encoder PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ IQ-8275 147.243 ms 125 - 129 MB NPU
encoder PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ QCS8550 (Proxy) 116.489 ms 1 - 258 MB NPU
encoder PRECOMPILED_QNN_ONNX float Qualcomm® QCS8450 179.204 ms 126 - 140 MB NPU
encoder PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ IQ-9075 141.352 ms 129 - 133 MB NPU
encoder PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ IQ-X7181 118.025 ms 254 - 254 MB NPU
encoder PRECOMPILED_QNN_ONNX float Qualcomm® Dragonwing™ Q-8750 66.052 ms 131 - 137 MB NPU
encoder QNN_CONTEXT_BINARY float Snapdragon® 8 Elite Gen 5 For Galaxy Mobile 50.921 ms 1 - 9 MB NPU
encoder QNN_CONTEXT_BINARY float Snapdragon® 8 Elite For Galaxy Mobile 65.435 ms 1 - 9 MB NPU
encoder QNN_CONTEXT_BINARY float Snapdragon® X2 Elite 53.141 ms 0 - 0 MB NPU
encoder QNN_CONTEXT_BINARY float Snapdragon® X Elite 117.168 ms 0 - 0 MB NPU
encoder QNN_CONTEXT_BINARY float Snapdragon® 8 Gen 3 Mobile 86.068 ms 1 - 9 MB NPU
encoder QNN_CONTEXT_BINARY float Snapdragon® 8 Gen 1 Mobile 176.281 ms 1 - 10 MB NPU
encoder QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ IQ-8275 146.985 ms 0 - 57 MB NPU
encoder QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ QCS8550 (Proxy) 114.916 ms 0 - 4 MB NPU
encoder QNN_CONTEXT_BINARY float Qualcomm® SA8775P 139.846 ms 0 - 9 MB NPU
encoder QNN_CONTEXT_BINARY float Qualcomm® SA8650P 139.846 ms 0 - 9 MB NPU
encoder QNN_CONTEXT_BINARY float Qualcomm® SA8255P 139.846 ms 0 - 9 MB NPU
encoder QNN_CONTEXT_BINARY float Qualcomm® QCS8450 176.281 ms 1 - 10 MB NPU
encoder QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ IQ-9075 139.37 ms 0 - 56 MB NPU
encoder QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ IQ-X7181 117.168 ms 0 - 0 MB NPU
encoder QNN_CONTEXT_BINARY float Qualcomm® Dragonwing™ Q-8750 65.435 ms 1 - 9 MB NPU
encoder QNN_CONTEXT_BINARY float Qualcomm® SA7255P 404.454 ms 1 - 10 MB NPU
encoder QNN_CONTEXT_BINARY float Qualcomm® SA8295P 175.209 ms 1 - 6 MB NPU
encoder VOICE_AI float Snapdragon® 8 Elite Gen 5 For Galaxy Mobile 50.825 ms 1 - 10 MB NPU
encoder VOICE_AI float Snapdragon® 8 Elite For Galaxy Mobile 65.624 ms 1 - 9 MB NPU
encoder VOICE_AI float Snapdragon® X2 Elite 53.155 ms 0 - 0 MB NPU
encoder VOICE_AI float Snapdragon® X Elite 117.886 ms 0 - 0 MB NPU
encoder VOICE_AI float Snapdragon® 8 Gen 3 Mobile 86.088 ms 1 - 9 MB NPU
encoder VOICE_AI float Snapdragon® 8 Gen 1 Mobile 173.015 ms 0 - 10 MB NPU
encoder VOICE_AI float Qualcomm® Dragonwing™ IQ-8275 146.9 ms 0 - 57 MB NPU
encoder VOICE_AI float Qualcomm® Dragonwing™ QCS8550 (Proxy) 115.999 ms 0 - 4 MB NPU
encoder VOICE_AI float Qualcomm® SA8775P 140.075 ms 1 - 11 MB NPU
encoder VOICE_AI float Qualcomm® SA8650P 140.075 ms 1 - 11 MB NPU
encoder VOICE_AI float Qualcomm® SA8255P 140.075 ms 1 - 11 MB NPU
encoder VOICE_AI float Qualcomm® QCS8450 173.015 ms 0 - 10 MB NPU
encoder VOICE_AI float Qualcomm® Dragonwing™ IQ-9075 140.046 ms 0 - 56 MB NPU
encoder VOICE_AI float Qualcomm® Dragonwing™ IQ-X7181 117.886 ms 0 - 0 MB NPU
encoder VOICE_AI float Qualcomm® Dragonwing™ Q-8750 65.624 ms 1 - 9 MB NPU
encoder VOICE_AI float Qualcomm® SA8295P 175.253 ms 0 - 5 MB NPU

License

  • The license for the original implementation of Whisper-Small can be found here.

References

Community

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support