Cactus Compute's Whistle brings speech-to-text to the edge
A compact on-device ASR model built for edge hardware and WebAssembly, with early support for English, German, and French.
Cactus Compute has released Whistle, a speech-to-text model designed to run directly on edge devices and in the browser via WebAssembly. Rather than chasing the largest possible accuracy ceiling, the project aims squarely at deployment where cloud round-trips aren't an option — phones, embedded hardware, and web apps that need transcription without a server.
The initial release covers three languages — English, German, and French — according to its Hugging Face repository. That focused scope fits the on-device brief: a tighter language set keeps the model small enough to load and run locally without specialized accelerators.
Why it matters
On-device automatic speech recognition addresses two concerns that dominate real-world deployments:
- Privacy: audio never leaves the device, which matters for regulated industries and consumer trust.
- Latency and cost: local inference removes network dependence and per-request API fees, making always-on transcription practical.
The WebAssembly angle is notable. Shipping ASR that runs in a standard browser sandbox lowers the barrier for web developers who want transcription features without maintaining backend infrastructure or asking users to install native software.
The release is an initial version, and key details — parameter count, context handling, and benchmark figures — aren't specified in the record. Developers evaluating Whistle for production should test transcription quality against their own audio and language needs, since edge-optimized models typically trade some accuracy for footprint and speed.
Sources
- Visit
Cactus-Compute/whistle
Hugging Face
More in Speech → Text
Phonon-2 brings on-device ASR to Apple Silicon
A low-bit quantized, Parakeet-based speech recognizer built to run locally on Mac hardware.
Audio8-ASR-Infinite brings streaming bilingual speech recognition
A new open model targets real-time transcription for Chinese and English audio.
Moondream shrinks Parakeet ASR for CPUs
A ternary-quantized take on the Parakeet TDT speech model aims to run transcription without a GPU.
0 comments
No comments yet. Be the first to weigh in.