Whistle ASR: 16.9 MB On-Device Speech Model

Introduction
Cactus Compute just open-sourced Whistle—a compact multilingual automatic speech recognition model plus a CPU inference engine called Needle—aimed at phones, wearables, robots, and MCU-class targets. Primary sources: the Oct 2, 2026 blog Whistle (Jakub Mroz, Henry Ndubuaku et al.) and the Hugging Face card Cactus-Compute/whistle.
Open weights ship as a single 16.9 MB whistle.cact file. You can run it today with pip install cactus-needle and a short Python call—no cloud ASR round-trip required for the default path.
Pro Tip
Treat Whistle as an on-device first-pass ASR, not a cloud Whisper replacement for long-form podcasts. Cap clips near the 30-second one-pass window, lean on keyword biasing for wake words and product names, and keep the tiny whistle.cact beside your Needle binary so mobiles and robots stay offline-capable.
What Shipped
Concrete facts from Cactus’s Oct 2 blog + HF card:
FactDetail
Model
Whistle ASR (open weights)
Vendor / authors
Cactus Compute — Jakub Mroz, Henry Ndubuaku et al.
Weights file
Single 16.9 MB whistle.cact
Engine
Needle (CPU) via cactus-needle
Languages
7: EN, DE, FR, ES, IT, NL, PL
Clip window
Up to ~30 s in one pass
Extras
Word timestamps, speech embeddings, keyword biasing
Claimed M4 Pro TTFT
~11.1 ms vs Whisper base ~73.2 ms
Claimed decode
~1319 tok/s
Size context
Whistle 16.9 MB vs Whisper base 145.3 MB / Moonshine tiny 41.9 MB
Those latency and throughput numbers are vendor-reported on Apple M4 Pro—useful as a sizing signal, not a cross-device SLA. Still, the headline is clear: a Whisper-base-class footprint that is roughly 9× smaller than Whisper base and well under Moonshine tiny, with a CPU engine packaged for real devices.
Why On-Device Builders Care
Cloud ASR is fine until you hit offline robots, wearables with spotty radios, or privacy-sensitive capture. Whistle’s pitch is a single small artifact plus Needle binaries that target mobiles, wearables, robots, and MCU paths—not only laptop demos.
Word timestamps matter for subtitle UIs and tool routing. Speech embeddings open “same speaker / similar clip” shortcuts without a second model. Keyword biasing is the practical win for wake phrases, brand names, and constrained vocab in noisy rooms.
How to Run It
Install Needle, drop a short WAV, and transcribe:
pip install cactus-needleimport needle
result = needle.transcribe("clip.wav")
print(result)Needle also exposes tooling so speech can drive tool calls—useful when ASR is the front door to an on-device agent rather than a transcript dump. Keep clips inside the ~30 s one-pass window for the happy path; longer audio needs your own chunking strategy until Cactus documents a streaming recipe.
For device packaging, start from the HF card and the blog for Needle binaries aimed at mobiles / wearables / robots / MCU targets.
Limits to Keep Honest
- Seven languages only (EN/DE/FR/ES/IT/NL/PL)—not a global Whisper multilingual substitute
- ~30 s one-pass window; plan chunking for long meetings
- Latency figures are M4 Pro vendor claims—re-bench on your silicon
- On-device MCU paths depend on Needle packaging maturity for your board
Conclusion
Whistle is the rare same-day ASR drop that leads with a shippable file size: 16.9 MB, seven European languages, word timestamps, embeddings, keyword biasing, and a CPU Needle path you can pip install this afternoon. Start from the Cactus blog and Cactus-Compute/whistle, then A/B against Whisper base / Moonshine tiny on your device before you freeze a wake-word or robot pipeline.
—Anabel ♡
