Skip to content
Rush Commerce
AI & Automation3 min read

Cactus Whistle: 16.9 MB on-device speech-to-text, Apache 2.0

Cactus Whistle is a 16.9 MB on-device speech-to-text model that beats Whisper Base on several benchmarks and runs on CPU. Here is where it fits.

On-device speech-to-text just got small enough to stop being a decision. Cactus Compute released Whistle on October 2: a 16.9 MB transcription model that runs on a plain CPU, ships under Apache 2.0, and beats OpenAI's Whisper Base on most of the benchmarks Cactus published. Whisper Base is 145.3 MB. Whistle is about one-ninth of that.

What actually happened

Per Cactus's launch post, Whistle is one file with no dependencies. It transcribes clips up to 30 seconds of 16 kHz mono audio in seven languages (English, German, French, Spanish, Italian, Dutch and Polish) and detects the language itself. It returns word-level timestamps with confidence scores, and it takes a keyword list so it favors your product names and jargon.

On accuracy, Cactus says Whistle is ahead of Whisper Base on LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22 and the FLEURS average. Whisper Base stays ahead on TED-LIUM, AMI meeting audio and the MLS average. So it is not a clean sweep, and meeting recordings are the weak spot.

On speed, Cactus measured 11.1 ms to first token against 73.2 ms for Whisper Base, and 1,319 tokens per second against 266, on an Apple M4 Pro CPU. Those are vendor numbers on vendor hardware. Test on your own devices.

Whistle runs on the same engine as Cactus's Needle project (Apache 2.0, 13.1k GitHub stars), with prebuilt targets for macOS, Linux, Windows on ARM, Android, iOS, watchOS, the browser and more. Weights are on Hugging Face. In Python it is pip install cactus-needle and one transcribe() call.

Why it matters for your business

Most small teams pay per minute for transcription because local models were too big or too slow to put in an app. A 17 MB model changes that math for short audio.

Voice notes and field capture. A technician dictates a job note on a phone with no signal. The text is there before they get back to the truck, and no audio leaves the device.

Voice commands in your own app. Keyword biasing means "SKU 4471" and your product names come through correctly. That is the part generic cloud APIs get wrong.

Privacy by default. Audio that never leaves the device is audio you never have to secure, store or disclose.

Know the limits. Thirty seconds per clip means you chunk long calls yourself. Seven languages means no Spanish-and-Mandarin support desk. And if your audio is meetings, check Whisper's lead on AMI before you switch.

Key takeaways

  • Cactus Whistle is a 16.9 MB speech-to-text model, Apache 2.0, CPU only
  • Cactus reports it beats Whisper Base on LibriSpeech, SPGISpeech, Earnings-22 and FLEURS
  • Whisper Base still wins on TED-LIUM, AMI meeting audio and MLS
  • Clips cap at 30 seconds; seven languages with auto-detection
  • Keyword biasing and word timestamps make it usable for commands and field notes

Paying per minute for short voice clips? We build voice capture that runs on the device, tuned to your product names, and wired into the systems you already use. See what we build, or check the numbers with our ROI calculator.

Sources: Cactus Compute: Whistle, Hugging Face: Cactus-Compute/whistle, GitHub: cactus-compute/needle.

  • #speech-to-text
  • #on-device-ai
  • #cactus-whistle
  • #voice-agents
  • #edge-ai
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.