# Whistle (Cactus Compute)

> Source: https://aiwiki.ai/wiki/whistle_cactus_compute
> Updated: 2026-10-10
> Fact-checked: 2026-10-10
> Categories: AI Models, Speech & Audio AI
> License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - attribute to "AI Wiki (aiwiki.ai)"
> Cite as: AI Wiki. "Whistle (Cactus Compute)." aiwiki.ai, 10 Oct 2026. https://aiwiki.ai/wiki/whistle_cactus_compute
> From AI Wiki (https://aiwiki.ai), the free encyclopedia of artificial intelligence. Reuse freely with attribution.

Whistle is a [speech recognition](https://aiwiki.ai/wiki/speech_recognition) model from Cactus Compute, released on October 2, 2026. Its 16.9 MB model file runs on a CPU through the C++ engine also used by Needle, Cactus Compute's tool-calling model.[1]

## Capabilities

The model weights are distributed under Apache-2.0. Whistle processes speech locally and provides the following outputs.[2]

| Feature | Specification |
| --- | --- |
| Input | 16 kHz mono audio; up to 30 seconds per pass |
| Languages | English, German, French, Spanish, Italian, Dutch and Polish |
| Transcription | Text with detected or specified language |
| Word alignment | Start time, end time and probability per word |
| Speech embeddings | Encoder output at 80 ms intervals |

## Architecture

Whistle converts audio into 80 log-mel channels. Three convolutional reductions leave one encoder frame per 80 ms. An eight-block encoder processes the clip non-causally; an eight-layer decoder reads it through gated cross-attention. The decoder uses five-beam search and an 8,192-piece text vocabulary with seven language tokens. Decoding stops at a 320-token cap.[1]

The decoder can run at depths from two to eight layers through `--audio-depth`. Every selection retains the complete encoder. Keyword biasing uses an Aho-Corasick automaton during decoding.[1]

## Streaming and keyword biasing

`needle.stream(chunks)` repeatedly transcribes buffered audio. It commits words when two consecutive passes agree, leaving an unconfirmed tail. Each update reports committed text, pending text and the pass duration. Word times use the start of the stream as their reference. A single non-streaming call refuses audio longer than 30 seconds.[3]

Keyword lists increase the probability of supplied names and phrases. Cactus reports the following contextual-biasing experiment on 2,611 LibriSpeech test-clean utterances, using per-utterance lists and scoring from Le and colleagues' 2021 work.[3]

| Keyword list | B-WER (%) | U-WER (%) | Overall WER (%) |
| --- | --- | --- | --- |
| None | 18.43 | 3.05 | 4.73 |
| 100 entries | 4.46 | 2.74 | 2.93 |
| 1,000 entries | 5.04 | 2.84 | 3.08 |

B-WER measures errors on words in the biasing list; U-WER covers other words. The original study constructs artificial lists containing rare reference words and distractors. These lists therefore provide information about the reference transcript that an arbitrary application might not possess.[4]

## Python interface

The package installs with `pip install cactus-needle`. A minimal transcription call is:[5]

```python
import needle

result = needle.transcribe("clip.wav")
print(result["text"])
```

`needle.transcribe` accepts a WAV path or 16 kHz mono floating-point samples. Its result contains `text`, `language`, `ttft_ms` and `decode_tps`. Optional arguments select a language, supply keywords or enable word timestamps. Language codes are `en`, `de`, `fr`, `es`, `it`, `nl` and `pl`.[5]

Microphone capture and resampling require the `[mic]` extra. The weights download separately from the engine on first use; `NEEDLE_WHISTLE_WEIGHTS` selects a local copy. `needle.Whistle()` provides an object interface for speech embeddings. The interface permits one loaded model per process and is not thread-safe.[5]

## Deployment and Needle integration

Native deployment folders provide a command-line runner, static library and C header. Supported targets include macOS, Windows, Linux, Android, iOS and watchOS, with ARM, x86, RISC-V and MIPS variants where listed. Browser deployment uses JavaScript and WebAssembly; WASI deployment uses a component and its interface definition.[6]

Whistle and Needle can share a runtime. Loading both `whistle.cact` and `needle3.cact` allows the engine to transcribe audio, submit the text against supplied tool schemas and return one JSON result containing calls and speech fields. The speech fields have an `audio_` prefix. This combines two models: Whistle supplies the transcript and Needle selects the tool calls.[6][7]

The Python package also permits an explicit sequence: transcribe with Whistle, then pass the text to a Needle object. Tool names and enumerated argument values can be supplied as transcription keywords. A transcription error can prevent Needle from matching the intended argument.[3]

Local audio processing does not imply that every package operation is network-free. The repository documents enabled-by-default telemetry and the opt-out settings `NEEDLE_TELEMETRY=0` and `DO_NOT_TRACK=1`. It also provides terminal transcription and comparison commands.[8]

## Benchmarks

Cactus reports 11.1 ms to the first token and 1,319 decoded tokens per second for ten seconds of audio on an Apple M4 Pro CPU. The decoder rate measures time after the first token, rather than complete transcription latency.[1]

The following [word error rates](https://aiwiki.ai/wiki/word_error_rate) are developer-reported results from Whistle's model card and chart.[2]

| Test | Whistle WER (%) |
| --- | --- |
| LibriSpeech test-clean | 4.31 |
| LibriSpeech test-other | 10.49 |
| SPGISpeech | 7.65 |
| Earnings-22 | 19.01 |
| AMI | 26.07 |
| AMI cleaned | 22.87 |
| TED-LIUM | 7.61 |
| FLEURS, seven-language average | 21.4 |
| MLS, six languages excluding English | 24.9 |

Cactus evaluated 86,174 utterances with Whisper normalizers. Comparisons with [Whisper](https://aiwiki.ai/wiki/whisper) and Moonshine combine published accuracy figures with runtime timing tests. Precision differs: Whistle uses 2-4 bits, Whisper uses CPU fp32 and Moonshine uses int8. Whisper's AMI figure uses AMI-IHM, a different subset. These conditions limit direct comparisons.[2]

## References

1. Jakub Mroz and Henry Ndubuaku. [Whistle: Speech to Text in 16.9 MB](https://cactuscompute.com/blog/whistle). Cactus Compute, October 2, 2026.
2. Cactus Compute. [Whistle model card](https://huggingface.co/Cactus-Compute/whistle), including its benchmark chart. Accessed October 10, 2026.
3. Jakub Mroz. [Getting the Most out of Whistle](https://cactuscompute.com/blog/whistle-best-practices). Cactus Compute, October 3, 2026.
4. Duc Le and colleagues. [Contextualized Streaming End-to-End Speech Recognition with Trie-Based Deep Biasing and Shallow Fusion](https://arxiv.org/abs/2104.02194). INTERSPEECH 2021; arXiv:2104.02194v2.
5. Roman Shemet. [Needle Python Docs](https://cactuscompute.com/blog/needle-python-docs). Cactus Compute, September 18, 2026.
6. Justin H. Lee. [What Devices Are Supported on Needle](https://cactuscompute.com/blog/needle-supported-devices). Cactus Compute, September 18, 2026.
7. Cactus Compute. [Needle 3 model card](https://huggingface.co/Cactus-Compute/needle3). Accessed October 10, 2026.
8. Cactus Compute. [Needle repository README](https://github.com/cactus-compute/needle). Accessed October 10, 2026.
