Local voice for AI coding agents — speech out and speech in, from a single static binary.
No Python, no API key, no cloud. Five TTS backends, local Whisper speech-to-text, and an MCP server that plugs into 14 AI tools.
English • Français • 中文 • 日本語 • 한국어 • Español
vox
|
+-----------------+-----------------+
| |
speak (TTS) hear (STT)
| |
+---+---+---+------+-------+------+ Whisper
| | | | | | | (Rust/candle)
say piper pocket qwen-native kokoro 99 languages
| CPU / Metal / CUDA
| |
+----------------- rodio ----------------------+
(playback) cpal (capture)
| Backend | Engine | Voice cloning | Latency (warm) | GPU | Platform |
|---|---|---|---|---|---|
say |
macOS native | No | 3s | No | macOS |
piper |
ONNX (Rust) | No | <1s | No | All |
pocket |
Candle (Rust, 100M) | Yes* | ~2s | No (CPU-first) | All |
qwen-native |
Candle (Rust) | Yes | ~3s | Metal/CUDA | All |
kokoro |
ONNX (Rust, opt-in) | No | <1s | No | macOS only |
*
HF_TOKENand the gated kyutai/pocket-tts license accepted. The bundled checkpoint is English-only (all 8 voices are English speakers); Kyutai's per-language checkpoints (fr/de/es/it/pt) need upstream support in the pocket-tts crate and are not wired up yet.
All times measured end-to-end (model loading + inference + audio playback). Cold = first CLI call.
| Backend | M2 Pro (CPU) | RTX 4070 Ti SUPER | Voice cloning | Quality |
|---|---|---|---|---|
say |
3s | macOS only | No | System voices |
piper |
<1s | <1s | No | Good |
pocket (Kyutai, 100M) default |
7–14s cold¹ | same (CPU-only model) | Yes | Very good (EN only) — generation faster than real-time |
kokoro |
<1s | macOS only | No | Fair (EN only) |
qwen-native (Qwen3-TTS, 0.6B) |
11m33s / 3s warm | 48s (CPU) | Yes | Excellent |
With daemon (vox daemon start — keeps model server warm):
| Backend | M2 Pro (CPU) | Notes |
|---|---|---|
qwen-native |
~3s | Model stays in RAM via global Mutex |
¹
say(macOS) orpiper(all platforms). For best quality + cloning:qwen-nativewith the daemon.
# Quick install (macOS ARM / Linux x86_64 & ARM64)
curl -fsSL https://raw.githubusercontent.com/rtk-ai/vox/main/install.sh | sh
# Custom install dir (no sudo needed)
curl -fsSL https://raw.githubusercontent.com/rtk-ai/vox/main/install.sh | VOX_INSTALL_DIR=~/.local/bin sh
# Homebrew (macOS)
brew install rtk-ai/tap/voxThe installer defaults to /usr/local/bin (with sudo if needed). When sudo is
not usable (CI, agents, no TTY), it falls back to ~/.local/bin automatically.
Pre-built binaries are available for each release:
| Platform | Binary | GPU |
|---|---|---|
| macOS (Apple Silicon) | vox-aarch64-apple-darwin.tar.gz |
Metal |
| Linux x86_64 | vox-x86_64-unknown-linux-gnu.tar.gz |
CPU |
| Linux x86_64 + CUDA | vox-x86_64-unknown-linux-gnu-cuda.tar.gz |
CUDA |
| Linux ARM64 | vox-aarch64-unknown-linux-gnu.tar.gz |
CPU |
| Linux (Debian/Ubuntu) | vox-{x86_64,aarch64}-unknown-linux-gnu.deb |
CPU |
| Linux (Fedora/RHEL) | vox-{x86_64,aarch64}-unknown-linux-gnu.rpm |
CPU |
| Windows x86_64 | vox-x86_64-pc-windows-msvc.zip |
CPU |
| Windows x86_64 + CUDA | vox-x86_64-pc-windows-msvc-cuda.zip |
CUDA |
Download from GitHub Releases.
cargo install --path . # CPU only
cargo install --path . --features metal # macOS Apple Silicon (GPU)
cargo install --path . --features cuda # Linux/Windows NVIDIA (GPU)Linux requires sudo apt install libasound2-dev.
| Platform | Default backend | Notes |
|---|---|---|
All (English or no -l) |
pocket |
Public weights auto-download on first use (~226 MB) |
| All (other languages) | piper |
The pocket checkpoint is English-only; piper has per-language voices |
vox "Hello, world." # Speak with default backend
vox -b qwen-native "Neural TTS." # Qwen3 (best quality)
vox -b piper "Fast TTS." # Piper (fastest)
vox --volume 2.0 "Louder!" # 2x volume (range: 0.0-5.0)
vox -l fr "Bonjour" # French
echo "Piped text" | vox # Read from stdin
vox --list-voices # List available voices
vox setup # Interactive TUI configurationFor humans — choose backend, voice, language, style, and volume interactively:
vox setup┌ Backend ──┐┌ Voice ─────┐┌ Lang ┐┌ Style ────┐┌ Volume ┐┌ Config ──────┐
│> say ││> Samantha ││> en ││> (default)││ 0.5 ││ Backend: say │
│ piper ││ Thomas ││ fr ││ calm ││> 1.0 ││ Voice: ... │
│ qwen-nat ││ Amelie ││ es ││ warm ││ 1.5 ││ Lang: en │
│ qwen ││ ││ ja ││ ││ 3.0 ││ [T]est [S]ave│
└───────────┘└────────────┘└──────┘└──────────┘└────────┘└──────────────┘
Navigate with arrow keys / hjkl, Tab to switch panel, T to test, S to save, Q to quit.
AI agents use CLI flags instead: vox -b qwen-native -l fr "text"
One command configures 14 AI tools (Claude Code, Cursor, VS Code, Zed, Codex, Gemini, Amazon Q, and more):
vox init # MCP server (default) — all AI tools
vox init -m cli # CLAUDE.md + Stop hook (recommended)
vox init -m skill # /speak slash command
vox init -m all # all of the aboveRunning vox init again is safe — it skips files that are already configured.
The generated instructions and Stop hook speak your language. By default vox
uses your vox config set lang preference, falling back to your system locale;
with neither, it tells the agent to match whatever language you write in.
Override it explicitly:
vox init -m cli --lang de # agent summaries in German, hook says "Fertig."
vox init -m cli --lang ja # JapaneseCLI mode is recommended for AI coding agents. Benchmarks show CLI tools are 10-32x cheaper and 100% reliable vs 72% for MCP due to MCP's TCP timeout overhead and JSON schema cost per call.
| Mode | Reliability | Token cost | Best for |
|---|---|---|---|
CLI (vox init -m cli) |
100% | Low (Bash call) | Claude Code, Codex, terminal agents |
MCP (vox init) |
~72% | Higher (JSON schema) | Cursor, VS Code, GUI-based tools |
vox clone add patrick --audio ~/voice.wav --text "Transcription"
vox clone record myvoice --duration 10
vox -v patrick "This speaks with your voice."
vox clone list
vox clone remove patrickWorks with the qwen-native and pocket backends, both pure Rust. A 3-second reference clip is enough.
vox config show
vox config set backend qwen-native
vox config set lang fr
vox config set voice Chelsie
vox config set gender feminine
vox config set style warm
vox config resetvox pack install peon # Install a pack
vox pack set peon # Activate it
vox pack play greeting # Play a sound
vox pack list # List available packsLocal Whisper on candle (pure Rust, 99 languages). The model is downloaded from Hugging Face on first use and kept warm by the daemon / MCP server.
vox hear # Listen, auto-detect language, print text
vox hear -l fr -t 60 -s 3.0 # French, max 60s, stop after 3s of silence
vox hear -m openai/whisper-large-v3-turbo # Best quality (GPU + 16 GB RAM recommended)
vox hear -f recording.wav # Transcribe a file instead of the mic| Env var | Description |
|---|---|
VOX_STT_MODEL |
Whisper repo (default openai/whisper-small, ~1 GB RAM; openai/whisper-large-v3-turbo ~3.5 GB) |
VOX_VAD_THRESHOLD |
Minimum RMS speech threshold, 0-1 (default 0.0125; adapts to ambient noise) |
VOX_VAD_DEBUG |
Set to 1 to print RMS levels and the chosen threshold |
No external tools needed: microphone capture uses cpal, so sox is no longer required.
export ANTHROPIC_API_KEY=sk-...
vox chat -l fr # Talk with ClaudeAll state is stored locally — no data sent to external servers (except vox chat which uses Claude API).
~/.config/vox/ # or ~/Library/Application Support/vox/ on macOS
vox.db # SQLite: preferences, voice clones, usage logs
clones/ # Audio files for voice clones
packs/ # Installed sound packs
| Env var | Description |
|---|---|
VOX_CONFIG_DIR |
Override config directory |
VOX_DB_PATH |
Override database path |
| Document | Description |
|---|---|
| Architecture | Technical architecture, backends, DB schema, MCP protocol, security |
| Features | All commands and features documented |
| Guide | Installation, quick start, troubleshooting |
