Skip to content

Repository files navigation

vox — Voice Command

vox

Local voice for AI coding agents — speech out and speech in, from a single static binary.

No Python, no API key, no cloud. Five TTS backends, local Whisper speech-to-text, and an MCP server that plugs into 14 AI tools.

CI Release License

English • Français • 中文 • 日本語 • 한국어 • Español


                          vox
                           |
         +-----------------+-----------------+
         |                                   |
     speak (TTS)                         hear (STT)
         |                                   |
  +---+---+---+------+-------+------+    Whisper
  |   |   |   |      |       |      |   (Rust/candle)
 say piper pocket qwen-native kokoro        99 languages
      |                                          CPU / Metal / CUDA
      |                                              |
      +----------------- rodio ----------------------+
              (playback)        cpal (capture)

Backends

Backend Engine Voice cloning Latency (warm) GPU Platform
say macOS native No 3s No macOS
piper ONNX (Rust) No <1s No All
pocket Candle (Rust, 100M) Yes* ~2s No (CPU-first) All
qwen-native Candle (Rust) Yes ~3s Metal/CUDA All
kokoro ONNX (Rust, opt-in) No <1s No macOS only

* pocket ships 8 predefined voices with zero setup (public weights). Voice cloning from a reference WAV needs HF_TOKEN and the gated kyutai/pocket-tts license accepted. The bundled checkpoint is English-only (all 8 voices are English speakers); Kyutai's per-language checkpoints (fr/de/es/it/pt) need upstream support in the pocket-tts crate and are not wired up yet.

Benchmark — single sentence (~50 chars)

All times measured end-to-end (model loading + inference + audio playback). Cold = first CLI call.

Backend M2 Pro (CPU) RTX 4070 Ti SUPER Voice cloning Quality
say 3s macOS only No System voices
piper <1s <1s No Good
pocket (Kyutai, 100M) default 7–14s cold¹ same (CPU-only model) Yes Very good (EN only) — generation faster than real-time
kokoro <1s macOS only No Fair (EN only)
qwen-native (Qwen3-TTS, 0.6B) 11m33s / 3s warm 48s (CPU) Yes Excellent

With daemon (vox daemon start — keeps model server warm):

Backend M2 Pro (CPU) Notes
qwen-native ~3s Model stays in RAM via global Mutex

¹ pocket end-to-end cold on an i7-1065G7 laptop CPU: model load + generation + full playback of 5–9s of audio. Generation alone runs faster than real-time. pocket is CPU-only by design, so a GPU does not change the figure. First run also downloads ~226 MB once. All CUDA benchmarks measured on RTX 4070 Ti SUPER (16GB). For lowest latency: say (macOS) or piper (all platforms). For best quality + cloning: qwen-native with the daemon.

Install

Pre-built binaries (recommended)

# Quick install (macOS ARM / Linux x86_64 & ARM64)
curl -fsSL https://raw.githubusercontent.com/rtk-ai/vox/main/install.sh | sh

# Custom install dir (no sudo needed)
curl -fsSL https://raw.githubusercontent.com/rtk-ai/vox/main/install.sh | VOX_INSTALL_DIR=~/.local/bin sh

# Homebrew (macOS)
brew install rtk-ai/tap/vox

The installer defaults to /usr/local/bin (with sudo if needed). When sudo is not usable (CI, agents, no TTY), it falls back to ~/.local/bin automatically.

Pre-built binaries are available for each release:

Platform Binary GPU
macOS (Apple Silicon) vox-aarch64-apple-darwin.tar.gz Metal
Linux x86_64 vox-x86_64-unknown-linux-gnu.tar.gz CPU
Linux x86_64 + CUDA vox-x86_64-unknown-linux-gnu-cuda.tar.gz CUDA
Linux ARM64 vox-aarch64-unknown-linux-gnu.tar.gz CPU
Linux (Debian/Ubuntu) vox-{x86_64,aarch64}-unknown-linux-gnu.deb CPU
Linux (Fedora/RHEL) vox-{x86_64,aarch64}-unknown-linux-gnu.rpm CPU
Windows x86_64 vox-x86_64-pc-windows-msvc.zip CPU
Windows x86_64 + CUDA vox-x86_64-pc-windows-msvc-cuda.zip CUDA

Download from GitHub Releases.

From source

cargo install --path .                   # CPU only
cargo install --path . --features metal  # macOS Apple Silicon (GPU)
cargo install --path . --features cuda   # Linux/Windows NVIDIA (GPU)

Linux requires sudo apt install libasound2-dev.

Platform defaults

Platform Default backend Notes
All (English or no -l) pocket Public weights auto-download on first use (~226 MB)
All (other languages) piper The pocket checkpoint is English-only; piper has per-language voices

Quick start

vox "Hello, world."                     # Speak with default backend
vox -b qwen-native "Neural TTS."        # Qwen3 (best quality)
vox -b piper "Fast TTS."                # Piper (fastest)
vox --volume 2.0 "Louder!"             # 2x volume (range: 0.0-5.0)
vox -l fr "Bonjour"                     # French
echo "Piped text" | vox                 # Read from stdin
vox --list-voices                       # List available voices
vox setup                               # Interactive TUI configuration

Interactive setup (TUI)

For humans — choose backend, voice, language, style, and volume interactively:

vox setup
┌ Backend ──┐┌ Voice ─────┐┌ Lang ┐┌ Style ────┐┌ Volume ┐┌ Config ──────┐
│> say      ││> Samantha  ││> en  ││> (default)││  0.5   ││ Backend: say │
│  piper    ││  Thomas    ││  fr  ││  calm     ││> 1.0   ││ Voice: ...   │
│  qwen-nat ││  Amelie    ││  es  ││  warm     ││  1.5   ││ Lang:  en    │
│  qwen     ││           ││  ja  ││          ││  3.0   ││ [T]est [S]ave│
└───────────┘└────────────┘└──────┘└──────────┘└────────┘└──────────────┘

Navigate with arrow keys / hjkl, Tab to switch panel, T to test, S to save, Q to quit.

AI agents use CLI flags instead: vox -b qwen-native -l fr "text"

AI assistant integration

One command configures 14 AI tools (Claude Code, Cursor, VS Code, Zed, Codex, Gemini, Amazon Q, and more):

vox init                # MCP server (default) — all AI tools
vox init -m cli         # CLAUDE.md + Stop hook (recommended)
vox init -m skill       # /speak slash command
vox init -m all         # all of the above

Running vox init again is safe — it skips files that are already configured.

The generated instructions and Stop hook speak your language. By default vox uses your vox config set lang preference, falling back to your system locale; with neither, it tells the agent to match whatever language you write in. Override it explicitly:

vox init -m cli --lang de   # agent summaries in German, hook says "Fertig."
vox init -m cli --lang ja   # Japanese

CLI mode vs MCP mode

CLI mode is recommended for AI coding agents. Benchmarks show CLI tools are 10-32x cheaper and 100% reliable vs 72% for MCP due to MCP's TCP timeout overhead and JSON schema cost per call.

Mode Reliability Token cost Best for
CLI (vox init -m cli) 100% Low (Bash call) Claude Code, Codex, terminal agents
MCP (vox init) ~72% Higher (JSON schema) Cursor, VS Code, GUI-based tools

Voice cloning

vox clone add patrick --audio ~/voice.wav --text "Transcription"
vox clone record myvoice --duration 10
vox -v patrick "This speaks with your voice."
vox clone list
vox clone remove patrick

Works with the qwen-native and pocket backends, both pure Rust. A 3-second reference clip is enough.

Preferences

vox config show
vox config set backend qwen-native
vox config set lang fr
vox config set voice Chelsie
vox config set gender feminine
vox config set style warm
vox config reset

Sound packs

vox pack install peon              # Install a pack
vox pack set peon                  # Activate it
vox pack play greeting             # Play a sound
vox pack list                      # List available packs

Speech-to-text (all platforms)

Local Whisper on candle (pure Rust, 99 languages). The model is downloaded from Hugging Face on first use and kept warm by the daemon / MCP server.

vox hear                                   # Listen, auto-detect language, print text
vox hear -l fr -t 60 -s 3.0                # French, max 60s, stop after 3s of silence
vox hear -m openai/whisper-large-v3-turbo  # Best quality (GPU + 16 GB RAM recommended)
vox hear -f recording.wav                  # Transcribe a file instead of the mic
Env var Description
VOX_STT_MODEL Whisper repo (default openai/whisper-small, ~1 GB RAM; openai/whisper-large-v3-turbo ~3.5 GB)
VOX_VAD_THRESHOLD Minimum RMS speech threshold, 0-1 (default 0.0125; adapts to ambient noise)
VOX_VAD_DEBUG Set to 1 to print RMS levels and the chosen threshold

No external tools needed: microphone capture uses cpal, so sox is no longer required.

Voice conversation (macOS)

export ANTHROPIC_API_KEY=sk-...
vox chat -l fr                     # Talk with Claude

Data

All state is stored locally — no data sent to external servers (except vox chat which uses Claude API).

~/.config/vox/           # or ~/Library/Application Support/vox/ on macOS
  vox.db                 # SQLite: preferences, voice clones, usage logs
  clones/                # Audio files for voice clones
  packs/                 # Installed sound packs
Env var Description
VOX_CONFIG_DIR Override config directory
VOX_DB_PATH Override database path

Documentation

Document Description
Architecture Technical architecture, backends, DB schema, MCP protocol, security
Features All commands and features documented
Guide Installation, quick start, troubleshooting

License

Apache-2.0

About

A universal AI toolkit for high-performance Speech-to-Text (STT) and Text-to-Speech (TTS) processing, designed for low-latency and easy model integration.

Topics

Resources

Stars

161 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages