Voice-to-text for your Linux desktop. Press a hotkey, speak, and your words appear wherever the cursor is. Or leave VAD mode on and let it transcribe continuously, hands-free.
Backend-agnostic — works with local whisper.cpp, faster-whisper, or any OpenAI-compatible ASR endpoint (Groq, OpenAI, your own server).
- Interactive TUI — Terminal UI with real-time status, VAD mode, live transcription, and configuration (default mode)
- Push-to-talk — bind
ears toggleto a keyboard shortcut for quick dictation - VAD mode — hands-free voice activity detection with auto-transcription (
ears vad) - Streaming transcription — real-time text output with LocalAgreement for stable progressive output
- Volume ducking — optionally lowers system volume while you're speaking
- Bash mode — constrain dictation to valid shell syntax via grammar-guided decoding (speak commands, get code)
- Profiles — switch between local whisper and cloud APIs (Groq, OpenAI, etc.) per-invocation
- Text filters — optional lowercase conversion and punctuation removal
- Language detection — auto-detects from keyboard layout (Hyprland + GNOME)
- Smart text input — uses
wtypeon Hyprland/Wayland, clipboard paste viaydotoolelsewhere - PipeWire audio — native support for the modern Linux audio stack
- Post-transcribe hooks — run custom scripts after each transcription
- Audio feedback — embedded cue sounds, customizable with custom sound override support
- State management — file-based locking and state with automatic crash recovery
- No telemetry — audio goes only to the server you configure
- Linux with PipeWire audio system
- A whisper.cpp or OpenAI-compatible ASR server running
- Text input tool:
wtype(Hyprland/Wayland) orydotool(other systems)
notify-sendfor desktop notificationspaplayfor audio feedbackfzffor interactive device selectionwl-clipboard(wl-copy) for clipboard-based text input on non-Hyprland systems
# Arch Linux (Hyprland/Omarchy)
sudo pacman -S pipewire wtype libnotify pulseaudio fzf
# Ubuntu/Debian
sudo apt install pipewire ydotool wl-clipboard libnotify-bin pulseaudio-utils fzf
# Fedora
sudo dnf install pipewire ydotool libnotify pulseaudio-utils fzf- Clone and build whisper.cpp:
git clone https://github.com/ggerganov/whisper.cpp
cd whisper.cpp
make server # CPU only
make server WHISPER_CUDA=1 # With NVIDIA GPU- Download a model:
bash ./models/download-ggml-model.sh base.en- Start the server:
./server -m models/ggml-base.en.bin -p 8178mkdir -p ~/.local/bin
gh release download latest --repo heiervang-technologies/ears --pattern 'ears' --dir ~/.local/bin --clobber
chmod +x ~/.local/bin/ears
export PATH="$HOME/.local/bin:$PATH"Building from source requires Rust 1.88 or newer.
git clone https://github.com/heiervang-technologies/ears
cd ears
cargo build --release
cargo install --path .git clone https://github.com/heiervang-technologies/ears
cd ears
./install.shConfiguration is stored in ~/.config/ears/config.toml:
server = "http://127.0.0.1:8178"
device = "alsa_input.usb-..."
# language = "en" # Optional (auto-detects from keyboard layout)
# api_key = "sk-..." # Optional (for authenticated ASR services)
# model = "whisper-large-v3-turbo" # Optional (for cloud APIs that require it)
# prompt = "vLLM, PyTorch" # Optional (context biasing: names, acronyms, jargon)
filter_silence_artifacts = true # Set false to retain exact phrases such as "Thank you"
[text_filters]
lowercase = false
remove_punctuation = false
[vad]
# Increase this for longer mid-sentence pauses; decrease it for faster dispatch.
max_silence_duration_ms = 1200Server URL: ears sends requests to
{server}/v1/audio/transcriptions, appending/v1/audio/transcriptionsfor you. Setserverto the host (and base path) without a trailing/v1— e.g. Groq ishttps://api.groq.com/openai, not.../openai/v1. A trailing/v1produces a doubled/v1/v1/...path that 404s;ears testwarns about this.
Secrets:
api_keyis stored in plaintext, so ears writes config files with0600permissions. TheEARS_API_KEY(and otherEARS_*) environment overrides only apply to interactive runs — a keybind-launchedears toggleinherits the graphical session environment, not your shell, so for push-to-talk the key must live in the config file.
Named profiles let you switch between ASR backends. Create config.{name}.toml alongside the default:
# ~/.config/ears/config.toml ← default (e.g. local whisper)
# ~/.config/ears/config.groq.toml ← Groq cloud API
ears -p groq # Launch TUI with Groq profile
ears -p groq toggle # Push-to-talk with Groq profileYou can also set the profile via environment variable:
export EARS_PROFILE=groq
ears togglePriority: -p flag > EARS_PROFILE env var > default config.toml.
ears server http://localhost:8178 # Set
ears server # Show currentValidate the active profile before binding a key to it. ears test prints a
summary (server, endpoint, model, device, language, and whether an API key is
set — masked), then runs a health check against the server:
ears test # Summary + server health check
ears -p groq test # Test a specific profile
ears test sample.wav # Also run a sample transcriptionIt exits non-zero if the server is unreachable or transcription fails, so it also works in scripts.
ears list # List available devices
ears select # Interactive selection (fzf)
ears current # Show current deviceEnvironment variables override config file values:
| Variable | Purpose |
|---|---|
EARS_SERVER |
Override whisper server URL |
EARS_DEVICE |
Override audio device |
EARS_LANGUAGE |
Override language code |
EARS_API_KEY |
Override API key |
EARS_MODEL |
Override model name |
EARS_PROFILE |
Set config profile |
earsLaunches an interactive terminal UI with status monitoring, VAD mode controls, configuration, and logs.
Bind ears toggle to a keyboard shortcut:
# Hyprland (~/.config/hypr/bindings.conf)
bind = SUPER SHIFT, V, exec, ears toggle
# i3/Sway
bindsym $mod+Shift+v exec ears toggleThen: press shortcut → speak → press again → text is typed.
Bash mode constrains the speech model's output to valid shell syntax, so spoken
commands land as code (ls → ls, not LS/Alice) instead of prose. You say
the command out loud; the grammar keeps it structurally valid bash. It is not
translation — say "git status", not "show me the git status".
Enable it per profile in ~/.config/ears/config.<name>.toml:
bash_mode = true # constrain output to the built-in bash grammar
auto_enter = false # recommended: type the command but DON'T run it
# guided_grammar = "..." # optional: override the built-in grammar (GBNF)Then use push-to-talk or VAD. Each VAD segment is treated as one discrete command and typed immediately, independently of progressive typing:
ears -p bash toggle # speak a command, toggle again → it's typed (not run)Notes:
- Requires a server with grammar-guided decoding. Bash mode routes requests to
the OpenAI-compatible
/v1/chat/completionsendpoint withstructured_outputs.grammar(e.g. vLLM); the plain transcription endpoint does not support it. Normal (non-bash) profiles are unaffected. - A configured
modelis required in bash mode. - The command allow-list lives in
grammars/bash.gbnf— extend it as needed.
ears vad # Start VAD (or stop if already running)Continuously listens and auto-transcribes when speech is detected. Toggle on/off by running the command again.
ears ghost listens hands-free like ears vad, but nothing is typed while
you speak. The transcript so far appears as inline ghost text at the cursor
of the focused app and is committed when you stop talking. It uses the
Wayland input-method preedit, delivered through a small fcitx5 addon, so it
works in any app with input-method support (terminals, browsers, GTK, Qt).
fcitx5-addon/install.sh # build + install the earsghost addon (needs fcitx5 headers)
ears ghost # hands-free (VAD) listening with ghost completion
ears toggle --ghost # push-to-talk: ghost text while recording, commit on the second pressHyprland binding example:
bindd = , F14, Ears ghost completion, exec, ears toggle --ghost
Notes:
- Apps draw preedit themselves. Most underline it; alacritty needs a small
patch adding
[colors.preedit] foreground/underlineto render grey, non-underlined ghost text. - If the addon is not reachable, final text is typed as usual.
- Before showing or committing, ears checks that the input method's app matches Hyprland's active window class. If they differ or either is unknown (e.g. a field that never enabled the input method), the ghost is cleared and the text is typed instead. This is an app-level check, not window identity: two windows of the same app (say two Hover windows) are indistinguishable.
- If fcitx5 stops answering mid-commit, ears cannot tell whether the text arrived, so it does not retype it (no duplicates, no Enter) and pauses typing.
ears typing offalso silences ghost output.ears ghostandears vadshare the toggle: either one stops the other.- Auto-Enter is not sent in VAD ghost mode; push-to-talk keeps
auto_enter.
The input method passes text and formatting hints; each app draws the ghost. Set the style once and ears writes it into the apps that support it:
[ghost]
color = "yellow" # #rrggbb, #rgb, or grey | blue-grey | orange | yellow | green-yellow
underline = false
frozen_color = "#ffffff" # the settled start, fixed during live decodingears ghost-style # show the style and where it is applied
ears ghost-style green-yellow # set a preset (or "#c8d44a", or "default")
ears ghost-style --underline
ears ghost-style --frozen "#ffffff" # colour the settled part (or "default")
With continuous live decoding, the start of the ghost settles while the last
few words are still open to correction. ears marks the settled part as the
input method's highlighted range (the fcitx5 addon must be current:
fcitx5-addon/install.sh), and frozen_color colours it, so you can see
what is fixed during live decoding while you speak. It remains uncommitted
preedit; with final_correction = true, a separate final transcription can
still replace it.
In the TUI's Configuration panel, o cycles the ghost colour presets and O
the frozen ones. Supported apps:
- Alacritty with the preedit-colors patch:
[colors.preedit](foreground,highlight_foreground) inalacritty.toml. Alacritty reloads it immediately. Withwrap = true(whichears ghost-stylesets), a ghost too long for its tmux pane row wraps over the pane's rows below instead of opening the fcitx popup. - Hover: the
hover.ime.ghost_preedit_colorandhover.ime.ghost_frozen_colorprefs in each profile'suser.js, which apply from the next Hover start.
Chromium, Firefox and GTK4 apps (Walker) keep their own preedit style. The
style is also re-applied whenever ears ghost or ears toggle --ghost starts.
By default the ghost re-transcribes the whole recording every 300 ms, so each update costs more than the last. With a Qwen3-ASR server, ears can decode continuously instead:
live_decoding = "continuous" # default: "repeat"
final_correction = false # commit the live result instead of re-transcribing
# live_rollback_words = 3 # words left open to revision; lower freezes sooner
# live_min_step_ms = 150 # streaming: least new audio between decodes (max 5000)The audio is sent as 8 s encoder windows (Qwen3-ASR's encoder never attends
across them), so vLLM's encoder and prefix caches reuse every finished window.
The text already settled is forced as the start of the answer, so the model
only decodes the new words; the last three stay open to revision. On a 33 s
clip, updates stayed at about 75 ms (versus climbing to 760 ms) and the final
text matched a full transcription. The server needs
--trust-request-chat-template; otherwise ears falls back to repeat.
Recordings longer than the server's context (about 90 s) are decoded in segments: after 60 s the current segment is finished at the next pause (at 80 s at the latest) and a new one starts there, so long dictation keeps its live preview.
With final_correction = true (the default) the committed text still comes
from a full transcription.
Words freeze (turn the frozen colour) once live_rollback_words newer words
follow them, so a name the model mishears is fixed in the live preview only
if it is spelled right in time. List names, acronyms and jargon you dictate
in prompt (for example prompt = "Heiervang, vLLM, Hyprland"): it is sent
as context with every live update and the final transcription.
When the server runs the ears_stream vLLM plugin, ears streams instead of
re-sending the audio on every tick: one WebSocket to /v1/ears/stream (the
server URL with ws:///wss://, same API key), over which only new audio
goes out and the growing transcript comes back as soon as each decode is done
(protocol: docs/STREAM_PROTOCOL.md). Push-to-talk
sends the recording every 50 ms; VAD mode streams one utterance per speech
segment, starting with the pre-speech buffer and cancelling it when the
segment is dropped. Nothing to configure: without the plugin (404, no answer
within 2 s, or a model it cannot serve) ears decodes per tick as above, and if
the connection drops mid-utterance that utterance is finished per tick from
the text already settled. The committed text comes from the same final path
either way, so nothing is ever committed twice.
The Qwen continuous path also tracks the exact frozen text prefix. Updated
ghost addons highlight the frozen prefix and underline both parts for inline
preedit; the app decides how those hints are rendered. The stream plugin supplies
model-tokenizer spans for inspection with stream_wav.py --tokens.
Use ears ghost-watch for the live split in any app, or --json for other visualizers.
Frozen means fixed during live decoding; full final correction can replace
it. See token freeze tracking for the guarantees,
visualization, and protocol fields.
ears Launch interactive TUI (default)
ears -p groq Launch TUI with named profile
ears toggle, t Toggle recording/transcription
ears vad, v Toggle VAD mode
ears list, l List audio devices
ears select, s Select device interactively
ears current, c Show current device
ears server [URL] Show or set whisper server URL
ears help Show help
States: Idle → Recording → Transcribing → Idle, plus VadActive for VAD mode.
State is persisted to $XDG_RUNTIME_DIR/ears/state and reconciled on startup.
- Stops
pw-recordprocess (SIGTERM) - Waits 300ms for file flush
- Validates WAV file (RIFF header check)
- Detects language from keyboard layout (if not configured)
- POSTs audio to
/v1/audio/transcriptionsendpoint - Filters silence artifacts ("Thank you.", etc.)
- Applies text filters (lowercase, punctuation removal)
- Types text via
wtypeor clipboard paste - Runs post-transcribe hook if configured
- Cleans up temporary files
Place an executable script at ~/.config/ears/hooks/post-transcribe. It receives:
$1- Path to a copy of the audio file$2- The transcribed text
The hook runs asynchronously and may outlive the Ears command. Its private audio copy is removed when the hook exits, including failure. A hook that delegates work to another background process must copy the audio before returning.
Place custom WAV files in ~/.local/share/ears-sounds/:
start.wav- Recording starteddone.wav- Transcription completebell.wav- Error occurred
Falls back to embedded sounds if not found. Embedded sounds are cached by content
in $XDG_CACHE_HOME/ears/sounds (normally ~/.cache/ears/sounds), so restarting
Ears reuses the same files. Corrupted entries are repaired automatically.
- Check server:
curl http://localhost:8178/health(local) orcurl -H "Authorization: Bearer $KEY" https://api.groq.com/openai/v1/models(cloud) - Check config:
ears server
- Recording may have timed out (2 minute limit)
- Check state:
cat $XDG_RUNTIME_DIR/ears/state - Check logs:
cat $XDG_RUNTIME_DIR/ears/debug.log
Desktop VAD publishes $XDG_RUNTIME_DIR/ears/vad-health.json once per second.
It identifies the desktop session, microphone, last audio/VAD progress, speech
probability, rejected-candidate count, audio backlog, and current processing
stage. It contains no audio or transcript text. A separate supervisor reports
missing audio or slow processing in debug.log, even if the processing task
is blocked. These are diagnostic warnings, not automatic restarts.
For detailed periodic measurements, start Ears with:
RUST_LOG=info,ears::health=debug ears vadears vad is a toggle: stop an existing VAD session before starting this way.
The health snapshot is always available; debug logging is selected at startup.
Only one desktop VAD health owner may run per state directory. Starting another
returns desktop VAD health owner already active; stop the existing desktop
session first. The operating system releases this ownership lock after a crash.
The updated HAIos ears bridge uses this snapshot to prevent Friend from showing
healthy listening for a stalled or unrelated Ears process. Existing audio cues
and existing IPC events keep their meanings.
If typing or Enter delivery fails, Ears pauses further keyboard input while continuing to transcribe. Check the target for partial text, then stop and restart VAD to resume typing. Friend reports this pause through the health snapshot. Clipboard paste leaves the transcript on the clipboard. Ears does not restore an older value, which could overwrite something copied during paste. If copying the transcript fails, Ears aborts before sending Ctrl+V.
A stopped snapshot is intentionally retained. If rolling back to an older Ears
binary that does not publish health, stop VAD and remove only
$XDG_RUNTIME_DIR/ears/vad-health.json before starting the older version; the
bridge then uses its legacy process/state check.
- Hyprland: ensure
wtypeis installed - Other: ensure
ydotooldis running (pgrep ydotoold) - Test manually:
wtype "test"orydotool type "test"
ears list # See all devices
ears select # Pick the right one
ears current # Verifyears/
├── src/
│ ├── main.rs # Entry point, command dispatch
│ ├── lib.rs # Library exports
│ ├── cli.rs # CLI argument parsing (clap)
│ ├── config.rs # Configuration management
│ ├── state.rs # Recording state machine
│ ├── lock.rs # File locking (single instance)
│ ├── process.rs # Child process management
│ ├── audio.rs # Audio device discovery
│ ├── whisper.rs # Whisper HTTP client
│ ├── desktop.rs # Notifications, feedback, text input, keyboard detection
│ ├── ducker.rs # System volume ducking during speech
│ ├── text_filters.rs # Text transformation filters
│ ├── streaming.rs # Streaming transcription + LocalAgreement
│ ├── streaming_engine.rs # Streaming engine coordinator
│ ├── vad.rs # Voice activity detection (Silero)
│ ├── continuous_capture.rs# Continuous audio capture
│ ├── progressive_typing.rs# Progressive text output
│ ├── ipc.rs # Unix domain socket IPC server
│ ├── ws_input.rs # WebSocket audio input mode
│ └── tui/
│ ├── mod.rs # TUI module exports
│ ├── app.rs # TUI application state
│ ├── ui.rs # TUI rendering
│ └── event.rs # TUI event handling
├── sounds/ # Embedded cue sounds
├── docs/ # Architecture and design docs
├── tests/ # Integration and snapshot tests
├── install.sh # Installation script
├── Cargo.toml # Rust package manifest
└── README.md # This file
cargo build # Debug build
cargo build --release # Release build
cargo test # Run all tests
cargo clippy # Lint
cargo fmt # Format
RUST_LOG=debug cargo run # Run with debug loggingRelease maintainers should follow docs/RELEASING.md. The
package version in Cargo.toml is the canonical version.
- Audio is sent to the configured whisper server only (defaults to localhost)
- API keys are stored in
config.toml(ensure appropriate file permissions) - Temporary audio files in
$XDG_RUNTIME_DIR(cleared on logout) - No audio is saved permanently
- No telemetry
MIT
Built with:
- whisper.cpp - Fast whisper inference
- PipeWire - Modern Linux audio
- ratatui - TUI framework
- wtype / ydotool - Text input automation
The profile-level filter_silence_artifacts setting controls known whole-transcript
hallucination heuristics in HTTP transcription, including grammar-constrained
requests. It defaults to true for compatibility. Set it to false if legitimate
short utterances such as “Thank you” are being discarded. This is separate from
[text_filters] transformations and alphabet filtering. Continuous forced-prefix
and vLLM streaming decoding do not use this HTTP heuristic. Restart an active
listener after changing the setting.
