Skip to content
heiervang-technologiesPublic

About

ASR tool for linux

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

ears

Voice-to-text for your Linux desktop. Press a hotkey, speak, and your words appear wherever the cursor is. Or leave VAD mode on and let it transcribe continuously, hands-free.

Backend-agnostic — works with local whisper.cpp, faster-whisper, or any OpenAI-compatible ASR endpoint (Groq, OpenAI, your own server).

ears TUI demo

Features

  • Interactive TUI — Terminal UI with real-time status, VAD mode, live transcription, and configuration (default mode)
  • Push-to-talk — bind ears toggle to a keyboard shortcut for quick dictation
  • VAD mode — hands-free voice activity detection with auto-transcription (ears vad)
  • Streaming transcription — real-time text output with LocalAgreement for stable progressive output
  • Volume ducking — optionally lowers system volume while you're speaking
  • Bash mode — constrain dictation to valid shell syntax via grammar-guided decoding (speak commands, get code)
  • Profiles — switch between local whisper and cloud APIs (Groq, OpenAI, etc.) per-invocation
  • Text filters — optional lowercase conversion and punctuation removal
  • Language detection — auto-detects from keyboard layout (Hyprland + GNOME)
  • Smart text input — uses wtype on Hyprland/Wayland, clipboard paste via ydotool elsewhere
  • PipeWire audio — native support for the modern Linux audio stack
  • Post-transcribe hooks — run custom scripts after each transcription
  • Audio feedback — embedded cue sounds, customizable with custom sound override support
  • State management — file-based locking and state with automatic crash recovery
  • No telemetry — audio goes only to the server you configure

Prerequisites

Required

  • Linux with PipeWire audio system
  • A whisper.cpp or OpenAI-compatible ASR server running
  • Text input tool: wtype (Hyprland/Wayland) or ydotool (other systems)

Optional

  • notify-send for desktop notifications
  • paplay for audio feedback
  • fzf for interactive device selection
  • wl-clipboard (wl-copy) for clipboard-based text input on non-Hyprland systems

Installing Dependencies

# Arch Linux (Hyprland/Omarchy)
sudo pacman -S pipewire wtype libnotify pulseaudio fzf

# Ubuntu/Debian
sudo apt install pipewire ydotool wl-clipboard libnotify-bin pulseaudio-utils fzf

# Fedora
sudo dnf install pipewire ydotool libnotify pulseaudio-utils fzf

Setting up a Whisper Server

  1. Clone and build whisper.cpp:
git clone https://github.com/ggerganov/whisper.cpp
cd whisper.cpp
make server            # CPU only
make server WHISPER_CUDA=1  # With NVIDIA GPU
  1. Download a model:
bash ./models/download-ggml-model.sh base.en
  1. Start the server:
./server -m models/ggml-base.en.bin -p 8178

Installation

From GitHub Releases

mkdir -p ~/.local/bin
gh release download latest --repo heiervang-technologies/ears --pattern 'ears' --dir ~/.local/bin --clobber
chmod +x ~/.local/bin/ears
export PATH="$HOME/.local/bin:$PATH"

From Source

Building from source requires Rust 1.88 or newer.

git clone https://github.com/heiervang-technologies/ears
cd ears
cargo build --release
cargo install --path .

Using install.sh

git clone https://github.com/heiervang-technologies/ears
cd ears
./install.sh

Configuration

Configuration is stored in ~/.config/ears/config.toml:

server = "http://127.0.0.1:8178"
device = "alsa_input.usb-..."
# language = "en"          # Optional (auto-detects from keyboard layout)
# api_key = "sk-..."       # Optional (for authenticated ASR services)
# model = "whisper-large-v3-turbo"  # Optional (for cloud APIs that require it)
# prompt = "vLLM, PyTorch" # Optional (context biasing: names, acronyms, jargon)
filter_silence_artifacts = true # Set false to retain exact phrases such as "Thank you"

[text_filters]
lowercase = false
remove_punctuation = false

[vad]
# Increase this for longer mid-sentence pauses; decrease it for faster dispatch.
max_silence_duration_ms = 1200

Server URL: ears sends requests to {server}/v1/audio/transcriptions, appending /v1/audio/transcriptions for you. Set server to the host (and base path) without a trailing /v1 — e.g. Groq is https://api.groq.com/openai, not .../openai/v1. A trailing /v1 produces a doubled /v1/v1/... path that 404s; ears test warns about this.

Secrets: api_key is stored in plaintext, so ears writes config files with 0600 permissions. The EARS_API_KEY (and other EARS_*) environment overrides only apply to interactive runs — a keybind-launched ears toggle inherits the graphical session environment, not your shell, so for push-to-talk the key must live in the config file.

Profiles

Named profiles let you switch between ASR backends. Create config.{name}.toml alongside the default:

# ~/.config/ears/config.toml        ← default (e.g. local whisper)
# ~/.config/ears/config.groq.toml   ← Groq cloud API

ears -p groq          # Launch TUI with Groq profile
ears -p groq toggle   # Push-to-talk with Groq profile

You can also set the profile via environment variable:

export EARS_PROFILE=groq
ears toggle

Priority: -p flag > EARS_PROFILE env var > default config.toml.

Set server URL

ears server http://localhost:8178   # Set
ears server                          # Show current

Test the configuration

Validate the active profile before binding a key to it. ears test prints a summary (server, endpoint, model, device, language, and whether an API key is set — masked), then runs a health check against the server:

ears test                 # Summary + server health check
ears -p groq test         # Test a specific profile
ears test sample.wav      # Also run a sample transcription

It exits non-zero if the server is unreachable or transcription fails, so it also works in scripts.

Select microphone

ears list      # List available devices
ears select    # Interactive selection (fzf)
ears current   # Show current device

Environment variables

Environment variables override config file values:

Variable Purpose
EARS_SERVER Override whisper server URL
EARS_DEVICE Override audio device
EARS_LANGUAGE Override language code
EARS_API_KEY Override API key
EARS_MODEL Override model name
EARS_PROFILE Set config profile

Usage

TUI Mode (default)

ears

Launches an interactive terminal UI with status monitoring, VAD mode controls, configuration, and logs.

Push-to-Talk (keyboard shortcut)

Bind ears toggle to a keyboard shortcut:

# Hyprland (~/.config/hypr/bindings.conf)
bind = SUPER SHIFT, V, exec, ears toggle

# i3/Sway
bindsym $mod+Shift+v exec ears toggle

Then: press shortcut → speak → press again → text is typed.

Bash Mode (dictate shell commands)

Bash mode constrains the speech model's output to valid shell syntax, so spoken commands land as code (ls → ls, not LS/Alice) instead of prose. You say the command out loud; the grammar keeps it structurally valid bash. It is not translation — say "git status", not "show me the git status".

Enable it per profile in ~/.config/ears/config.<name>.toml:

bash_mode = true             # constrain output to the built-in bash grammar
auto_enter = false           # recommended: type the command but DON'T run it
# guided_grammar = "..."     # optional: override the built-in grammar (GBNF)

Then use push-to-talk or VAD. Each VAD segment is treated as one discrete command and typed immediately, independently of progressive typing:

ears -p bash toggle    # speak a command, toggle again → it's typed (not run)

Notes:

  • Requires a server with grammar-guided decoding. Bash mode routes requests to the OpenAI-compatible /v1/chat/completions endpoint with structured_outputs.grammar (e.g. vLLM); the plain transcription endpoint does not support it. Normal (non-bash) profiles are unaffected.
  • A configured model is required in bash mode.
  • The command allow-list lives in grammars/bash.gbnf — extend it as needed.

VAD Mode (headless)

ears vad    # Start VAD (or stop if already running)

Continuously listens and auto-transcribes when speech is detected. Toggle on/off by running the command again.

Ghost Completion (inline preview)

ears ghost listens hands-free like ears vad, but nothing is typed while you speak. The transcript so far appears as inline ghost text at the cursor of the focused app and is committed when you stop talking. It uses the Wayland input-method preedit, delivered through a small fcitx5 addon, so it works in any app with input-method support (terminals, browsers, GTK, Qt).

fcitx5-addon/install.sh   # build + install the earsghost addon (needs fcitx5 headers)
ears ghost                # hands-free (VAD) listening with ghost completion
ears toggle --ghost       # push-to-talk: ghost text while recording, commit on the second press

Hyprland binding example:

bindd = , F14, Ears ghost completion, exec, ears toggle --ghost

Notes:

  • Apps draw preedit themselves. Most underline it; alacritty needs a small patch adding [colors.preedit] foreground/underline to render grey, non-underlined ghost text.
  • If the addon is not reachable, final text is typed as usual.
  • Before showing or committing, ears checks that the input method's app matches Hyprland's active window class. If they differ or either is unknown (e.g. a field that never enabled the input method), the ghost is cleared and the text is typed instead. This is an app-level check, not window identity: two windows of the same app (say two Hover windows) are indistinguishable.
  • If fcitx5 stops answering mid-commit, ears cannot tell whether the text arrived, so it does not retype it (no duplicates, no Enter) and pauses typing.
  • ears typing off also silences ghost output.
  • ears ghost and ears vad share the toggle: either one stops the other.
  • Auto-Enter is not sent in VAD ghost mode; push-to-talk keeps auto_enter.

Ghost text style

The input method passes text and formatting hints; each app draws the ghost. Set the style once and ears writes it into the apps that support it:

[ghost]
color = "yellow"   # #rrggbb, #rgb, or grey | blue-grey | orange | yellow | green-yellow
underline = false
frozen_color = "#ffffff"   # the settled start, fixed during live decoding
ears ghost-style                  # show the style and where it is applied
ears ghost-style green-yellow     # set a preset (or "#c8d44a", or "default")
ears ghost-style --underline
ears ghost-style --frozen "#ffffff"  # colour the settled part (or "default")

With continuous live decoding, the start of the ghost settles while the last few words are still open to correction. ears marks the settled part as the input method's highlighted range (the fcitx5 addon must be current: fcitx5-addon/install.sh), and frozen_color colours it, so you can see what is fixed during live decoding while you speak. It remains uncommitted preedit; with final_correction = true, a separate final transcription can still replace it.

In the TUI's Configuration panel, o cycles the ghost colour presets and O the frozen ones. Supported apps:

  • Alacritty with the preedit-colors patch: [colors.preedit] (foreground, highlight_foreground) in alacritty.toml. Alacritty reloads it immediately. With wrap = true (which ears ghost-style sets), a ghost too long for its tmux pane row wraps over the pane's rows below instead of opening the fcitx popup.
  • Hover: the hover.ime.ghost_preedit_color and hover.ime.ghost_frozen_color prefs in each profile's user.js, which apply from the next Hover start.

Chromium, Firefox and GTK4 apps (Walker) keep their own preedit style. The style is also re-applied whenever ears ghost or ears toggle --ghost starts.

Continuous live decoding (Qwen3-ASR on vLLM)

By default the ghost re-transcribes the whole recording every 300 ms, so each update costs more than the last. With a Qwen3-ASR server, ears can decode continuously instead:

live_decoding = "continuous"   # default: "repeat"
final_correction = false       # commit the live result instead of re-transcribing
# live_rollback_words = 3      # words left open to revision; lower freezes sooner
# live_min_step_ms = 150       # streaming: least new audio between decodes (max 5000)

The audio is sent as 8 s encoder windows (Qwen3-ASR's encoder never attends across them), so vLLM's encoder and prefix caches reuse every finished window. The text already settled is forced as the start of the answer, so the model only decodes the new words; the last three stay open to revision. On a 33 s clip, updates stayed at about 75 ms (versus climbing to 760 ms) and the final text matched a full transcription. The server needs --trust-request-chat-template; otherwise ears falls back to repeat. Recordings longer than the server's context (about 90 s) are decoded in segments: after 60 s the current segment is finished at the next pause (at 80 s at the latest) and a new one starts there, so long dictation keeps its live preview. With final_correction = true (the default) the committed text still comes from a full transcription.

Words freeze (turn the frozen colour) once live_rollback_words newer words follow them, so a name the model mishears is fixed in the live preview only if it is spelled right in time. List names, acronyms and jargon you dictate in prompt (for example prompt = "Heiervang, vLLM, Hyprland"): it is sent as context with every live update and the final transcription.

When the server runs the ears_stream vLLM plugin, ears streams instead of re-sending the audio on every tick: one WebSocket to /v1/ears/stream (the server URL with ws:///wss://, same API key), over which only new audio goes out and the growing transcript comes back as soon as each decode is done (protocol: docs/STREAM_PROTOCOL.md). Push-to-talk sends the recording every 50 ms; VAD mode streams one utterance per speech segment, starting with the pre-speech buffer and cancelling it when the segment is dropped. Nothing to configure: without the plugin (404, no answer within 2 s, or a model it cannot serve) ears decodes per tick as above, and if the connection drops mid-utterance that utterance is finished per tick from the text already settled. The committed text comes from the same final path either way, so nothing is ever committed twice.

The Qwen continuous path also tracks the exact frozen text prefix. Updated ghost addons highlight the frozen prefix and underline both parts for inline preedit; the app decides how those hints are rendered. The stream plugin supplies model-tokenizer spans for inspection with stream_wav.py --tokens. Use ears ghost-watch for the live split in any app, or --json for other visualizers. Frozen means fixed during live decoding; full final correction can replace it. See token freeze tracking for the guarantees, visualization, and protocol fields.

All Commands

ears                   Launch interactive TUI (default)
ears -p groq           Launch TUI with named profile
ears toggle, t         Toggle recording/transcription
ears vad, v            Toggle VAD mode
ears list, l           List audio devices
ears select, s         Select device interactively
ears current, c        Show current device
ears server [URL]      Show or set whisper server URL
ears help              Show help

How It Works

State Machine

States: Idle → Recording → Transcribing → Idle, plus VadActive for VAD mode.

State is persisted to $XDG_RUNTIME_DIR/ears/state and reconciled on startup.

Transcription Flow

  1. Stops pw-record process (SIGTERM)
  2. Waits 300ms for file flush
  3. Validates WAV file (RIFF header check)
  4. Detects language from keyboard layout (if not configured)
  5. POSTs audio to /v1/audio/transcriptions endpoint
  6. Filters silence artifacts ("Thank you.", etc.)
  7. Applies text filters (lowercase, punctuation removal)
  8. Types text via wtype or clipboard paste
  9. Runs post-transcribe hook if configured
  10. Cleans up temporary files

Post-Transcribe Hook

Place an executable script at ~/.config/ears/hooks/post-transcribe. It receives:

  • $1 - Path to a copy of the audio file
  • $2 - The transcribed text

The hook runs asynchronously and may outlive the Ears command. Its private audio copy is removed when the hook exits, including failure. A hook that delegates work to another background process must copy the audio before returning.

Custom Sounds

Place custom WAV files in ~/.local/share/ears-sounds/:

  • start.wav - Recording started
  • done.wav - Transcription complete
  • bell.wav - Error occurred

Falls back to embedded sounds if not found. Embedded sounds are cached by content in $XDG_CACHE_HOME/ears/sounds (normally ~/.cache/ears/sounds), so restarting Ears reuses the same files. Corrupted entries are repaired automatically.

Troubleshooting

"Whisper server not running!"

  • Check server: curl http://localhost:8178/health (local) or curl -H "Authorization: Bearer $KEY" https://api.groq.com/openai/v1/models (cloud)
  • Check config: ears server

"No active recording"

  • Recording may have timed out (2 minute limit)
  • Check state: cat $XDG_RUNTIME_DIR/ears/state
  • Check logs: cat $XDG_RUNTIME_DIR/ears/debug.log

VAD appears unresponsive

Desktop VAD publishes $XDG_RUNTIME_DIR/ears/vad-health.json once per second. It identifies the desktop session, microphone, last audio/VAD progress, speech probability, rejected-candidate count, audio backlog, and current processing stage. It contains no audio or transcript text. A separate supervisor reports missing audio or slow processing in debug.log, even if the processing task is blocked. These are diagnostic warnings, not automatic restarts.

For detailed periodic measurements, start Ears with:

RUST_LOG=info,ears::health=debug ears vad

ears vad is a toggle: stop an existing VAD session before starting this way. The health snapshot is always available; debug logging is selected at startup. Only one desktop VAD health owner may run per state directory. Starting another returns desktop VAD health owner already active; stop the existing desktop session first. The operating system releases this ownership lock after a crash. The updated HAIos ears bridge uses this snapshot to prevent Friend from showing healthy listening for a stalled or unrelated Ears process. Existing audio cues and existing IPC events keep their meanings.

If typing or Enter delivery fails, Ears pauses further keyboard input while continuing to transcribe. Check the target for partial text, then stop and restart VAD to resume typing. Friend reports this pause through the health snapshot. Clipboard paste leaves the transcript on the clipboard. Ears does not restore an older value, which could overwrite something copied during paste. If copying the transcript fails, Ears aborts before sending Ctrl+V.

A stopped snapshot is intentionally retained. If rolling back to an older Ears binary that does not publish health, stop VAD and remove only $XDG_RUNTIME_DIR/ears/vad-health.json before starting the older version; the bridge then uses its legacy process/state check.

Text isn't being typed

  • Hyprland: ensure wtype is installed
  • Other: ensure ydotoold is running (pgrep ydotoold)
  • Test manually: wtype "test" or ydotool type "test"

Wrong microphone

ears list       # See all devices
ears select     # Pick the right one
ears current    # Verify

Development

Project Structure

ears/
├── src/
│   ├── main.rs              # Entry point, command dispatch
│   ├── lib.rs               # Library exports
│   ├── cli.rs               # CLI argument parsing (clap)
│   ├── config.rs            # Configuration management
│   ├── state.rs             # Recording state machine
│   ├── lock.rs              # File locking (single instance)
│   ├── process.rs           # Child process management
│   ├── audio.rs             # Audio device discovery
│   ├── whisper.rs           # Whisper HTTP client
│   ├── desktop.rs           # Notifications, feedback, text input, keyboard detection
│   ├── ducker.rs            # System volume ducking during speech
│   ├── text_filters.rs      # Text transformation filters
│   ├── streaming.rs         # Streaming transcription + LocalAgreement
│   ├── streaming_engine.rs  # Streaming engine coordinator
│   ├── vad.rs               # Voice activity detection (Silero)
│   ├── continuous_capture.rs# Continuous audio capture
│   ├── progressive_typing.rs# Progressive text output
│   ├── ipc.rs               # Unix domain socket IPC server
│   ├── ws_input.rs          # WebSocket audio input mode
│   └── tui/
│       ├── mod.rs           # TUI module exports
│       ├── app.rs           # TUI application state
│       ├── ui.rs            # TUI rendering
│       └── event.rs         # TUI event handling
├── sounds/                  # Embedded cue sounds
├── docs/                    # Architecture and design docs
├── tests/                   # Integration and snapshot tests
├── install.sh               # Installation script
├── Cargo.toml               # Rust package manifest
└── README.md                # This file

Build & Test

cargo build                    # Debug build
cargo build --release          # Release build
cargo test                     # Run all tests
cargo clippy                   # Lint
cargo fmt                      # Format
RUST_LOG=debug cargo run       # Run with debug logging

Release maintainers should follow docs/RELEASING.md. The package version in Cargo.toml is the canonical version.

Security

  • Audio is sent to the configured whisper server only (defaults to localhost)
  • API keys are stored in config.toml (ensure appropriate file permissions)
  • Temporary audio files in $XDG_RUNTIME_DIR (cleared on logout)
  • No audio is saved permanently
  • No telemetry

License

MIT

Credits

Built with:

The profile-level filter_silence_artifacts setting controls known whole-transcript hallucination heuristics in HTTP transcription, including grammar-constrained requests. It defaults to true for compatibility. Set it to false if legitimate short utterances such as “Thank you” are being discarded. This is separate from [text_filters] transformations and alphabet filtering. Continuous forced-prefix and vLLM streaming decoding do not use this HTTP heuristic. Restart an active listener after changing the setting.

About

ASR tool for linux

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages