A unified speech recognition API that integrates multiple ASR engines.
- File Recognition — Upload audio/video files or provide a URL, get complete transcription results
- Streaming File Recognition — Upload files or provide a URL, receive sentence-by-sentence results via SSE
- Real-time Streaming — Send audio streams via WebSocket, get real-time transcription results
- File Upload & Dedup — Upload files with MD5 fingerprint, instant upload for duplicate files
- Web Interface — Vue.js frontend with file/URL recognition, real-time streaming display, and SRT export
- Python 3.12+ (managed by uv)
- uv — fast Python package manager
- Node.js 18+
- FFmpeg (required for audio processing)
# 1. Create virtual environment (requires uv: https://docs.astral.sh/uv/)
uv venv --python 3.12
# 2. Install dependencies
uv pip install -r requirements.txt
# or install with dev dependencies:
uv pip install -e ".[dev]"
# 3. Start the server
uv run uvicorn server.main:app --host 0.0.0.0 --port 8020The server runs at http://localhost:8020. Visit http://localhost:8020/docs for interactive API documentation.
# 1. Navigate to web directory
cd web
# 2. Install dependencies
npm install
# 3. Start development server
npm run devThe frontend runs at http://localhost:3020 and automatically proxies API requests to the backend.
Terminal 1 - Backend:
cd OneASR
uv run uvicorn server.main:app --host 0.0.0.0 --port 8020Terminal 2 - Frontend:
cd OneASR/web
npm run dev# Check backend health
curl http://localhost:8020/health
# List available engines
curl -H "X-API-Key: oneasr-key" http://localhost:8020/api/v1/enginesAfter startup:
- Frontend UI:
http://localhost:3020 - API Docs:
http://localhost:8020/docs - Health Check:
http://localhost:8020/health
The project includes a media clipping tool for splitting long videos into shorter segments:
# Clip a specific duration (default 2 minutes)
python cli/clip.py input_video.mp4 120
# Auto-clip entire video into 2-minute segments
python cli/clip.py input_video.mp4Python API usage:
from cli.clip import MediaClipper
clipper = MediaClipper("video.mp4")
clipper.clip(start=0, duration=120, output="clip.mp4")
clips = clipper.auto_clip(clip_duration=120, output_dir="clips/")OneASR supports various local ASR engines (Qwen3-ASR, Faster-Whisper, FireRedASR, etc.). Models are stored by default under the models/ directory in OneASR/. You can download models using either Hugging Face CLI or ModelScope.
# Option A: Hugging Face CLI (provides `hf` command)
pip install -U "huggingface_hub[cli]"
# Optional: China mirror acceleration
export HF_ENDPOINT=https://hf-mirror.com
# Option B: ModelScope CLI (recommended for users in China)
pip install -U modelscopeNote: Run these commands from the
OneASR/directory to save models directly tomodels/{model_dir}.
Recommended: Qwen/Qwen3-ASR-1.7B (High quality) or Qwen/Qwen3-ASR-0.6B (Lightweight)
- Hugging Face (
hf):hf download Qwen/Qwen3-ASR-1.7B --local-dir models/Qwen3-ASR-1.7B
- ModelScope CLI:
modelscope download --model Qwen/Qwen3-ASR-1.7B --local_dir models/Qwen3-ASR-1.7B
Used for word/character-level timestamps: Qwen/Qwen3-ForcedAligner-0.6B
- Hugging Face (
hf):hf download Qwen/Qwen3-ForcedAligner-0.6B --local-dir models/Qwen3-ForcedAligner-0.6B
- ModelScope CLI:
modelscope download --model Qwen/Qwen3-ForcedAligner-0.6B --local_dir models/Qwen3-ForcedAligner-0.6B
Supported: Systran/faster-whisper-medium, large-v3, small, etc.
- Hugging Face (
hf):hf download Systran/faster-whisper-medium --local-dir models/faster-whisper-medium
- ModelScope CLI:
modelscope download --model Systran/faster-whisper-medium --local_dir models/faster-whisper-medium
Supported: FireRedTeam/FireRedASR-AED-L, FireRedTeam/FireRedASR-LLM-L
- Hugging Face (
hf):hf download FireRedTeam/FireRedASR-AED-L --local-dir models/FireRedASR-AED-L
- ModelScope CLI:
modelscope download --model FireRedTeam/FireRedASR-AED-L --local_dir models/FireRedASR-AED-L
Model: GilgameshWind/X-ASR-zh-en — Chinese/English streaming recognition using zipformer2 transducer
-
Hugging Face (
hf):hf download GilgameshWind/X-ASR-zh-en \ --include "deployment/models/chunk-160ms-model/*" \ --local-dir models mv models/deployment/models/chunk-160ms-model models/chunk-160ms-model -
ModelScope CLI:
modelscope download --model Gilgamesh-J/X-ASR-zh-en \ --include "deployment/models/chunk-160ms-model/*" \ --local_dir models mv models/deployment/models/chunk-160ms-model models/chunk-160ms-model
# download_models.py
from modelscope import snapshot_download
MODELS = {
"models/Qwen3-ASR-1.7B": "Qwen/Qwen3-ASR-1.7B",
"models/Qwen3-ForcedAligner-0.6B": "Qwen/Qwen3-ForcedAligner-0.6B",
"models/faster-whisper-medium": "Systran/faster-whisper-medium",
}
for local_path, model_id in MODELS.items():
print(f"Downloading {model_id} to {local_path} via ModelScope...")
snapshot_download(model_id, local_dir=local_path)
print("All models downloaded successfully!")After downloading, configure the corresponding provider in OneASR/config.yaml:
ASR-Providers:
# Qwen3-ASR Configuration
qwen:
enable: true
engine: qwen
load:
model_name: Qwen/Qwen3-ASR-1.7B
model_path: models/Qwen3-ASR-1.7B
device: cpu # or cuda:0 / mps
dtype: float32
max_new_tokens: 256
max_inference_batch_size: 32
# Optional: forced aligner
# forced_aligner_name: Qwen/Qwen3-ForcedAligner-0.6B
# forced_aligner_path: models/Qwen3-ForcedAligner-0.6B
# Faster-Whisper Configuration
faster-whisper:
enable: false
engine: faster-whisper
load:
model_name: medium
model_path: models/faster-whisper-medium
device: cpu
compute_type: int8All API endpoints (except /health) require an API Key via the X-API-Key request header.
Configure the API Key in config.yaml:
api_key: oneasr-key# Call API with API Key
curl -H "X-API-Key: oneasr-key" http://localhost:8020/api/v1/engines
# Upload file for recognition
curl -X POST -H "X-API-Key: oneasr-key" \
http://localhost:8020/api/v1/transcribe/file \
-F "file=@audio.mp3" -F "format=srt" -o subtitle.srtWebSocket streaming uses query parameters (browser WebSocket API doesn't support custom headers): ws://localhost:8020/ws/transcribe/stream?api_key=oneasr-key
| Endpoint | Method | Description |
|---|---|---|
/api/v1/audio/transcriptions |
POST | Create transcription (file upload or file_uuid) |
/api/v1/audio/transcriptions/stream |
POST | Create streaming transcription (SSE) |
/api/v1/audio/models |
GET | List available models |
/api/v1/files/upload |
POST | Upload file (supports MD5 instant upload) |
/api/v1/files/list |
GET | List all uploaded files |
/api/v1/files/{file_id} |
GET | Get file info |
/api/v1/files/{file_id} |
DELETE | Delete uploaded file |
/health |
GET | Health check |
| Endpoint | Method | Description |
|---|---|---|
/v1/realtime?api_key= |
WebSocket | Real-time streaming (OpenAI Realtime Transcription 协议) |
/ws/transcribe/stream?api_key= |
WebSocket | Real-time streaming (Legacy WhisperLiveKit 协议) |
Upload files with optional MD5 fingerprint for instant upload of duplicate files:
# First upload: file is saved and metadata stored in DB
curl -X POST -H "X-API-Key: oneasr-key" \
"http://localhost:8020/api/v1/files/upload?file_md5=abc123&file_size=1024000" \
-F "file=@audio.mp3"
# Response: {"duplicate": false, "file_id": "uuid-1", ...}
# Second upload with same file: instant return (no re-upload)
curl -X POST -H "X-API-Key: oneasr-key" \
"http://localhost:8020/api/v1/files/upload?file_md5=abc123&file_size=1024000" \
-F "file=@audio_copy.mp3"
# Response: {"duplicate": true, "file_id": "uuid-1", ...}Real-time speech recognition powered by WhisperLiveKit:
- VAD/VAC — Silero Voice Activity Detection, automatic speech/silence boundary detection
- SimulStreaming — Low-latency streaming strategy (default), or LocalAgreement high-accuracy strategy
- Timestamp Alignment — Each line has precise start/end timestamps
- Speaker Diarization — Optional speaker identification (requires additional models)
- Multi-language — Automatic language detection or manual specification
const ws = new WebSocket("ws://localhost:8020/v1/realtime?api_key=oneasr-key");
// 1. 配置会话
ws.send(JSON.stringify({
type: "session.update",
session: {
type: "transcription",
audio: {
input: {
format: { type: "audio/pcm", rate: 16000 },
transcription: { model: "wlk-live", language: "zh" },
},
},
},
}));
ws.onmessage = (event) => {
const data = JSON.parse(event.data);
if (data.type === "session.updated") {
console.log("Session configured:", data.session.id);
} else if (data.type === "conversation.item.input_audio_transcription.delta") {
process.stdout.write(data.delta); // 增量文本
} else if (data.type === "conversation.item.input_audio_transcription.completed") {
console.log("\nFinal:", data.transcript); // 完整文本
}
};
// 2. 发送 base64 编码的 PCM 音频
ws.send(JSON.stringify({
type: "input_audio_buffer.append",
audio: btoa(String.fromCharCode(...pcmBytes)),
}));
// 3. 提交音频缓冲区
ws.send(JSON.stringify({ type: "input_audio_buffer.commit" }));CLI 工具,将本地音视频文件模拟为麦克风实时输入:
# 基本用法
python -m cli.stream_simulation_client test.wav
# 指定中文和引擎
python -m cli.stream_simulation_client test.wav --language zh --model wlk-live
# 10 倍速快速测试
python -m cli.stream_simulation_client test.wav --speed 10
# 连接远程服务
python -m cli.stream_simulation_client test.wav --url ws://remote:8020/v1/realtime# Run all tests
uv run pytest tests/ -v
# Run API tests only (fast, no external dependencies)
uv run pytest tests/api/ -v
# Run general/unit tests only
uv run pytest tests/general/ -v
# Run a specific test file
uv run pytest tests/api/test_file_upload.py -v
# Skip integration tests (need real audio files and running services)
uv run pytest tests/ -m "not integration" -v
# Run with short traceback
uv run pytest tests/api/ -v --tb=shorttests/
├── conftest.py # Shared fixtures (client, DB cleanup)
├── api/ # API endpoint tests (TestClient)
│ ├── test_audio.py # Audio transcription API
│ ├── test_file_upload.py # File upload, MD5 dedup, CRUD
│ ├── test_models.py # Models listing, format params
│ └── test_streaming.py # SSE streaming transcription
└── general/ # Unit tests & integration tests
├── test_format.py # Output format conversion (SRT/VTT/JSON/TSV)
├── test_engines.py # Engine loading and transcription
├── test_mimo.py # MiMo audio understanding engine
├── test_stream_websocket.py # WebSocket streaming integration
├── test_stream_realtime.py # OpenAI Realtime Transcription protocol tests
└── test_stream_whisperlivekit.py # WhisperLiveKit engine tests
Edit config.yaml to configure API Key and engines:
api_key: oneasr-key
default_provider: whisper1
providers:
whisper1:
engine: faster-whisper
type: local
model_name: medium
device: cpu
compute_type: int8
wlk-live:
engine: whisperlivekit
type: local
model_name: base
device: cpu
compute_type: int8
backend: auto
backend_policy: simulstreaming
language: auto
vac: true
diarization: false
pcm_input: trueDetailed designs and guides are located in the docs/ directory:
- API Design & Reference
- API Design — Architecture, audio chunking, and streaming design
- API Reference — Endpoint specifications, parameters, and examples
- Architecture & Toolkits
- Project Design — System layering, state machine, and scheduling
- ASR Toolkit Design — VAD, resampling, and audio chunking
- Guides & Best Practices
- Model Guide — FireRedASR, SenseVoice, Whisper, and cloud engines
- uv Guide — Fast dependency and environment management
- Development Guide — Code style and testing conventions
- References
- yt-dlp Reference — Media extraction reference
OneASR/
├── docs/ # Documentation center
│ ├── api/ # API design and reference
│ ├── architecture/ # Architectural specifications
│ ├── guides/ # Model and environment guides
│ ├── references/ # Upstream third-party references
│ └── assets/ # Icons and visual assets
├── app/ # Core backend application
│ ├── main.py # FastAPI entry point & lifespan
│ ├── api/ # Routing layer (controllers)
│ │ ├── auth.py # API Key authentication
│ │ ├── audio.py # OpenAI-compatible audio API
│ │ ├── file_upload.py # File upload, MD5 dedup, & URL import
│ │ ├── file_transcription.py # Async file transcription & SSE stream
│ │ ├── media.py # Media download & parse API
│ │ ├── model.py # Models listing API
│ │ ├── provider.py # Providers listing API
│ │ ├── realtime.py # Realtime WebSocket transcription
│ │ └── realtime_ext.py # Extended realtime WebSocket API
│ ├── core/ # Core configs, errors, & storage
│ │ ├── config.py # YAML config management
│ │ ├── errors.py # Unified exceptions
│ │ └── file_storage.py # File storage utilities
│ ├── db/ # Database session & engine
│ │ ├── base.py # SQLAlchemy declarative base
│ │ └── session.py # Async engine & sessionmaker
│ ├── engines/ # ASR engine adapters
│ │ ├── base.py # Abstract engine base class
│ │ ├── whisper_engine.py # faster-whisper adapter
│ │ ├── firered_engine.py # FireRedASR adapter
│ │ ├── qwen_engine.py # Qwen ASR adapter
│ │ ├── xasr_engine.py # Sherpa / SenseVoice adapter
│ │ ├── openai_engine.py # OpenAI-compatible engine
│ │ ├── mimo_engine.py # Xiaomi MiMo adapter
│ │ └── registry.py # Engine registry & factory
│ ├── models/ # SQLAlchemy ORM models
│ │ └── orm_models.py # Database entity tables
│ ├── schemas/ # Pydantic request/response schemas (DTOs)
│ │ ├── audio.py # Audio transcription schemas
│ │ ├── file.py # File asset & URL import schemas
│ │ ├── media.py # Media parse & task schemas
│ │ └── transcription.py # Async transcription task schemas
│ ├── services/ # Business logic service layer
│ │ ├── file_service.py # File transcription task service
│ │ ├── media_service.py # Media download & state machine
│ │ └── record_service.py # Transcription records service
│ └── utils/ # Audio & processing utilities
│ ├── audio.py # Format conversion & duration probe
│ ├── asr_toolkit.py # VAD segmentation & streaming chunks
│ ├── audio_converter.py # Resampling & channel conversion
│ ├── download.py # URL download utility
│ ├── format.py # Output formatting (SRT/VTT/JSON/TSV)
│ ├── stream.py # PCM streaming utilities
│ ├── vad.py # Silero VAD integration
│ └── video_url.py # yt-dlp & direct URL extractor
├── cli/ # CLI & client test tools
│ ├── clip.py # Media clipping tool
│ ├── converter.py # Audio converter
│ └── stream_simulation_client.py # Simulation client
├── web/ # Vue 3 frontend web UI
├── tests/ # Automated test suites
├── models/ # Model storage directory
├── data/ # SQLite database directory
├── uploads/ # Uploaded files storage
├── download/ # Media download temporary directory
├── config.yaml # Engine & server configuration
├── pyproject.toml # Project metadata & dependencies
└── requirements.txt # Pinned requirements
This project is licensed under the Apache 2.0 License - see the LICENSE file for details.