Skip to content
codeAndxvPublic

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

OneASR

OneASR Logo

A unified speech recognition API that integrates multiple ASR engines.

中文文档

Features

  • File Recognition — Upload audio/video files or provide a URL, get complete transcription results
  • Streaming File Recognition — Upload files or provide a URL, receive sentence-by-sentence results via SSE
  • Real-time Streaming — Send audio streams via WebSocket, get real-time transcription results
  • File Upload & Dedup — Upload files with MD5 fingerprint, instant upload for duplicate files
  • Web Interface — Vue.js frontend with file/URL recognition, real-time streaming display, and SRT export

Quick Start

Prerequisites

  • Python 3.12+ (managed by uv)
  • uv — fast Python package manager
  • Node.js 18+
  • FFmpeg (required for audio processing)

Backend Setup

# 1. Create virtual environment (requires uv: https://docs.astral.sh/uv/)
uv venv --python 3.12

# 2. Install dependencies
uv pip install -r requirements.txt
# or install with dev dependencies:
uv pip install -e ".[dev]"

# 3. Start the server
uv run uvicorn server.main:app --host 0.0.0.0 --port 8020

The server runs at http://localhost:8020. Visit http://localhost:8020/docs for interactive API documentation.

Frontend Setup

# 1. Navigate to web directory
cd web

# 2. Install dependencies
npm install

# 3. Start development server
npm run dev

The frontend runs at http://localhost:3020 and automatically proxies API requests to the backend.

Quick Launch (Two Terminals)

Terminal 1 - Backend:

cd OneASR
uv run uvicorn server.main:app --host 0.0.0.0 --port 8020

Terminal 2 - Frontend:

cd OneASR/web
npm run dev

Verify Installation

# Check backend health
curl http://localhost:8020/health

# List available engines
curl -H "X-API-Key: oneasr-key" http://localhost:8020/api/v1/engines

After startup:

  • Frontend UI: http://localhost:3020
  • API Docs: http://localhost:8020/docs
  • Health Check: http://localhost:8020/health

Media Clipping Tool

The project includes a media clipping tool for splitting long videos into shorter segments:

# Clip a specific duration (default 2 minutes)
python cli/clip.py input_video.mp4 120

# Auto-clip entire video into 2-minute segments
python cli/clip.py input_video.mp4

Python API usage:

from cli.clip import MediaClipper

clipper = MediaClipper("video.mp4")
clipper.clip(start=0, duration=120, output="clip.mp4")
clips = clipper.auto_clip(clip_duration=120, output_dir="clips/")

Model Download Guide (Hugging Face / ModelScope)

OneASR supports various local ASR engines (Qwen3-ASR, Faster-Whisper, FireRedASR, etc.). Models are stored by default under the models/ directory in OneASR/. You can download models using either Hugging Face CLI or ModelScope.

1. Install Download Tools

# Option A: Hugging Face CLI (provides `hf` command)
pip install -U "huggingface_hub[cli]"

# Optional: China mirror acceleration
export HF_ENDPOINT=https://hf-mirror.com

# Option B: ModelScope CLI (recommended for users in China)
pip install -U modelscope

2. Model Download Commands

Note: Run these commands from the OneASR/ directory to save models directly to models/{model_dir}.

① Qwen3-ASR (Main ASR Model)

Recommended: Qwen/Qwen3-ASR-1.7B (High quality) or Qwen/Qwen3-ASR-0.6B (Lightweight)

  • Hugging Face (hf):
    hf download Qwen/Qwen3-ASR-1.7B --local-dir models/Qwen3-ASR-1.7B
  • ModelScope CLI:
    modelscope download --model Qwen/Qwen3-ASR-1.7B --local_dir models/Qwen3-ASR-1.7B

② Qwen3-ForcedAligner (Timestamp Forced Alignment Model)

Used for word/character-level timestamps: Qwen/Qwen3-ForcedAligner-0.6B

  • Hugging Face (hf):
    hf download Qwen/Qwen3-ForcedAligner-0.6B --local-dir models/Qwen3-ForcedAligner-0.6B
  • ModelScope CLI:
    modelscope download --model Qwen/Qwen3-ForcedAligner-0.6B --local_dir models/Qwen3-ForcedAligner-0.6B

③ Faster-Whisper

Supported: Systran/faster-whisper-medium, large-v3, small, etc.

  • Hugging Face (hf):
    hf download Systran/faster-whisper-medium --local-dir models/faster-whisper-medium
  • ModelScope CLI:
    modelscope download --model Systran/faster-whisper-medium --local_dir models/faster-whisper-medium

④ FireRedASR Models

Supported: FireRedTeam/FireRedASR-AED-L, FireRedTeam/FireRedASR-LLM-L

  • Hugging Face (hf):
    hf download FireRedTeam/FireRedASR-AED-L --local-dir models/FireRedASR-AED-L
  • ModelScope CLI:
    modelscope download --model FireRedTeam/FireRedASR-AED-L --local_dir models/FireRedASR-AED-L

⑤ X-ASR (Real-time Streaming Model, sherpa-onnx based)

Model: GilgameshWind/X-ASR-zh-en — Chinese/English streaming recognition using zipformer2 transducer

  • Hugging Face (hf):

    hf download GilgameshWind/X-ASR-zh-en \
      --include "deployment/models/chunk-160ms-model/*" \
      --local-dir models
    mv models/deployment/models/chunk-160ms-model models/chunk-160ms-model
  • ModelScope CLI:

    modelscope download --model Gilgamesh-J/X-ASR-zh-en \
      --include "deployment/models/chunk-160ms-model/*" \
      --local_dir models
    mv models/deployment/models/chunk-160ms-model models/chunk-160ms-model

3. Batch Download via Python Script (Optional)

# download_models.py
from modelscope import snapshot_download

MODELS = {
    "models/Qwen3-ASR-1.7B": "Qwen/Qwen3-ASR-1.7B",
    "models/Qwen3-ForcedAligner-0.6B": "Qwen/Qwen3-ForcedAligner-0.6B",
    "models/faster-whisper-medium": "Systran/faster-whisper-medium",
}

for local_path, model_id in MODELS.items():
    print(f"Downloading {model_id} to {local_path} via ModelScope...")
    snapshot_download(model_id, local_dir=local_path)
print("All models downloaded successfully!")

4. Enable Models in config.yaml

After downloading, configure the corresponding provider in OneASR/config.yaml:

ASR-Providers:
  # Qwen3-ASR Configuration
  qwen:
    enable: true
    engine: qwen
    load:
      model_name: Qwen/Qwen3-ASR-1.7B
      model_path: models/Qwen3-ASR-1.7B
      device: cpu  # or cuda:0 / mps
      dtype: float32
      max_new_tokens: 256
      max_inference_batch_size: 32
      # Optional: forced aligner
      # forced_aligner_name: Qwen/Qwen3-ForcedAligner-0.6B
      # forced_aligner_path: models/Qwen3-ForcedAligner-0.6B

  # Faster-Whisper Configuration
  faster-whisper:
    enable: false
    engine: faster-whisper
    load:
      model_name: medium
      model_path: models/faster-whisper-medium
      device: cpu
      compute_type: int8

Authentication

All API endpoints (except /health) require an API Key via the X-API-Key request header.

Configure the API Key in config.yaml:

api_key: oneasr-key
# Call API with API Key
curl -H "X-API-Key: oneasr-key" http://localhost:8020/api/v1/engines

# Upload file for recognition
curl -X POST -H "X-API-Key: oneasr-key" \
  http://localhost:8020/api/v1/transcribe/file \
  -F "file=@audio.mp3" -F "format=srt" -o subtitle.srt

WebSocket streaming uses query parameters (browser WebSocket API doesn't support custom headers): ws://localhost:8020/ws/transcribe/stream?api_key=oneasr-key

API Endpoints

Unified API (OpenAI-compatible)

Endpoint Method Description
/api/v1/audio/transcriptions POST Create transcription (file upload or file_uuid)
/api/v1/audio/transcriptions/stream POST Create streaming transcription (SSE)
/api/v1/audio/models GET List available models
/api/v1/files/upload POST Upload file (supports MD5 instant upload)
/api/v1/files/list GET List all uploaded files
/api/v1/files/{file_id} GET Get file info
/api/v1/files/{file_id} DELETE Delete uploaded file
/health GET Health check

WebSocket Streaming

Endpoint Method Description
/v1/realtime?api_key= WebSocket Real-time streaming (OpenAI Realtime Transcription 协议)
/ws/transcribe/stream?api_key= WebSocket Real-time streaming (Legacy WhisperLiveKit 协议)

File Upload & Instant Upload (MD5 Dedup)

Upload files with optional MD5 fingerprint for instant upload of duplicate files:

# First upload: file is saved and metadata stored in DB
curl -X POST -H "X-API-Key: oneasr-key" \
  "http://localhost:8020/api/v1/files/upload?file_md5=abc123&file_size=1024000" \
  -F "file=@audio.mp3"
# Response: {"duplicate": false, "file_id": "uuid-1", ...}

# Second upload with same file: instant return (no re-upload)
curl -X POST -H "X-API-Key: oneasr-key" \
  "http://localhost:8020/api/v1/files/upload?file_md5=abc123&file_size=1024000" \
  -F "file=@audio_copy.mp3"
# Response: {"duplicate": true, "file_id": "uuid-1", ...}

Streaming Recognition

Real-time speech recognition powered by WhisperLiveKit:

  • VAD/VAC — Silero Voice Activity Detection, automatic speech/silence boundary detection
  • SimulStreaming — Low-latency streaming strategy (default), or LocalAgreement high-accuracy strategy
  • Timestamp Alignment — Each line has precise start/end timestamps
  • Speaker Diarization — Optional speaker identification (requires additional models)
  • Multi-language — Automatic language detection or manual specification

WebSocket Connection (OpenAI Realtime Transcription 协议)

const ws = new WebSocket("ws://localhost:8020/v1/realtime?api_key=oneasr-key");

// 1. 配置会话
ws.send(JSON.stringify({
  type: "session.update",
  session: {
    type: "transcription",
    audio: {
      input: {
        format: { type: "audio/pcm", rate: 16000 },
        transcription: { model: "wlk-live", language: "zh" },
      },
    },
  },
}));

ws.onmessage = (event) => {
  const data = JSON.parse(event.data);
  if (data.type === "session.updated") {
    console.log("Session configured:", data.session.id);
  } else if (data.type === "conversation.item.input_audio_transcription.delta") {
    process.stdout.write(data.delta);  // 增量文本
  } else if (data.type === "conversation.item.input_audio_transcription.completed") {
    console.log("\nFinal:", data.transcript);  // 完整文本
  }
};

// 2. 发送 base64 编码的 PCM 音频
ws.send(JSON.stringify({
  type: "input_audio_buffer.append",
  audio: btoa(String.fromCharCode(...pcmBytes)),
}));

// 3. 提交音频缓冲区
ws.send(JSON.stringify({ type: "input_audio_buffer.commit" }));

Stream Simulation Client

CLI 工具,将本地音视频文件模拟为麦克风实时输入:

# 基本用法
python -m cli.stream_simulation_client test.wav

# 指定中文和引擎
python -m cli.stream_simulation_client test.wav --language zh --model wlk-live

# 10 倍速快速测试
python -m cli.stream_simulation_client test.wav --speed 10

# 连接远程服务
python -m cli.stream_simulation_client test.wav --url ws://remote:8020/v1/realtime

Running Tests

# Run all tests
uv run pytest tests/ -v

# Run API tests only (fast, no external dependencies)
uv run pytest tests/api/ -v

# Run general/unit tests only
uv run pytest tests/general/ -v

# Run a specific test file
uv run pytest tests/api/test_file_upload.py -v

# Skip integration tests (need real audio files and running services)
uv run pytest tests/ -m "not integration" -v

# Run with short traceback
uv run pytest tests/api/ -v --tb=short

Test Structure

tests/
├── conftest.py                  # Shared fixtures (client, DB cleanup)
├── api/                         # API endpoint tests (TestClient)
│   ├── test_audio.py            # Audio transcription API
│   ├── test_file_upload.py      # File upload, MD5 dedup, CRUD
│   ├── test_models.py           # Models listing, format params
│   └── test_streaming.py        # SSE streaming transcription
└── general/                     # Unit tests & integration tests
    ├── test_format.py           # Output format conversion (SRT/VTT/JSON/TSV)
    ├── test_engines.py          # Engine loading and transcription
    ├── test_mimo.py             # MiMo audio understanding engine
    ├── test_stream_websocket.py # WebSocket streaming integration
    ├── test_stream_realtime.py  # OpenAI Realtime Transcription protocol tests
    └── test_stream_whisperlivekit.py   # WhisperLiveKit engine tests

Configuration

Edit config.yaml to configure API Key and engines:

api_key: oneasr-key

default_provider: whisper1

providers:
  whisper1:
    engine: faster-whisper
    type: local
    model_name: medium
    device: cpu
    compute_type: int8
  wlk-live:
    engine: whisperlivekit
    type: local
    model_name: base
    device: cpu
    compute_type: int8
    backend: auto
    backend_policy: simulstreaming
    language: auto
    vac: true
    diarization: false
    pcm_input: true

Documentation Navigation

Detailed designs and guides are located in the docs/ directory:

  • API Design & Reference
    • API Design — Architecture, audio chunking, and streaming design
    • API Reference — Endpoint specifications, parameters, and examples
  • Architecture & Toolkits
  • Guides & Best Practices
    • Model Guide — FireRedASR, SenseVoice, Whisper, and cloud engines
    • uv Guide — Fast dependency and environment management
    • Development Guide — Code style and testing conventions
  • References

Project Structure

OneASR/
├── docs/                         # Documentation center
│   ├── api/                      # API design and reference
│   ├── architecture/             # Architectural specifications
│   ├── guides/                   # Model and environment guides
│   ├── references/               # Upstream third-party references
│   └── assets/                   # Icons and visual assets
├── app/                          # Core backend application
│   ├── main.py                   # FastAPI entry point & lifespan
│   ├── api/                      # Routing layer (controllers)
│   │   ├── auth.py               # API Key authentication
│   │   ├── audio.py              # OpenAI-compatible audio API
│   │   ├── file_upload.py        # File upload, MD5 dedup, & URL import
│   │   ├── file_transcription.py # Async file transcription & SSE stream
│   │   ├── media.py              # Media download & parse API
│   │   ├── model.py              # Models listing API
│   │   ├── provider.py           # Providers listing API
│   │   ├── realtime.py           # Realtime WebSocket transcription
│   │   └── realtime_ext.py       # Extended realtime WebSocket API
│   ├── core/                     # Core configs, errors, & storage
│   │   ├── config.py             # YAML config management
│   │   ├── errors.py             # Unified exceptions
│   │   └── file_storage.py       # File storage utilities
│   ├── db/                       # Database session & engine
│   │   ├── base.py               # SQLAlchemy declarative base
│   │   └── session.py            # Async engine & sessionmaker
│   ├── engines/                  # ASR engine adapters
│   │   ├── base.py               # Abstract engine base class
│   │   ├── whisper_engine.py     # faster-whisper adapter
│   │   ├── firered_engine.py     # FireRedASR adapter
│   │   ├── qwen_engine.py        # Qwen ASR adapter
│   │   ├── xasr_engine.py        # Sherpa / SenseVoice adapter
│   │   ├── openai_engine.py      # OpenAI-compatible engine
│   │   ├── mimo_engine.py        # Xiaomi MiMo adapter
│   │   └── registry.py           # Engine registry & factory
│   ├── models/                   # SQLAlchemy ORM models
│   │   └── orm_models.py         # Database entity tables
│   ├── schemas/                  # Pydantic request/response schemas (DTOs)
│   │   ├── audio.py              # Audio transcription schemas
│   │   ├── file.py               # File asset & URL import schemas
│   │   ├── media.py              # Media parse & task schemas
│   │   └── transcription.py      # Async transcription task schemas
│   ├── services/                 # Business logic service layer
│   │   ├── file_service.py       # File transcription task service
│   │   ├── media_service.py      # Media download & state machine
│   │   └── record_service.py     # Transcription records service
│   └── utils/                    # Audio & processing utilities
│       ├── audio.py              # Format conversion & duration probe
│       ├── asr_toolkit.py        # VAD segmentation & streaming chunks
│       ├── audio_converter.py    # Resampling & channel conversion
│       ├── download.py           # URL download utility
│       ├── format.py             # Output formatting (SRT/VTT/JSON/TSV)
│       ├── stream.py             # PCM streaming utilities
│       ├── vad.py                # Silero VAD integration
│       └── video_url.py          # yt-dlp & direct URL extractor
├── cli/                          # CLI & client test tools
│   ├── clip.py                   # Media clipping tool
│   ├── converter.py              # Audio converter
│   └── stream_simulation_client.py # Simulation client
├── web/                          # Vue 3 frontend web UI
├── tests/                        # Automated test suites
├── models/                       # Model storage directory
├── data/                         # SQLite database directory
├── uploads/                      # Uploaded files storage
├── download/                     # Media download temporary directory
├── config.yaml                   # Engine & server configuration
├── pyproject.toml                # Project metadata & dependencies
└── requirements.txt              # Pinned requirements

License

This project is licensed under the Apache 2.0 License - see the LICENSE file for details.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages