Skip to content

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

PotPlayer-FunASR-Plugin

为 PotPlayer 提供本地语音识别字幕的插件,让无字幕视频也能获得字幕。 基于阿里达摩院 SenseVoice Small 多语言模型, 通过 FunASR llama.cpp runtime 在 CPU 上推理,无需 Python、无需 GPU、无需编译。

特性

  • 极小体积:总部署约 247 MB(二进制 4.7 MB + ASR 模型 242 MB + VAD 模型 1.6 MB + main.exe 9.5 KB)
  • 零依赖:自带预编译 Windows x64 二进制,不需要 Python / CUDA / PyTorch / .NET SDK
  • 多语言:SenseVoice Small 支持中、英、日、韩、粤语自动检测
  • CPU 推理:基于 ggml / llama.cpp,AVX2 加速;i5-12400 上 RTF ≈ 0.045(22 倍实时)
  • 内置 VAD:FSMN-VAD 自动分段,并输出带时间戳的 SRT 字幕
  • 无缝替换:部署时替换 PotPlayer 的 Const-me 引擎,用户在引擎下拉框选择「Const-me」即可使用

工作原理

PotPlayer 的「声音生成字幕」功能内置了固定的引擎列表(BLAS / CPU / CUDA / Const-me / Faster-Whisper-XX),不会自动扫描新目录。因此本插件的做法是替换 Const-me 引擎目录下的 main.exe,让 PotPlayer 在用户选择「Const-me」引擎时实际调用 FunASR: 注:由于potplayer硬编码的限制,Const-me必须要下载一个模型才能启动,一般下载最小的tiny模型就行,实际上不会调用该模型,而是使用funASR进行识别。

Module\Whisper\Const-me\
├── main.exe                       # ← 本插件的 C# 入口(替换原始 Const-me main.exe)
├── main.exe.constme_orig          # ← 原始 main.exe 备份
├── Whisper.dll.constme_disabled   # ← 原始 Whisper.dll 重命名禁用
├── Whisper.dll.constme_orig       # ← 原始 Whisper.dll 备份
└── bin\
    ├── llama-funasr-sensevoice.exe  # SenseVoice 推理二进制 (v0.1.4 AVX2)
    ├── llama-funasr-vad.exe          # FSMN-VAD 二进制(输出分段时间戳)
    ├── sensevoice-small-q8.gguf      # SenseVoice Small Q8 量化模型 (242 MB)
    └── fsmn-vad.gguf                 # FSMN VAD 模型 (1.6 MB)

main.exe 接收 PotPlayer 传入的 whisper.cpp 标准 CLI 参数(-m -f -l -osrt 等),内部:

  1. 调用 llama-funasr-vad.exe 获取每段语音的起止时间戳(毫秒)
  2. 调用 llama-funasr-sensevoice.exe --vad 获取每段的识别文本
  3. 将时间戳与文本合并为 SRT 格式输出到 stdout(PotPlayer 读取 stdout 作为字幕)

项目结构

Potplayer-FunASR-Plugin/
├── bin/                            # 预编译二进制 + GGUF 模型
│   ├── llama-funasr-sensevoice.exe # SenseVoice 推理二进制 (v0.1.4 AVX2)
│   ├── llama-funasr-vad.exe        # FSMN-VAD 二进制(输出分段时间戳)
│   ├── sensevoice-small-q8.gguf   # SenseVoice Small Q8 量化模型 (242 MB)
│   └── fsmn-vad.gguf               # FSMN VAD 模型 (1.6 MB)
├── run/
│   ├── main.exe                    # C# 编译的引擎入口(部署到 PotPlayer)
│   └── main.bat                    # 旧版 .bat 入口(保留备用)
├── scripts/
│   ├── download.ps1                # 下载二进制 + 模型
│   ├── compile.ps1                 # 编译 main.cs -> main.exe(用 Windows 自带 csc.exe)
│   ├── deploy.ps1                  # 部署到 PotPlayer\Module\Whisper\FunASR\
│   └── install_all.ps1             # 一键安装:下载 -> 编译 -> 部署
├── src/
│   └── main.cs                     # C# 源码(引擎入口逻辑)
├── tests/
│   ├── zh.mp3                      # 中文测试音频
│   └── en.mp3                      # 英文测试音频
└── README.md

快速开始

模型文件下载

由于 GGUF 模型文件超过 GitHub 单文件 100MB 限制,仓库不包含模型文件,需手动下载。

方式一:用脚本自动下载(推荐)

powershell.exe -NoProfile -ExecutionPolicy Bypass -File scripts\download.ps1

脚本会自动下载以下文件到 bin\ 目录:

文件 大小 下载地址
sensevoice-small-q8.gguf ~242 MB HuggingFace: SenseVoiceSmall-GGUF
fsmn-vad.gguf ~1.6 MB HuggingFace: fsmn-vad-GGUF
llama-funasr-sensevoice.exe ~1.5 MB GitHub Release: runtime-llamacpp-v0.1.4

方式二:手动下载

从上表链接下载对应文件,放到 bin\ 目录下即可。如需其他量化版本或模型(如 Paraformer),访问 FunAudioLLM HuggingFace 主页。

前置条件

  • Windows 10/11 x64
  • CPU 支持 AVX2(2013 年后的 x86-64 CPU 基本都支持)
  • PowerShell 5.1+(Windows 自带)
  • PotPlayer 64 位

一键安装

powershell.exe -NoProfile -ExecutionPolicy Bypass -File scripts\install_all.ps1

脚本会自动完成:

  1. 下载 llama-funasr-sensevoice.exe (v0.1.4 AVX2)、llama-funasr-vad.exe、sensevoice-small-q8.gguf、fsmn-vad.gguf
  2. 用 Windows 自带的 csc.exe 编译 src\main.cs -> run\main.exe
  3. 把 main.exe + bin\ 复制到 <PotPlayer>\Module\Whisper\Const-me\(自动备份原有 Const-me 文件)

分步安装

如果一键安装失败,可以分步执行:

# 1. 下载二进制和模型
powershell.exe -NoProfile -ExecutionPolicy Bypass -File scripts\download.ps1

# 2. 编译 main.exe
powershell.exe -NoProfile -ExecutionPolicy Bypass -File scripts\compile.ps1

# 3. 部署到 PotPlayer
powershell.exe -NoProfile -ExecutionPolicy Bypass -File scripts\deploy.ps1 -Force

在 PotPlayer 中使用

  1. 重启 PotPlayer(部署后必须重启,让 PotPlayer 加载新的 main.exe)
  2. 按 F5 打开「选项」
  3. 左侧选择「字幕」→「声音生成字幕」
  4. 「Whisper AI 引擎」下拉框中选择 Const-me
  5. 「Whisper AI 语言」选择 auto(自动检测)或指定语种
  6. 点击「确定」保存
  7. 播放任意带音轨的无字幕视频,PotPlayer 会自动调用 main.exe 生成字幕

命令行手动测试

# 测试部署后的 main.exe
& "A:\PotPlayer_20181126\PotPlayer\Module\Whisper\Const-me\main.exe" -l auto "C:\path\to\audio.wav"

# 测试仓库中的 main.exe
& ".\run\main.exe" -l auto ".\tests\zh.mp3"

输出示例(SRT 格式到 stdout):

1
00:00:00,410 --> 00:00:05,600
开放时间早上9点至下午5点。

命令行参数

main.exe 兼容 whisper.cpp 标准 CLI(Const-me/Whisper 接口),支持以下参数:

参数 说明 默认值
-l LANG 语种:auto/zh/en/yue/ja/ko auto
-f FILE 输入音频路径(也可作为位置参数直接传)
-m MODEL 覆盖默认模型路径 bin\sensevoice-small-q8.gguf
-t N CPU 线程数(保留兼容,runtime 内部决定) 4
-osrt 同时输出 .srt 文件到音频同目录 off
-otxt 同时输出 .txt 文件 off
-ovtt 同时输出 .vtt 文件 off
-nt 不打印时间戳(纯文本输出) off
-ot DIR 输出目录 音频同目录

未识别的参数会被忽略(--device、--gpu-mem 等 PotPlayer 传入的参数)。

性能参考

在 12th Gen Intel Core i5-12400(6 核 12 线程,AVX2)上测试:

音频 时长 VAD 段数 处理时间 RTF
zh.mp3(中文) 5.6s 1 0.28s 0.05
en.mp3(英文) 7.2s 1 0.35s 0.05

已知限制

  • 替换 Const-me 引擎:本插件通过替换 PotPlayer 内置的 Const-me 引擎工作,不是作为独立引擎注册。如需恢复原始 Const-me/Whisper,将 main.exe.constme_orig 重命名回 main.exe、Whisper.dll.constme_orig 重命名回 Whisper.dll 即可。
  • 非流式:当前实现为整段推理,不是边播放边出字幕。PotPlayer 会等待 main.exe 输出完成后再显示字幕。
  • v0.1.0 已知崩溃:早期 v0.1.0 预编译二进制在部分 CPU 上崩溃(STATUS_ILLEGAL_INSTRUCTION),已在 v0.1.4 修复。如遇崩溃请确认 bin\ 下是 v0.1.4 版本。

开发说明

修改 main.cs 后重新部署

# 1. 重新编译
powershell.exe -NoProfile -ExecutionPolicy Bypass -File scripts\compile.ps1

# 2. 重新部署(会自动检测 main.cs 比 main.exe 新,触发重新编译)
powershell.exe -NoProfile -ExecutionPolicy Bypass -File scripts\deploy.ps1 -Force

# 3. 重启 PotPlayer

切换模型

修改 bin\ 下的模型文件,或用 -m 参数指定其他模型路径。例如改用 Paraformer:

# 下载 paraformer GGUF 模型到 bin\
# 然后在 main.exe 调用时指定:
& ".\run\main.exe" -m ".\bin\paraformer-q8.gguf" -l auto "audio.wav"

许可证

见 LICENSE。

致谢

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages