跳到主要内容
Supermarket
返回能力市场
MCP Server
programming
MIT

FunASR

Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.

modelscopemodelscope
91/ 100

公开评测 · 综合采用结论

证据充分,整体质量与安全表现优秀

查看评测依据 评测我的项目基于公开项目证据,非安全认证或安装推荐
20.6kstars
2.1kforks
最近更新 4小时前
评测生成时间(北京时间)
本报告引擎
v3.9.1
当前引擎
v3.16.0

本报告与当前引擎使用不同规则;原分数不会自动更新,不同版本的分数不宜直接对比。

重新评测此项目

进入后确认来源与额度,提交才会创建任务。

Evaluation report

综合采用结论

91
A
满分 100
值得推荐低风险
决策摘要

证据充分,整体质量与安全表现优秀

84%
高置信度
100
文档
100
安全
80
质量
100
活跃
70
采用
  • 基础评测完成+25/25确定性评分与静态安全扫描已完成
  • README 有效证据+25/2516,924 个去重后的有效字符
  • 独立证据来源+4/201 类非重复证据,重复文件不叠加
  • 仓库元数据+10/10已取得仓库状态与采用数据
  • 活跃记录+5/5已取得最近提交时间
  • AI 复核+15/15已完成结构化 AI 证据复核
How it works · 流程图

FunASR 语音识别流水线

README 展示了从输入音频到输出文本的连续处理步骤,包括VAD、ASR、标点和说话人分离。

AI 提取 · 证据约束

左右滑动查看完整图示

FunASR 语音识别流水线README 展示了从输入音频到输出文本的连续处理步骤,包括VAD、ASR、标点和说话人分离。音频流语音段文本结构化结果输入音频VAD 检测ASR 识别后处理输出结果
图示依据
  • • AutoModel 流水线调用配置的 ASR、VAD 和说话人模型并返回组合结果
  • • 示例输出显示带说话人标签和时间戳的文本
  • • 流式示例按块处理音频并输出文本
五维表现
FunASR 作为开源语音识别工具包,解决真实问题,提供多种模型和部署方式,示例丰富,文档详尽。最大缺口是模型许可证和具体错误处理细节需进一步明确。
质量证据
  • Quick Start 部分提供安装命令和示例代码
  • Model Zoo 表格列出模型、任务、语言和参数
  • Deploy 部分提供API服务器和Docker命令
  • Benchmark 部分提供CER和速度对比
  • License 部分说明工具包MIT,模型单独许可
采用建议
优势
  • 问题与用途描述
  • 有效 README
  • 安装或接入步骤
  • 可执行示例
  • 未发现已知高风险模式
关注点
  • 模型许可证需单独查看,未在README中明确
  • 错误处理和排障信息分散,需查阅专门文档
  • 部分高级功能(如vLLM)依赖额外安装,未详细说明
  • 流式处理示例较复杂,需更多参数说明
适合

需要离线或流式语音识别的开发者、需要多语言和说话人分离的场景、希望自托管语音识别服务的团队、需要边缘设备部署的用户

不建议直接用于

需要完全无GPU且要求极低延迟的场景(但CPU可行)、需要统一许可证的商用项目(模型许可证各异)

也有自己的公开项目?先看完证据,再用当前规则生成独立报告。

评测我的项目 →
文档证据
100/100
问题与用途描述10 分
有效 README12 分
安装或接入步骤14 分
可执行示例16 分
输入、参数或工具说明11 分
输出或结果说明9 分
限制、权限或边界12 分
错误处理或排障8 分
许可证信息5 分
结构化章节3 分
安全证据
低风险
未发现已知高风险模式

静态扫描不是安全保证,生产接入前仍应人工复核权限和数据边界。

方法、证据与局限展开
数据来源

GitHub Repository API

扫描范围

1 个文件 · 20,965 字符

评测引擎

v3.9.1 · AI 复核已启用(deepseek-chat)

局限
  • 静态评测不会安装或执行项目代码
  • 安全扫描基于高信号文件与已知模式,不能替代人工审计
  • 流行度只反映采用程度,不代表安全或工程质量

30 天热度趋势

README

(简体中文|English|日本語|한국어)

FunASR

Industrial speech recognition toolkit for offline, streaming, and edge deployment.
ASR · VAD · punctuation · speaker pipelines · emotion and audio-event models · OpenAI-compatible serving

PyPI Stars Downloads Docs MCP Toplist

modelscope%2FFunASR | Trendshift

Quick Start · Model selection · Models · Deployment matrix · Deployment hub · Docs · Benchmark · Contribute


Quick Start

Native Transformers

For Fun-ASR-Nano transcription with the Hugging Face API, start with the Transformers 5.17.0 CPU quickstart. No FunASR toolkit or remote Python code is needed.

Space · Notebook · Python / batch examples

FunASR toolkit and pipelines

Open In Colab

No local setup? Open the Colab quickstart to transcribe a public sample or upload your own audio in a browser.

Found FunASR useful? Star the project so more builders can find it.

# CPU-only installs can use the default PyPI wheels.
pip install torch torchaudio
pip install funasr

For GPU quickstarts, install the PyTorch and torchaudio wheels that match your NVIDIA driver from pytorch.org before installing FunASR. After installation, confirm the GPU is visible:

python - <<'PY'
import torch
print(torch.cuda.is_available())
PY

Only use device="cuda" when this prints True; otherwise use device="cpu" or reinstall PyTorch with the correct CUDA wheel.

FunASR toolkit GPU example: Fun-ASR-Nano (Chinese, English, Japanese, and Chinese dialect groups and regional accents; the separate native Transformers CPU path is linked above):

from funasr import AutoModel

model = AutoModel(model="FunAudioLLM/Fun-ASR-Nano-2512", device="cuda")
result = model.generate(input="https://isv-data.oss-cn-hangzhou.aliyuncs.com/ics/MaaS/ASR/test_audio/asr_example_zh.wav")
print(result[0]["text"])

For the separate 31-language checkpoint, use Fun-ASR-MLT-Nano-2512. Language coverage is checkpoint-specific, so Nano and MLT-Nano should be treated as distinct model choices.

For a CPU-first example with five-language ASR plus emotion and audio-event tags, use SenseVoiceSmall. The pipeline below combines it with FSMN-VAD and CAM++ for speaker-aware VAD segments; these are not native speaker outputs of the SenseVoiceSmall checkpoint. See the SenseVoice paper, Hugging Face checkpoint, and GGUF edge checkpoint.

from funasr import AutoModel
from funasr.utils.postprocess_utils import rich_transcription_postprocess

model = AutoModel(model="iic/SenseVoiceSmall", vad_model="fsmn-vad", spk_model="cam++", device="cpu")
result = model.generate(
    input="https://isv-data.oss-cn-hangzhou.aliyuncs.com/ics/MaaS/ASR/test_audio/asr_example_zh.wav",
    batch_size_s=300,
)

# The AutoModel pipeline returns VAD segments with speaker ids and timestamps:
for seg in result[0]["sentence_info"]:
    print(f"[{seg['start']/1000:.1f}s] Speaker {seg['spk']}: {rich_transcription_postprocess(seg['sentence'])}")

This prints each returned segment's start time in seconds, anonymous speaker index, and text with SenseVoice tags removed. Text and segment boundaries depend on the audio and checkpoint; no fixed transcript is asserted here.

CAM++ extracts spk_embedding vectors. AutoModel clusters those embeddings and assigns speaker indices to VAD segments. Indices are local to a recording, not known-person identities. See the SDK contract for the component and result boundaries. Change to device="cuda" only after verifying a compatible GPU environment as described above.

Scale & deploy the flagship

At scale, accelerate Fun-ASR-Nano with vLLM (batch processing):

from funasr.auto.auto_model_vllm import AutoModelVLLM

model = AutoModelVLLM(model="FunAudioLLM/Fun-ASR-Nano-2512", tensor_parallel_size=1)
results = model.generate(["audio1.wav", "audio2.wav"], language="auto")

Deploy as API server: Local SenseVoice CPU recipe · Nano GPU serving and pinned vLLM setup

Use with AI agents: MCP Server for Claude/Cursor · OpenAI API for LangChain/Dify/AutoGen

Use with voice agents: OpenClaw realtime plugin for self-hosted Talk and Voice Call transcription

Why FunASR?

FunASR is a toolkit: choose the task, checkpoint, and runtime separately. Support in one model or adapter does not imply support in every serving backend.

TaskCheckpoint or pipelineRuntime entrypointImportant limitation
File transcription with emotion/event tagsSenseVoiceSmallPython AutoModel, CPU or GPUFive-language checkpoint; tags do not identify speakers.
LLM-based file transcriptionFun-ASR-NanoAutoModel; split-engine AutoModelVLLM for the documented GPU pathBase Nano covers zh/en/ja and Chinese dialects/accents; timestamp support depends on checkpoint and path.
Broader multilingual transcriptionFun-ASR-MLT-NanoPython AutoModelSeparate 31-language checkpoint; do not transfer its coverage to base Nano.
Chunked live transcriptionParaformer-zh-streamingStreaming SDK or runtime WebSocket serviceUse the streaming checkpoint and per-session cache, not an offline checkpoint.
Speaker-aware file transcriptionSenseVoiceSmall + FSMN-VAD + CAM++AutoModel with VAD and embedding clusteringAnonymous indices within a recording, not enrolled-speaker identification.
Joint text, timestamps, and speakersMOSS-Transcribe-Diarize, third-party OpenMOSSFunASR adapter or upstream backend in the MOSS guideOffline, recording-local anonymous labels; no external VAD/speaker pipeline for its unified path.
Native CPU/edge transcriptionFun-ASR-Nano or SenseVoiceSmall GGUFllama.cpp runtimeRequires matching converted weights; GGUF is not a Python AutoModel checkpoint.

See the Model Zoo and deployment matrix for checkpoint, interface, and licensing boundaries. Benchmark on your own audio and hardware before choosing a runtime.

Trying FunASR for the first time? Use the Colab quickstart before setting up a local environment. Choosing a first model? Start with the model selection guide. Planning a switch from Whisper or a cloud ASR provider? Use the migration guide and benchmark example to test representative audio, map features, and roll out safely.


Installation

pip install funasr
From source / Requirements
git clone https://github.com/modelscope/FunASR.git && cd FunASR
pip install -e ./

Requirements: Python ≥ 3.8. Install PyTorch + torchaudio first (pytorch.org), then pip install funasr.


Model Zoo

This list includes third-party models. OpenMOSS publishes MOSS-Transcribe-Diarize; FunASR provides an adapter, not ownership of its weights. Its unified path is offline, with anonymous labels scoped to each recording, not realtime or known-person identification. Model licenses are separate from the toolkit's MIT license.

ModelTaskLanguagesParamsLinks
Fun-ASR-NanoASRzh/en/ja + Chinese dialects and accents800M⭐ HF / Transformers · HF / FunASR GGUF
Fun-ASR-MLT-NanoASR31 languages800M⭐ 🤗
SenseVoiceSmallASR + emotion + eventszh/en/ja/ko/yue234M⭐ 🤗 GGUF paper
MOSS-Transcribe-DiarizeThird-party OpenMOSS: offline ASR + timestamps + anonymous speakersSee official cardSee official card🤗 guide
Paraformer-zhASR + timestampszh/en220M⭐ 🤗
Paraformer-zh-streamingStreaming ASRzh/en220M⭐ 🤗
Qwen3-ASRASR, 52 languagesmultilingual1.7Busage
GLM-ASR-NanoASR, 17 languagesmultilingual1.5Busage
Whisper-large-v3ASR + translationmultilingual1550Musage
Whisper-large-v3-turboASR + translationmultilingual809Musage
ct-puncPunctuationzh/en290M⭐ 🤗
fsmn-vadVADzh/en0.4M⭐ 🤗
cam++Speaker embeddings (pipeline component)—7.2M⭐ 🤗
emotion2vec+largeEmotion recognition—300M⭐ 🤗

Usage

Python tutorial · SDK parameters and outputs · Training · Model registration

from funasr import AutoModel

# Chinese production (VAD + ASR + punctuation + speaker)
model = AutoModel(model="paraformer-zh", vad_model="fsmn-vad", punc_model="ct-punc", spk_model="cam++", device="cuda")
result = model.generate(input="https://isv-data.oss-cn-hangzhou.aliyuncs.com/ics/MaaS/ASR/test_audio/asr_example_zh.wav", hotword="关键词 20")

# Optional Silero VAD (install first: python -m pip install "funasr[silero]")
model = AutoModel(
    model="paraformer-zh", vad_model="silero-vad", device="cuda",
    vad_kwargs={"silero_threshold": 0.5, "silero_min_silence_duration_ms": 100},
)
result = model.generate(input="audio.wav")

# Streaming real-time (feed audio chunk by chunk)
import soundfile as sf
model = AutoModel(model="paraformer-zh-streaming", device="cuda")
audio, sr = sf.read("speech.wav", dtype="float32")   # 16 kHz mono
chunk_size = [0, 10, 5]                               # 600 ms chunks
chunk_stride = chunk_size[1] * 960
cache = {}
n_chunks = (len(audio) - 1) // chunk_stride + 1
for i in range(n_chunks):
    chunk = audio[i * chunk_stride : (i + 1) * chunk_stride]
    res = model.generate(input=chunk, cache=cache, is_final=(i == n_chunks - 1),
                         chunk_size=chunk_size, encoder_chunk_look_back=4, decoder_chunk_look_back=1)
    if res[0]["text"]:
        print(res[0]["text"], end="", flush=True)

# Emotion recognition
model = AutoModel(model="emotion2vec_plus_large", device="cuda")
result = model.generate(input="audio.wav", granularity="utterance")

CLI (Agent-Friendly)

# Transcribe audio (simplest)
funasr audio.wav

# JSON output (for AI agents)
funasr audio.wav --output-format json

# SRT subtitles
funasr audio.wav --output-format srt --output-dir ./subs

# Speaker diarization + timestamps
funasr audio.wav --spk --timestamps -f json

# Choose model and language
funasr audio.wav --model paraformer --language zh

# Batch transcribe
funasr *.wav --output-format srt --output-dir ./output

Available models: sensevoice (default), paraformer, paraformer-en, fun-asr-nano


Deploy

Start a local SenseVoice CPU service from a fresh directory in a POSIX shell with Python 3.11. This installs the PyPI release into a separate environment, not this source checkout. Keep the unauthenticated service on loopback; use the security guide before exposing it to other clients.

python3.11 -m venv .venv-funasr-http
. .venv-funasr-http/bin/activate
python -m pip install torch torchaudio
python -m pip install funasr fastapi uvicorn python-multipart
python -m pip check
funasr-server --host 127.0.0.1 --port 8000 --model sensevoice --device cpu

Wait for model download and server startup. In a second terminal, use the same directory and curl 7.76+ to download a public Chinese audio sample and transcribe it. The request uses the preloaded model; no fixed transcript or speaker labels are promised.

curl --fail --location https://isv-data.oss-cn-hangzhou.aliyuncs.com/ics/MaaS/ASR/test_audio/BAC009S0764W0121.wav -o sample.wav && \
curl --fail-with-body http://127.0.0.1:8000/v1/audio/transcriptions \
  -F file=@sample.wav \
  -F model=sensevoice \
  -F response_format=verbose_json

For offline joint ASR and anonymous speaker labels (moss-transcribe-diarize), prepare the separate environment in the MOSS service, Docker, Kubernetes, vLLM, SGLang, LocalAI, and FunClip guide →. It is an alternative service, not another command in the CPU environment. Stop the CPU service before reusing port 8000. For Nano GPU serving, follow the pinned split-engine guide and inspect the actual backend logs; selecting a model does not by itself prove that vLLM was loaded.

# Docker streaming service
docker pull registry.cn-hangzhou.aliyuncs.com/funasr_repo/funasr:funasr-runtime-sdk-online-cpu-0.1.12

CPU / Edge — llama.cpp / GGUF (no GPU, no Python)

Run SenseVoice / Paraformer / Fun-ASR-Nano as a single self-contained binary on CPU and edge devices — this is to FunASR what whisper.cpp is to Whisper, but with ~3× lower CER than whisper.cpp on Chinese. Built-in FSMN-VAD, no Python at runtime.

# Linux / macOS: run from the extracted release directory
bash download-funasr-model.sh sensevoice ./gguf        # or: paraformer | nano
./llama-funasr-sensevoice -m ./gguf/sensevoice-small-q8.gguf --vad ./gguf/fsmn-vad.gguf -a audio.wav
# → 欢迎大家来体验达摩院推出的语音识别模型
# Windows PowerShell: run from the extracted archive root (with the `hf` CLI installed)
hf download FunAudioLLM/SenseVoiceSmall-GGUF sensevoice-small-q8.gguf --local-dir .\gguf
hf download FunAudioLLM/fsmn-vad-GGUF fsmn-vad.gguf --local-dir .\gguf
.\llama-funasr-sensevoice.exe -m .\gguf\sensevoice-small-q8.gguf --vad .\gguf\fsmn-vad.gguf -a audio.wav
# Use the windows-x64-vulkan package with a current AMD, Intel, or NVIDIA Vulkan driver:
.\llama-funasr-sensevoice.exe -m .\gguf\sensevoice-small-q8.gguf --vad .\gguf\fsmn-vad.gguf -a audio.wav --backend vulkan
# Use the windows-x64-cuda package on RTX 30-class GPUs:
.\llama-funasr-sensevoice.exe -m .\gguf\sensevoice-small-q8.gguf --vad .\gguf\fsmn-vad.gguf -a audio.wav --backend cuda

Use funasr-llamacpp-linux-x64-vulkan.tar.gz on Linux GPU systems with a working Vulkan driver/ICD:

./llama-funasr-sensevoice -m ./gguf/sensevoice-small-q8.gguf --vad ./gguf/fsmn-vad.gguf -a audio.wav --backend vulkan

The Windows Vulkan ZIP uses the system Vulkan loader supplied by the GPU driver; installing the Vulkan SDK is only necessary when building from source. Both Vulkan packages currently accelerate SenseVoiceSmall.

Tagged releases provide two Windows CUDA packages. The standard windows-x64-cuda ZIP targets CUDA architecture 86, while windows-x64-cuda-blackwell targets architecture 120 (sm_120) for RTX 50 / Blackwell GPUs. Both ZIPs bundle the required cuBLAS DLLs and use the static MSVC runtime, so users need a compatible NVIDIA driver but not a separate CUDA Toolkit installation. CI verifies the architecture and package boundary; it does not prove inference on physical Blackwell hardware.

Prebuilt binaries: Releases · v0.2.6 · Linux Vulkan tarball · Windows Vulkan zip · Windows CUDA zip · Windows Blackwell CUDA zip · Download & quickstart: funasr.com/deploy/llama-cpp · GGUF models: Hugging Face · Docs & benchmarks: runtime/llama.cpp/

OpenAI API example → · Gradio demo → · Client recipes → · JavaScript/TypeScript recipes → · Kubernetes template → · Workflow recipes → · Postman collection → · OpenAPI spec → · Security guide → · Deployment matrix → · Deployment docs → · Agent integration →


Benchmark

The historical benchmark report and split-engine measurements retain their original results. They are separate records, not universal speed rankings or production capacity guarantees.

Use the RTFx and reproducibility notes to compare checkpoint/revision, audio set, hardware, batching, warmup, timing scope, and CER/WER. Offline throughput is not streaming latency. The migration benchmark example helps measure your own recordings with the same evaluation scope.


What's new

  • MOSS-Transcribe-Diarize brings long-form ASR, timestamps, and anonymous speaker labels to FunASR services, Docker, Kubernetes, vLLM/SGLang workflows, and FunClip. Deploy MOSS ->
  • FunASR 1.4.16 preserves explicit VAD, punctuation, and speaker-model device placement across inference calls, and adds tested native Transformers adoption guides. Install with python -m pip install -U "funasr==1.4.16". Release and verification scope ->
  • Native Transformers: Released 5.17.0 supports Fun-ASR-Nano with the official -hf checkpoint, CPU examples and a notebook. Get started ->

See GitHub Releases for the complete changelog and downloadable assets.


Community

Start with troubleshooting before reporting a problem. Include your exact model, runtime, environment and a minimal reproduction.

📖 Documentation🐛 Issues
💬 Discussions🤗 HuggingFace
🤝 Contributing🌐 funasr.com
🗺️ Repository roles & roadmap📈 Growth plan
🧩 Community projects💡 Use-case showcase

Star History

Star History Chart

License

Citations

@inproceedings{gao2023funasr,
  author={Zhifu Gao and others},
  title={FunASR: A Fundamental End-to-End Speech Recognition Toolkit},
  booktitle={INTERSPEECH},
  year={2023}
}