FunAudioLLM/Fun-ASR-Nano-GGUF overview
Fun ASR Nano · GGUF FunASR llama.cpp runtime GGUF build of Fun ASR Nano SenseVoice SAN M encoder + adaptor + Qwen3 0.6B LLM decoder for the zero Python, CPU/ed…
Runs locally from ~447.6 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | FunAudioLLM/Fun-ASR-Nano-GGUF |
|---|---|
| Author | FunAudioLLM |
| Pipeline | automatic-speech-recognition |
| License | apache-2.0 |
| Base model | — |
| Last modified | 2026-07-26T18:04:00.000Z |
Model README
---
license: apache-2.0
language:
- zh
- en
library_name: gguf
tags:
- automatic-speech-recognition
- asr
- fun-asr
- funasr
- qwen3
- llama.cpp
- ggml
- cpu
- chinese
- arxiv:2407.04051
pipeline_tag: automatic-speech-recognition
---
Fun-ASR-Nano · GGUF (FunASR llama.cpp runtime)
GGUF build of Fun-ASR-Nano (SenseVoice SAN-M encoder + adaptor + Qwen3-0.6B LLM decoder) for the zero-Python, CPU/edge FunASR llama.cpp runtime — the accuracy leader (LLM decoder), single C++ binary.
LLM quantization (pick by size vs accuracy)
The Fun-ASR-Nano LLM (Qwen3-0.6B) ships in three tiers — all within 0.1% CER (184-file micro-CER). Pair any with funasr-encoder-f16.gguf (470 MB).
| LLM file | size | CER ↓ | speed |
|---|---|---|---|
| qwen3-0.6b-q4km.gguf | 484 MB | 8.35% | 6.1× | smallest |
| qwen3-0.6b-q5km.gguf | 551 MB | 8.25% | 5.7× | best accuracy |
| qwen3-0.6b-q8_0.gguf | 805 MB | 8.30% | 6.0× | |
Recommended: q4_K_M (smallest) or q5_K_M (best).
Get it running (no Python, no build)
These are GGUF weights for the FunASR llama.cpp runtime — a whisper.cpp-style, single self-contained binary for CPU / edge. Grab a prebuilt binary, then fetch this model and run:
- Prebuilt binaries (Linux / macOS / Windows) → GitHub Releases (tag
runtime-llamacpp-v*) - Deployment guide & qualified benchmarks → funasr.com/deploy/llama-cpp
bash download-funasr-model.sh nano ./gguf
llama-funasr-cli --enc ./gguf/funasr-encoder-f16.gguf -m ./gguf/qwen3-0.6b-q8_0.gguf --vad ./gguf/fsmn-vad.gguf -a audio.wav
Files
| file | size | notes |
|---|---|---|
| funasr-encoder-f16.gguf | 470 MB | audio encoder + adaptor (f16) |
| qwen3-0.6b-q8_0.gguf | 805 MB | LLM decoder, recommended (Q8_0) |
| qwen3-0.6b-q4km.gguf | 484 MB | LLM decoder, smaller (Q4_K_M) |
Usage (needs both the encoder and the LLM gguf)
llama-funasr-cli --enc funasr-encoder-f16.gguf -m qwen3-0.6b-q8_0.gguf -a audio.wav --vad fsmn-vad.gguf
On CPU: 8.30 % CER on the 184-clip Mandarin benchmark (vs whisper.cpp 22–31 %).
Links
- 🧩 Runtime & build: Fun-ASR · runtime/llama.cpp — ⭐ Star Fun-ASR!
- Source model: FunAudioLLM/Fun-ASR-Nano-2512
Run FunAudioLLM/Fun-ASR-Nano-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models