handy-computer/moss-transcribe-diarize-gguf overview
MOSS Transcribe Diarize: transcribe.cpp GGUF GGUF conversions of OpenMOSS Team/MOSS Transcribe Diarize https://huggingface.co/OpenMOSS Team/MOSS Transcribe Dia…
Runs locally from ~588.7 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| MOSS-Transcribe-Diarize-BF16.gguf | GGUF | BF16 | 1.70 GB | Download |
| MOSS-Transcribe-Diarize-F16.gguf | GGUF | F16 | 1.71 GB | Download |
| MOSS-Transcribe-Diarize-Q4_K_M.gguf | GGUF | Q4_K_M | 588.7 MB | Download |
| MOSS-Transcribe-Diarize-Q5_K_M.gguf | GGUF | Q5_K_M | 667.9 MB | Download |
| MOSS-Transcribe-Diarize-Q6_K.gguf | GGUF | Q6_K | 732.6 MB | Download |
| MOSS-Transcribe-Diarize-Q8_0.gguf | GGUF | Q8_0 | 941.2 MB | Download |
Model Details
| Model ID | handy-computer/moss-transcribe-diarize-gguf |
|---|---|
| Author | handy-computer |
| Pipeline | automatic-speech-recognition |
| License | apache-2.0 |
| Base model | OpenMOSS-Team/MOSS-Transcribe-Diarize |
| Last modified | 2026-09-15T07:05:40.000Z |
Model README
---
license: apache-2.0
base_model: OpenMOSS-Team/MOSS-Transcribe-Diarize
base_model_relation: quantized
library_name: transcribe.cpp
pipeline_tag: automatic-speech-recognition
language:
- en
- zh
tags:
- gguf
- transcribe.cpp
- asr
- speech-to-text
- moss
- audio-llm
- whisper-encoder
- qwen3
- diarization
transcribe_cpp:
schema_version: 2
wer_fleurs_en:
q8_0: 5.13
cer_fleurs_zh:
q8_0: 9.23
wer_librispeech_test_clean:
bf16: 2.08
f16: 2.07
q8_0: 1.93
q6_k: 1.96
q5_k_m: 1.99
q4_k_m: 2.59
rtf_m4_max:
cpu: 5.89
metal: 26.58
rtf_ryzen_4750u:
cpu: 1.57
vulkan: 3.07
streaming: false
diarize: true
translate: false
lang_detect: false
timestamps: segment
---
MOSS-Transcribe-Diarize: transcribe.cpp GGUF
GGUF conversions of OpenMOSS-Team/MOSS-Transcribe-Diarize for use
with transcribe.cpp.
Ported from upstream commit
pinned 2026-07-12.
Validated against the MOSS author repo (OpenMOSS/MOSS-Transcribe-Diarize) reference at transcribe.cpp commit
on 2026-07-12.
Offline English/Chinese speech-to-text with speaker diarization. A 0.9B
audio-LLM: a Whisper-Medium encoder (24 layers, d_model=1024) feeds a
4x temporal merge + VQAdaptor bridge into a Qwen3-0.6B decoder (28 layers)
via audio-token injection. The model emits [start][Sxx]text[end]; the
runtime parses those generated markers into clean text and segment rows.
Speaker attribution is opt-in (--diarize) and returns structured speaker
ids/turns. Not a streaming model.
Downloads
| Quantization | Download | Size | WER (LibriSpeech test-clean) |
| --- | --- | ---: | ---: |
| BF16 | MOSS-Transcribe-Diarize-BF16.gguf | 1.83 GB | 2.08% |
| F16 | MOSS-Transcribe-Diarize-F16.gguf | 1.83 GB | 2.07% |
| Q8_0 | MOSS-Transcribe-Diarize-Q8_0.gguf | 987 MB | 1.93% |
| Q6_K | MOSS-Transcribe-Diarize-Q6_K.gguf | 768 MB | 1.96% |
| Q5_K_M | MOSS-Transcribe-Diarize-Q5_K_M.gguf | 700 MB | 1.99% |
| Q4_K_M | MOSS-Transcribe-Diarize-Q4_K_M.gguf | 617 MB | 2.59% |
WER on the full LibriSpeech test-clean split (2,620 utterances), batch size 1, timestamps none. Figures without a commit were published before provenance was recorded.
Scored with the Whisper-style English text normalizer and jiwer 3.x. MOSS emits the
diarized format [start][Sxx]text[end]; the bracket spans are metadata and are
de-diarized to a space (for both hypothesis and reference) before scoring, matching
the author-repo reference runner. These values describe this dataset only, not a
general quality ranking: a quant that scores slightly better here is not necessarily
better in real-world use, because dataset-specific decoding near-ties can make
quantization noise help or hurt individual utterances. The same-manifest MOSS
author-repo reference (bf16, greedy) lands at 2.07% with 95% bootstrap CI
[1.82%, 2.40%]. The BF16 port lands at 2.08% (within +0.01 of the reference, well
inside the CI band); the lower-bit presets sit between 1.93% and 1.99% (statistical
noise) except Q4_K_M at 2.59%, whose excess is a handful of 4-bit tail failures (6
empty outputs, 5 English->Chinese language-drift utterances, 1 timestamp-token
repetition loop) rather than broad degradation. Prefer Q5_K_M or higher if those
tail failures matter. Reproduce with scripts/wer/run.py + `scripts/wer/score.py
--dediarize; public full_text` applies equivalent marker removal.
Usage
Build transcribe.cpp from source:
git clone git@github.com:handy-computer/transcribe.cpp.git
cd transcribe.cpp
cmake -B build && cmake --build build
Run on a 16 kHz mono WAV:
build/bin/transcribe-cli \
-m MOSS-Transcribe-Diarize-Q8_0.gguf \
input.wav
If your audio isn't already 16 kHz mono WAV, convert it first:
ffmpeg -i input.mp3 -ar 16000 -ac 1 output.wav
See the transcribe.cpp model page for performance
numbers, numerical validation, and reproduction steps.
License
Inherited from the base model: Apache-2.0. See the
upstream model card for full terms.
---
Original Model Card
> The section below is reproduced from
> OpenMOSS-Team/MOSS-Transcribe-Diarize at commit
> d7231bb for offline reference. The upstream card is the
> authoritative source.
MOSS-Transcribe-Diarize
<div align="center">
<a href="https://github.com/OpenMOSS/MOSS-Transcribe-Diarize"><img src="https://img.shields.io/badge/GitHub-OpenMOSS%2FMOSS--Transcribe--Diarize-black?logo=github"></a>
<a href="https://huggingface.co/OpenMOSS-Team/MOSS-Transcribe-Diarize"><img src="https://img.shields.io/badge/HuggingFace-Model-orange?logo=huggingface"></a>
<a href="https://arxiv.org/abs/2601.01554"><img src="https://img.shields.io/badge/arXiv-2601.01554-b31b1b?logo=arxiv"></a>
</div>
MOSS-Transcribe-Diarize 0.9B is an end-to-end audio understanding model for long-form multi-speaker transcription, diarization, timestamps, and acoustic event awareness.
Given an audio or video file, the model generates a compact speaker-aware transcript in one pass, including timestamps and anonymous speaker labels such as [S01], [S02], and beyond.
News
- 2026-07-09: Released MOSS-Transcribe-Diarize 0.9B.
Contents
Introduction
MOSS-Transcribe-Diarize 0.9B turns real-world long-form audio into structured, speaker-aware transcripts in one pass. Instead of stitching together separate ASR and diarization systems, it jointly performs speech transcription and speaker diarization, producing time-aligned text with consistent speaker labels.
The model is built for meetings, calls, podcasts, interviews, lectures, videos, and other long or messy multi-speaker recordings. It can also emit acoustic event annotations, giving downstream systems a richer view of what happened, who spoke, and when.
Core capabilities:
- Long-form transcription: Converts long audio or video recordings into timestamped text.
- Speaker-aware diarization: Assigns anonymous speaker labels such as
[S01]and[S02]without a separate diarization pipeline. - Promptable generation: Supports custom transcription instructions, hotwords, and acoustic event annotations.
Model Architecture
<p align="center">
<img src="Model_Architecture.png" alt="MOSS-Transcribe-Diarize 0.9B model architecture" width="900">
</p>
| Component | Specification |
|---|---|
| Text backbone | Qwen3-0.6B style causal decoder |
| Audio encoder | Whisper-Medium encoder configuration |
| Audio frontend | WhisperFeatureExtractor, 16 kHz, 80 mel bins, 30 s chunks |
| Audio-text bridge | 4x temporal merge + MLP adaptor |
| Fusion | Audio features replace <code><|audio_pad|></code> embeddings via masked_scatter |
| Output format | Compact [start][Sxx]text[end] transcript with speaker tags such as [S01] |
This Hugging Face repository includes the custom Transformers remote code required to load the model with trust_remote_code=True.
Evaluation
We evaluate MOSS-Transcribe-Diarize using three objective metrics: Character Error Rate (CER), concatenated minimum-permutation Character Error Rate (cpCER), and Delta-cp. Lower is better for all metrics. A dash (-) indicates that the result is unavailable.
<div style="overflow-x: auto;">
<table style="white-space: nowrap;">
<thead>
<tr>
<th rowspan="2" style="min-width: 220px;">Model</th>
<th colspan="3" style="text-align:center;">AISHELL‑4</th>
<th colspan="3" style="text-align:center;">Alimeeting</th>
<th colspan="3" style="text-align:center;">Podcast</th>
<th colspan="3" style="text-align:center;">Movies</th>
</tr>
<tr>
<th>CER↓</th><th>cpCER↓</th><th>Δcp↓</th>
<th>CER↓</th><th>cpCER↓</th><th>Δcp↓</th>
<th>CER↓</th><th>cpCER↓</th><th>Δcp↓</th>
<th>CER↓</th><th>cpCER↓</th><th>Δcp↓</th>
</tr>
</thead>
<tbody>
<tr>
<td style="white-space: nowrap;">Doubao</td>
<td>18.18</td><td>27.86</td><td>9.68</td>
<td>25.25</td><td>37.57</td><td>12.31</td>
<td>7.93</td><td>10.54</td><td>2.61</td>
<td>9.94</td><td>30.88</td><td>20.94</td>
</tr>
<tr>
<td style="white-space: nowrap;">ElevenLabs</td>
<td>19.58</td><td>37.95</td><td>18.36</td>
<td>25.70</td><td>36.69</td><td>10.99</td>
<td>8.50</td><td>11.34</td><td>2.85</td>
<td>11.49</td><td>17.85</td><td>6.37</td>
</tr>
<tr>
<td style="white-space: nowrap;">GPT-4o</td>
<td>-</td><td>-</td><td>-</td>
<td>-</td><td>-</td><td>-</td>
<td>-</td><td>-</td><td>-</td>
<td>14.37</td><td>23.67</td><td>9.31</td>
</tr>
<tr>
<td style="white-space: nowrap;">Gemini 2.5 Pro</td>
<td>42.70</td><td>53.42</td><td>10.72</td>
<td>27.43</td><td>41.64</td><td>14.21</td>
<td>7.38</td><td>10.23</td><td>2.85</td>
<td>15.46</td><td>24.15</td><td>8.69</td>
</tr>
<tr>
<td style="white-space: nowrap;">Gemini 3 Pro</td>
<td>22.75</td><td>27.43</td><td>4.68</td>
<td>26.75</td><td>32.84</td><td>6.09</td>
<td>-</td><td>-</td><td>-</td>
<td>8.62</td><td>14.73</td><td><u>6.11</u></td>
</tr>
<tr>
<td style="white-space: nowrap;">VIBEVOICE ASR</td>
<td>21.40</td><td>24.99</td><td>3.59</td>
<td>27.40</td><td>29.33</td><td>1.93</td>
<td>27.94</td><td>48.30</td><td>20.36</td>
<td>14.59</td><td>42.54</td><td>27.94</td>
</tr>
<tr>
<td style="white-space: nowrap;"><b>MOSS Transcribe Diarize 0.9B</b></td>
<td><u>14.84</u></td><td><u>15.83</u></td><td><u>0.99</u></td>
<td><u>24.86</u></td><td><u>22.17</u></td><td><u>-2.69</u></td>
<td><u>5.97</u></td><td><u>7.37</u></td><td><b>1.40</b></td>
<td><u>6.36</u></td><td><u>12.76</u></td><td>6.40</td>
</tr>
<tr>
<td style="white-space: nowrap;"><b>MOSS Transcribe Diarize Pro</b></td>
<td><b>13.78</b></td><td><b>14.02</b></td><td><b>0.24</b></td>
<td><b>18.22</b></td><td><b>13.94</b></td><td><b>-4.27</b></td>
<td><b>4.46</b></td><td><b>6.97</b></td><td><u>2.51</u></td>
<td><b>5.86</b></td><td><b>11.78</b></td><td><b>5.92</b></td>
</tr>
</tbody>
</table>
</div>
Quickstart
Environment Setup
Use a clean Python environment. The model uses custom Transformers code, so load the model and processor with trust_remote_code=True.
conda create -n moss-transcribe-diarize python=3.12 -y
conda activate moss-transcribe-diarize
git clone https://github.com/OpenMOSS/MOSS-Transcribe-Diarize.git
cd MOSS-Transcribe-Diarize
pip install --index-url https://download.pytorch.org/whl/cu128 torch torchaudio
pip install -e .
The GitHub package provides helper utilities such as audio/video loading, transcription message construction, transcript parsing, CLI inference, and the subtitle web app. The model weights and remote-code model files are loaded from this Hugging Face repository.
Python Usage
import torch
from transformers import AutoModelForCausalLM, AutoProcessor
from moss_transcribe_diarize import parse_transcript
from moss_transcribe_diarize.inference_utils import (
build_transcription_messages,
generate_transcription,
resolve_device,
)
model_id = "OpenMOSS-Team/MOSS-Transcribe-Diarize"
audio_path = "audio.wav"
device = resolve_device("auto")
dtype = torch.bfloat16 if device.type == "cuda" else torch.float32
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
dtype="auto",
).to(dtype=dtype).to(device).eval()
processor = AutoProcessor.from_pretrained(
model_id,
trust_remote_code=True,
)
messages = build_transcription_messages(audio_path)
result = generate_transcription(
model,
processor,
messages,
max_new_tokens=2048,
do_sample=False,
device=device,
dtype=dtype,
)
print(result["text"])
for segment in parse_transcript(result["text"]):
print(segment.start, segment.end, segment.speaker, segment.text)
The message flow follows the common Qwen multimodal pattern:
processor.apply_chat_template(messages, tokenize=False)renders text with audio placeholders.- The helper utilities load audio waveforms from the same messages.
processor(text=text, audio=audios)computes Whisper input features and expands audio placeholders.model.generate(...)produces timestamped transcription and diarization text.
Custom Prompt and Hotwords
The default prompt is optimized for timestamped transcription and speaker diarization:
请将音频转写为文本,每一段需以起始时间戳和说话人编号([S01]、[S02]、[S03]…)开头,正文为对应的语音内容,并在段末标注结束时间戳,以清晰标明该段语音范围。
To add hotwords, append a short hint to the default prompt:
请将音频转写为文本,每一段需以起始时间戳和说话人编号([S01]、[S02]、[S03]…)开头,正文为对应的语音内容,并在段末标注结束时间戳,以清晰标明该段语音范围。热词提示:热词1, 热词2, 热词3
More prompt recipes are available in the GitHub repository: <https://github.com/OpenMOSS/MOSS-Transcribe-Diarize/blob/main/examples/prompts.md>
Serve with vLLM and SGLang
MOSS-Transcribe-Diarize supports vLLM serving through the OpenAI-compatible transcription API. Use a pinned vLLM nightly build that includes the MOSS-Transcribe-Diarize model registration. Choose one of the following commands: for CUDA 12 environments, use cu129; for CUDA 13 environments, use cu130.
uv pip install -U vllm \
--torch-backend=auto \
--extra-index-url https://wheels.vllm.ai/68b4a1d582818e67adc903bf1b8fc5a5447da2fa/cu129
or:
uv pip install -U vllm \
--torch-backend=auto \
--extra-index-url https://wheels.vllm.ai/68b4a1d582818e67adc903bf1b8fc5a5447da2fa/cu130
vllm serve OpenMOSS-Team/MOSS-Transcribe-Diarize --trust-remote-code
curl http://localhost:8000/v1/audio/transcriptions \
-F model="OpenMOSS-Team/MOSS-Transcribe-Diarize" \
-F file=@"audio.wav" \
-F response_format="json" \
-F temperature="0"
The recommended way to serve MOSS-Transcribe-Diarize is SGLang Omni through the OpenAI-compatible /v1/audio/transcriptions endpoint. Install sglang-omni by following the installation guide, then download the model:
hf download OpenMOSS-Team/MOSS-Transcribe-Diarize
Serve the model:
sgl-omni serve \
--model-path OpenMOSS-Team/MOSS-Transcribe-Diarize \
--port 8000 \
--max-running-requests 16 \
--cuda-graph-max-bs 16 \
--mem-fraction-static 0.80
Use response_format=verbose_json when you need parsed speaker segments. json returns the raw transcript text only.
curl -X POST http://localhost:8000/v1/audio/transcriptions \
-F model=OpenMOSS-Team/MOSS-Transcribe-Diarize \
-F file=@audio.wav \
-F response_format=verbose_json
import requests
with open("audio.wav", "rb") as f:
resp = requests.post(
"http://localhost:8000/v1/audio/transcriptions",
data={
"model": "OpenMOSS-Team/MOSS-Transcribe-Diarize",
"response_format": "verbose_json",
},
files={"file": ("audio.wav", f, "audio/wav")},
timeout=300,
)
resp.raise_for_status()
payload = resp.json()
print(payload["text"])
for segment in payload.get("segments", []):
print(f"[{segment['start']:.2f}-{segment['end']:.2f}] {segment['text']}")
For longer multi-speaker audio, raise max_new_tokens so the decoder can finish the full diarized transcript:
curl -X POST http://localhost:8000/v1/audio/transcriptions \
-F model=OpenMOSS-Team/MOSS-Transcribe-Diarize \
-F file=@audio.wav \
-F response_format=verbose_json \
-F max_new_tokens=65536
| Parameter | Type | Default | Description |
|---|---|---|---|
| file | file | required | Audio file uploaded as multipart form data |
| model | string | server default | Model identifier |
| language | string | unset | Optional language hint |
| response_format | string | json | json, verbose_json, or text |
| temperature | float | model default (0.0) | Sampling temperature |
| max_new_tokens | int | 5120 | Max generated tokens; raise for long audio, for example 65536 |
| prompt | string | unset | Optional instruction override; omit to use the built-in transcribe+diarize prompt |
For benchmarking, performance numbers, and implementation details, see the SGLang Omni cookbook. The following single-H100 results are reported for short- and long-sequence multi-speaker ASR tasks.
movies short-sequence ASR:
| Concurrency | Throughput (req/s) | Mean latency (s) | RTF mean | audio_s/s |
|---:|---:|---:|---:|---:|
| 1 | 2.57 | 0.388 | 0.0612 | 29.76 |
| 2 | 4.89 | 0.409 | 0.0659 | 56.55 |
| 4 | 6.62 | 0.513 | 0.0790 | 76.64 |
| 8 | 6.80 | 0.533 | 0.0810 | 78.70 |
| 16 | 7.08 | 0.659 | 0.0922 | 81.98 |
aishell4_long long-sequence ASR:
| Concurrency | Throughput (req/s) | Mean latency (s) | RTF mean | audio_s/s |
|---:|---:|---:|---:|---:|
| 1 | 0.022 | 45.2 | 0.0197 | 50.64 |
| 2 | 0.032 | 60.7 | 0.0265 | 74.25 |
| 4 | 0.036 | 105.6 | 0.0461 | 81.64 |
| 8 | 0.040 | 172.6 | 0.0754 | 90.62 |
| 16 | 0.043 | 282.8 | 0.1237 | 98.83 |
Subtitle Web App
The source package includes a local subtitle workflow for upload, review, subtitle export, and optional FFmpeg burn-in:
mtd-subtitle-web \
--model OpenMOSS-Team/MOSS-Transcribe-Diarize \
--host 127.0.0.1 \
--port 7860
Open http://127.0.0.1:7860, upload an audio/video file, review the parsed subtitle segments, then download JSON/SRT/ASS or burn an MP4 if ffmpeg and ffprobe are available on PATH.
For batch processing:
mtd-subtitle /path/to/input.mp4 \
--model OpenMOSS-Team/MOSS-Transcribe-Diarize \
--out-dir runs/example \
--render
Output Format
The canonical output format is:
[start_time][Sxx]transcribed speech[end_time]
Example:
[0.48][S01]Welcome everyone[1.66][12.26][S02]The new transcription pipeline is ready for evaluation[13.81][14.36][S01]Great, include the diarization results in the report[18.76]
In this format:
start_timeandend_timeare timestamps in seconds.[S01],[S02], and similar labels are anonymous model-generated speaker labels.- Speaker labels are relative labels within the input audio and should not be interpreted as real speaker identities.
More Information
- GitHub: <https://github.com/OpenMOSS/MOSS-Transcribe-Diarize>
- MOSI.AI: <https://mosi.cn>
- OpenMOSS: <https://www.open-moss.com>
License
MOSS-Transcribe-Diarize 0.9B is licensed under the Apache License 2.0.
Citation
If you use MOSS-Transcribe-Diarize 0.9B, please cite the technical report:
@misc{moss_transcribe_diarize_2026,
title={MOSS Transcribe Diarize Technical Report},
author={{MOSI.AI}},
year={2026},
eprint={2601.01554},
archivePrefix={arXiv},
primaryClass={cs.SD},
url={https://arxiv.org/abs/2601.01554}
}Run handy-computer/moss-transcribe-diarize-gguf with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models