scholzmx/moss-transcribe-preview-2b-gguf overview
MOSS Transcribe preview 2B: calibrated GGUF Starling engine A block quantized GGUF of OpenMOSS Team/MOSS Transcribe preview 2B Apache 2.0 , built with Starling…
Runs locally from ~1.45 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| moss-transcribe-preview-2b-q4e8-fullimx.gguf | GGUF | Q4E8 | 1.45 GB | Download |
Model Details
| Model ID | scholzmx/moss-transcribe-preview-2b-gguf |
|---|---|
| Author | scholzmx |
| Pipeline | automatic-speech-recognition |
| License | apache-2.0 |
| Base model | OpenMOSS-Team/MOSS-Transcribe-preview-2B |
| Last modified | 2026-09-24T13:41:17.000Z |
Model README
---
license: apache-2.0
base_model: OpenMOSS-Team/MOSS-Transcribe-preview-2B
base_model_relation: quantized
pipeline_tag: automatic-speech-recognition
tags:
- gguf
- speech-recognition
- starling
---
MOSS-Transcribe-preview-2B: calibrated GGUF (Starling engine)
A block-quantized GGUF of OpenMOSS-Team/MOSS-Transcribe-preview-2B
(Apache-2.0), built with Starling's in-tree quantization pipeline. The linear
layers use importance-matrix-weighted Q4_0, and the tied embedding / lm_head
uses Q8_0.
> Runtime note: this file follows the Starling GGUF tensor contract and runs
> on the native starling-serve binary / libstarling_ggml engine from the
> starling repository, including its
> Vulkan fast engine. It is not a llama.cpp / whisper.cpp GGUF.
File
| file | size | linears | embed (tied head) |
|------|------|---------|-------------------|
| moss-transcribe-preview-2b-q4e8-fullimx.gguf | 1.55 GB | Q4_0 + imatrix | Q8_0 |
SHA-256: 5658f3107a72bc7d74c3c428615fde9c95b439a436b9a3b1f387cf2f82a439ff
Recipe: quants/recipes/moss-q4e8-fullimx.recipe in the starling repository
(starling-quantize --recipe … --imatrix … --f32-1d from the BF16-exact
conversion).
Measured quality
FLEURS en_us test, first 100 clips, corpus WER:
| engine | WER |
|--------|-----|
| starling ggml (CPU) | 7.92 % |
| starling fast engine (Vulkan) | 7.87 % |
Speed (Starling fast engine)
| device | 7.4 s clip | decode |
|--------|------------|--------|
| Pixel 10 Pro (Tensor G5, PowerVR) | 5.6 s | 76 ms/token |
| Ryzen 5650U, Radeon Vega 7 iGPU | 1.5 s | ~33 ms/token |
Usage
hf download scholzmx/moss-transcribe-preview-2b-gguf \
moss-transcribe-preview-2b-q4e8-fullimx.gguf --local-dir ./models
starling-serve --model moss \
--gguf ./models/moss-transcribe-preview-2b-q4e8-fullimx.gguf --port 8181
License
Apache-2.0, inherited from the base model. Quantization by the Starling project.
Run scholzmx/moss-transcribe-preview-2b-gguf with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models