GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

scholzmx/moss-transcribe-preview-2b-gguf overview

MOSS Transcribe preview 2B: calibrated GGUF Starling engine A block quantized GGUF of OpenMOSS Team/MOSS Transcribe preview 2B Apache 2.0 , built with Starling…

ggufspeech-recognitionstarlingautomatic-speech-recognitionbase_model:OpenMOSS-Team/MOSS-Transcribe-preview-2Bbase_model:quantized:OpenMOSS-Team/MOSS-Transcribe-preview-2Blicense:apache-2.0region:us

Runs locally from ~1.45 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
automatic-speech-recognition
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
moss-transcribe-preview-2b-q4e8-fullimx.ggufGGUFQ4E81.45 GBDownload

Model Details

Model IDscholzmx/moss-transcribe-preview-2b-gguf
Authorscholzmx
Pipelineautomatic-speech-recognition
Licenseapache-2.0
Base modelOpenMOSS-Team/MOSS-Transcribe-preview-2B
Last modified2026-09-24T13:41:17.000Z

Model README

---

license: apache-2.0

base_model: OpenMOSS-Team/MOSS-Transcribe-preview-2B

base_model_relation: quantized

pipeline_tag: automatic-speech-recognition

tags:

- gguf

- speech-recognition

- starling

---

MOSS-Transcribe-preview-2B: calibrated GGUF (Starling engine)

A block-quantized GGUF of OpenMOSS-Team/MOSS-Transcribe-preview-2B

(Apache-2.0), built with Starling's in-tree quantization pipeline. The linear

layers use importance-matrix-weighted Q4_0, and the tied embedding / lm_head

uses Q8_0.

> Runtime note: this file follows the Starling GGUF tensor contract and runs

> on the native starling-serve binary / libstarling_ggml engine from the

> starling repository, including its

> Vulkan fast engine. It is not a llama.cpp / whisper.cpp GGUF.

File

| file | size | linears | embed (tied head) |

|------|------|---------|-------------------|

| moss-transcribe-preview-2b-q4e8-fullimx.gguf | 1.55 GB | Q4_0 + imatrix | Q8_0 |

SHA-256: 5658f3107a72bc7d74c3c428615fde9c95b439a436b9a3b1f387cf2f82a439ff

Recipe: quants/recipes/moss-q4e8-fullimx.recipe in the starling repository

(starling-quantize --recipe … --imatrix … --f32-1d from the BF16-exact

conversion).

Measured quality

FLEURS en_us test, first 100 clips, corpus WER:

| engine | WER |

|--------|-----|

| starling ggml (CPU) | 7.92 % |

| starling fast engine (Vulkan) | 7.87 % |

Speed (Starling fast engine)

| device | 7.4 s clip | decode |

|--------|------------|--------|

| Pixel 10 Pro (Tensor G5, PowerVR) | 5.6 s | 76 ms/token |

| Ryzen 5650U, Radeon Vega 7 iGPU | 1.5 s | ~33 ms/token |

Usage

hf download scholzmx/moss-transcribe-preview-2b-gguf \
  moss-transcribe-preview-2b-q4e8-fullimx.gguf --local-dir ./models
starling-serve --model moss \
  --gguf ./models/moss-transcribe-preview-2b-q4e8-fullimx.gguf --port 8181

License

Apache-2.0, inherited from the base model. Quantization by the Starling project.

Run scholzmx/moss-transcribe-preview-2b-gguf with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models