scholzmx/parakeet-tdt-0.6b-v3-gguf overview
parakeet tdt 0.6b v3 — calibrated GGUF quants Starling engine Calibrated block quantized GGUFs of nvidia/parakeet tdt 0.6b v3 FastConformer + TDT transducer, 2…
Runs locally from ~310.1 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| parakeet-tdt-0.6b-v3-iq2_xxs-imx-shrink16.gguf | GGUF | IQ2_XXS | 310.1 MB | Download |
| parakeet-tdt-0.6b-v3-q2_k.gguf | GGUF | Q2_K | 546.9 MB | Download |
| parakeet-tdt-0.6b-v3-q3_k_m.gguf | GGUF | Q3_K_M | 604.2 MB | Download |
| parakeet-tdt-0.6b-v3-q4_k_m-shrink16.gguf | GGUF | Q4_K_M | 527.1 MB | Download |
| parakeet-tdt-0.6b-v3-q4_k_m.gguf | GGUF | Q4_K_M | 671.8 MB | Download |
| parakeet-tdt-0.6b-v3-q4_k_s.gguf | GGUF | Q4_K_S | 609.7 MB | Download |
| parakeet-tdt-0.6b-v3-q5_k_m.gguf | GGUF | Q5_K_M | 705.3 MB | Download |
| parakeet-tdt-0.6b-v3-q6_k.gguf | GGUF | Q6_K | 740.8 MB | Download |
| parakeet-tdt-0.6b-v3-q8_0.gguf | GGUF | Q8_0 | 864.0 MB | Download |
Model Details
Model README
---
license: cc-by-4.0
base_model: nvidia/parakeet-tdt-0.6b-v3
library_name: ggml
tags:
- speech-recognition
- asr
- gguf
- quantized
- multilingual
language:
- en
- de
- fr
- es
- it
- nl
- pl
- pt
- ru
- uk
- sv
- da
- fi
- cs
- sk
- sl
- hr
- bg
- el
- hu
- ro
- et
- lv
- lt
- mt
---
parakeet-tdt-0.6b-v3 — calibrated GGUF quants (Starling engine)
Calibrated block-quantized GGUFs of nvidia/parakeet-tdt-0.6b-v3 (FastConformer
- TDT transducer, 25 European languages, CC-BY-4.0), built with Starling's
in-tree quantization pipeline: llama.cpp-style quant levels (Q8_0 → IQ2) whose
block scales are weighted by an importance matrix collected on real speech
across all 25 supported languages (FLEURS train, 48 clips/language), verified
by WER with bootstrap confidence intervals.
> Runtime note: these files follow the Starling GGUF tensor contract and run
> on the native starling-serve binary / libstarling_ggml engine from the
> starling repository. They are not
> llama.cpp/whisper.cpp/parakeet.cpp GGUFs. (The converse works, though: the
> starling engine also loads community parakeet GGUFs — the cstr/parakeet.cpp
> and transcribe.cpp dialects — directly, via its loader compat layer.)
Files
| file | size | level | notes |
|------|------|-------|-------|
| parakeet-tdt-0.6b-v3-q8_0.gguf | ~906 MB | Q8_0 | default; lossless in practice |
| parakeet-tdt-0.6b-v3-q6_k.gguf | ~777 MB | Q6_K | |
| parakeet-tdt-0.6b-v3-q5_k_m.gguf | ~740 MB | Q5_K/Q6_K mix | |
| parakeet-tdt-0.6b-v3-q4_k_m.gguf | ~704 MB | Q4_K/Q6_K mix | free — no measurable WER loss vs f32 |
| parakeet-tdt-0.6b-v3-q4_k_s.gguf | ~639 MB | Q4_K | |
| parakeet-tdt-0.6b-v3-q3_k_m.gguf | ~634 MB | Q3_K/Q5_K mix | |
| parakeet-tdt-0.6b-v3-q2_k.gguf | ~574 MB | Q2_K/Q4_K mix | even multilingual option |
| parakeet-tdt-0.6b-v3-q4_k_m-shrink16.gguf | ~553 MB | Q4_K/Q6_K mix + F16 convs | size-competitive q4 (community-file parity, see below) |
| parakeet-tdt-0.6b-v3-iq2_xxs-imx-shrink16.gguf | ~325 MB | IQ2_XXS + F16 convs | smallest; English-first trade (see below) |
All levels ≥ Q4 were verified indistinguishable from f32 on clean and 5 dB-noised
English audio and on 300-clip German/English test sets; the sub-2-bit IQ file
trades tail-language accuracy for size (details below). The 640-row joint/LSTM
linears automatically stay at Q8_0 (block-size rule) in every level.
Measured quality (FLEURS test, mean WER [95% bootstrap CI])
| build | German | English |
|-------|--------|---------|
| f32 (2508 MB) | 5.30 [4.48–6.14] | 6.50 [5.57–7.36] |
| q4_k_m (704 MB) | 5.31 [4.49–6.16] | 6.41 [5.55–7.24] |
| q2_k (574 MB) | 6.03 [5.19–6.91] | 6.36 [5.51–7.24] |
| iq2_xxs + shrink16 (325 MB) | 9.38 [8.30–10.44] | 8.39 [7.34–9.36] |
25-language mean: f32 14.0 → q2_k 16.9 → iq2_xxs 26.4 (48 clips/language; the
iq2_xxs gap is concentrated in lt/lv/sl/mt/hu/sk). Choose by deployment:
**q4_k_m if quality matters, q2_k for even multilingual coverage at 574 MB,
iq2_xxs only for English-first deployments at 325 MB.**
Calibration-data notes (measured, not assumed): importance-matrix statistics
saturate at roughly a dozen clips per language (5× more calibration audio
changed the matrix by <1% and WER not at all); the calibration language mix
is not a measurable lever for English/German once a matrix exists (EN-only vs
25-language matrices are statistically indistinguishable on EN/DE at 300
clips); what matters is calibrated-vs-uniform (uniform Q2_K degrades to
30%/17% WER where the calibrated build matches f32) and the bit width itself.
Usage
hf download scholzmx/parakeet-tdt-0.6b-v3-gguf \
parakeet-tdt-0.6b-v3-q4_k_m.gguf --local-dir ./models
starling-serve --model parakeet --gguf ./models/parakeet-tdt-0.6b-v3-q4_k_m.gguf --port 8181
Provenance
- Base weights: nvidia/parakeet-tdt-0.6b-v3 (CC-BY-4.0)
- Conversion: starling
convert_parakeet_gguf.py(F32 base) - Calibration audio: FLEURS train split, 24 of the 25 languages × 48 clips
(Slovenian's shard was unavailable during collection; the measured
language-mix insensitivity makes this immaterial)
- Quantizer: starling
starling-quantize(ggml block quantization + imatrix) - Evaluation: starling
wer_quant.py, FLEURS test (300 clips EN/DE, 48/language elsewhere)
Quantization pipeline and full result tables: docs/quantization.md in the
starling repository.
Run scholzmx/parakeet-tdt-0.6b-v3-gguf with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models