vcruz305/Ternary-Bonsai-27B-GGUF overview
Ternary Bonsai 27B — Standard GGUF Ladder Q3 K M → Q8 0 Imatrix GGUF quants of prism ml's Ternary Bonsai 27B https://huggingface.co/prism ml/Ternary Bonsai 27B…
Runs locally from ~600.1 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Ternary-Bonsai-27B-IQ4_XS.gguf | GGUF | IQ4_XS | 14.05 GB | Download |
| Ternary-Bonsai-27B-Q3_K_M.gguf | GGUF | Q3_K_M | 12.39 GB | Download |
| Ternary-Bonsai-27B-Q4_K_M.gguf | GGUF | Q4_K_M | 15.41 GB | Download |
| Ternary-Bonsai-27B-Q5_K_M.gguf | GGUF | Q5_K_M | 17.91 GB | Download |
| Ternary-Bonsai-27B-Q6_K.gguf | GGUF | Q6_K | 20.57 GB | Download |
| Ternary-Bonsai-27B-Q8_0.gguf | GGUF | Q8_0 | 26.63 GB | Download |
| Ternary-Bonsai-27B-mmproj-BF16.gguf | GGUF | BF16 | 888.0 MB | Download |
| Ternary-Bonsai-27B-mmproj-Q8_0.gguf | GGUF | Q8_0 | 600.1 MB | Download |
Model Details
| Model ID | vcruz305/Ternary-Bonsai-27B-GGUF |
|---|---|
| Author | vcruz305 |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | prism-ml/Ternary-Bonsai-27B-unpacked |
| Last modified | 2026-07-14T20:51:37.000Z |
Model README
---
license: apache-2.0
library_name: llama.cpp
pipeline_tag: text-generation
tags:
- gguf
- llama-cpp
- imatrix
- bonsai
base_model:
- prism-ml/Ternary-Bonsai-27B-unpacked
---
Ternary-Bonsai-27B — Standard GGUF Ladder (Q3_K_M → Q8_0)
Imatrix GGUF quants of prism-ml's Ternary-Bonsai-27B (Qwen3.6-27B backbone, 262K context, hybrid attention, vision), covering the quality tiers between the official QAT release and F16.
Which repo should you use?
If you have ≤8GB: use prism-ml's official QAT quants, not these.
- Q2_0 / PQ2_0 (7.2GB), Q2_g64 (7.6GB) — 95% of F16 quality, quantization-aware-trained with custom kernels. At those sizes they beat anything post-training quantization can produce, including anything in this repo.
This repo covers the gap above them: the standard ladder for 12–32GB setups, quantized from prism-ml's own F16 GGUF with an importance matrix (63KB coding/reasoning calibration corpus).
Quants
| File | Size | Measured (GB10, 273GB/s) |
|---|---|---|
| Q8_0 | 28.6 GB | 7.8 tok/s |
| Q6_K | 22.1 GB | 9.1 tok/s |
| Q5_K_M | 19.2 GB | 10.4 tok/s |
| Q4_K_M | 16.6 GB | 12.3 tok/s |
| IQ4_XS | 15.1 GB | 14.0 tok/s |
| Q3_K_M | 13.3 GB | 13.6 tok/s |
All tiers individually smoke-tested (coherent code generation, chat template engages via --jinja). Q8_0 quantized without imatrix (unneeded at 8-bit); all others use the included Ternary-Bonsai-27B.imatrix. Decode rates scale with memory bandwidth — a 936GB/s GPU (RTX 3090-class) will run ~3× these numbers.
Both official mmproj files are included for vision (--mmproj Ternary-Bonsai-27B-mmproj-Q8_0.gguf).
DSpark drafter: tested, NOT recommended with these quants
We verified prism-ml's Ternary DSpark drafter against this repo's Q4_K_M using their llama.cpp fork (branch prism):
| Config | tok/s | Draft acceptance |
|---|---|---|
| Q4_K_M base | 11.9 | — |
| Q4_K_M + DSpark drafter | 11.9–12.7 | 54% (217/404) |
Acceptance drops to ~54% (vs 79% for the Bonsai-27B pairing, where it delivers +31%) — at that rate the draft overhead roughly cancels the win on bandwidth-bound hardware. The drafter appears to track the ternary QAT weights more tightly than the F16 latent this ladder is quantized from. On high-bandwidth GPUs (where drafting is cheaper) it may still net positive — measure before adopting, and check timings.draft_n_accepted in any API response to verify. For guaranteed drafter gains at small sizes, use prism-ml's official QAT Q2_0 + drafter combination.
Provenance
- Source:
prism-ml/Ternary-Bonsai-27B-ggufF16 (53.8GB) — quantized directly from their official F16 GGUF, no re-conversion - llama.cpp: mainline, commit
cecbf5fb0(quantization + baseline numbers); PrismML fork62061f91branchprism(drafter test) - imatrix: 63KB ChatML coding/debugging/reasoning corpus, 4096 ctx
- LICENSE/NOTICE carried over from the upstream repo (Apache-2.0)
- Quantized on NVIDIA DGX Spark (GB10, aarch64) — day-zero release, same-day as the upstream drop
- Companion repo: vcruz305/Bonsai-27B-GGUF — the non-ternary ladder, where the DSpark drafter pairing IS verified worthwhile (+31%)
Report issues in the community tab — smoke-test failures, incoherence, or numbers that don't reproduce. Community benchmark reports welcome (include hardware, backend, and full launch command).
Run vcruz305/Ternary-Bonsai-27B-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models