GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

vcruz305/Ternary-Bonsai-27B-GGUF overview

Ternary Bonsai 27B — Standard GGUF Ladder Q3 K M → Q8 0 Imatrix GGUF quants of prism ml's Ternary Bonsai 27B https://huggingface.co/prism ml/Ternary Bonsai 27B…

llama.cppggufllama-cppimatrixbonsaitext-generationbase_model:prism-ml/Ternary-Bonsai-27B-unpackedbase_model:quantized:prism-ml/Ternary-Bonsai-27B-unpackedlicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~600.1 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

8 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Ternary-Bonsai-27B-IQ4_XS.ggufGGUFIQ4_XS14.05 GBDownload
Ternary-Bonsai-27B-Q3_K_M.ggufGGUFQ3_K_M12.39 GBDownload
Ternary-Bonsai-27B-Q4_K_M.ggufGGUFQ4_K_M15.41 GBDownload
Ternary-Bonsai-27B-Q5_K_M.ggufGGUFQ5_K_M17.91 GBDownload
Ternary-Bonsai-27B-Q6_K.ggufGGUFQ6_K20.57 GBDownload
Ternary-Bonsai-27B-Q8_0.ggufGGUFQ8_026.63 GBDownload
Ternary-Bonsai-27B-mmproj-BF16.ggufGGUFBF16888.0 MBDownload
Ternary-Bonsai-27B-mmproj-Q8_0.ggufGGUFQ8_0600.1 MBDownload

Model Details

Model IDvcruz305/Ternary-Bonsai-27B-GGUF
Authorvcruz305
Pipelinetext-generation
Licenseapache-2.0
Base modelprism-ml/Ternary-Bonsai-27B-unpacked
Last modified2026-07-14T20:51:37.000Z

Model README

---

license: apache-2.0

library_name: llama.cpp

pipeline_tag: text-generation

tags:

  • gguf
  • llama-cpp
  • imatrix
  • bonsai

base_model:

  • prism-ml/Ternary-Bonsai-27B-unpacked

---

Ternary-Bonsai-27B — Standard GGUF Ladder (Q3_K_M → Q8_0)

Imatrix GGUF quants of prism-ml's Ternary-Bonsai-27B (Qwen3.6-27B backbone, 262K context, hybrid attention, vision), covering the quality tiers between the official QAT release and F16.

Which repo should you use?

If you have ≤8GB: use prism-ml's official QAT quants, not these.

  • Q2_0 / PQ2_0 (7.2GB), Q2_g64 (7.6GB) — 95% of F16 quality, quantization-aware-trained with custom kernels. At those sizes they beat anything post-training quantization can produce, including anything in this repo.

This repo covers the gap above them: the standard ladder for 12–32GB setups, quantized from prism-ml's own F16 GGUF with an importance matrix (63KB coding/reasoning calibration corpus).

Quants

| File | Size | Measured (GB10, 273GB/s) |

|---|---|---|

| Q8_0 | 28.6 GB | 7.8 tok/s |

| Q6_K | 22.1 GB | 9.1 tok/s |

| Q5_K_M | 19.2 GB | 10.4 tok/s |

| Q4_K_M | 16.6 GB | 12.3 tok/s |

| IQ4_XS | 15.1 GB | 14.0 tok/s |

| Q3_K_M | 13.3 GB | 13.6 tok/s |

All tiers individually smoke-tested (coherent code generation, chat template engages via --jinja). Q8_0 quantized without imatrix (unneeded at 8-bit); all others use the included Ternary-Bonsai-27B.imatrix. Decode rates scale with memory bandwidth — a 936GB/s GPU (RTX 3090-class) will run ~3× these numbers.

Both official mmproj files are included for vision (--mmproj Ternary-Bonsai-27B-mmproj-Q8_0.gguf).

DSpark drafter: tested, NOT recommended with these quants

We verified prism-ml's Ternary DSpark drafter against this repo's Q4_K_M using their llama.cpp fork (branch prism):

| Config | tok/s | Draft acceptance |

|---|---|---|

| Q4_K_M base | 11.9 | — |

| Q4_K_M + DSpark drafter | 11.9–12.7 | 54% (217/404) |

Acceptance drops to ~54% (vs 79% for the Bonsai-27B pairing, where it delivers +31%) — at that rate the draft overhead roughly cancels the win on bandwidth-bound hardware. The drafter appears to track the ternary QAT weights more tightly than the F16 latent this ladder is quantized from. On high-bandwidth GPUs (where drafting is cheaper) it may still net positive — measure before adopting, and check timings.draft_n_accepted in any API response to verify. For guaranteed drafter gains at small sizes, use prism-ml's official QAT Q2_0 + drafter combination.

Provenance

  • Source: prism-ml/Ternary-Bonsai-27B-gguf F16 (53.8GB) — quantized directly from their official F16 GGUF, no re-conversion
  • llama.cpp: mainline, commit cecbf5fb0 (quantization + baseline numbers); PrismML fork 62061f91 branch prism (drafter test)
  • imatrix: 63KB ChatML coding/debugging/reasoning corpus, 4096 ctx
  • LICENSE/NOTICE carried over from the upstream repo (Apache-2.0)
  • Quantized on NVIDIA DGX Spark (GB10, aarch64) — day-zero release, same-day as the upstream drop
  • Companion repo: vcruz305/Bonsai-27B-GGUF — the non-ternary ladder, where the DSpark drafter pairing IS verified worthwhile (+31%)

Report issues in the community tab — smoke-test failures, incoherence, or numbers that don't reproduce. Community benchmark reports welcome (include hardware, backend, and full launch command).

Run vcruz305/Ternary-Bonsai-27B-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models