ukisai/Swift-Bonsai-2-GGUF overview
license: apache 2.0 license link: https://huggingface.co/ukisai/Swift Bonsai 2 GGUF/blob/main/LICENSE library name: llama.cpp pipeline tag: text generation bas…
Runs locally from ~5.54 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | ukisai/Swift-Bonsai-2-GGUF |
|---|---|
| Author | ukisai |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | prism-ml/Ternary-Bonsai-2-27B-gguf |
| Last modified | 2026-09-24T16:16:58.000Z |
Model README
---
license: apache-2.0
license_link: https://huggingface.co/ukisai/Swift-Bonsai-2-GGUF/blob/main/LICENSE
library_name: llama.cpp
pipeline_tag: text-generation
base_model:
- prism-ml/Ternary-Bonsai-2-27B-gguf
base_model_relation: finetune
tags:
- gguf
- ternary
- 1-bit
- 2-bit
- pq2
- reasoning
- swift
- experimental
---
<div align="center">
<a href="https://ukisai.com"><img src="ukisai-banner.png" alt="UkisAI" style="width:100%;max-width:100%;height:auto;display:block;margin-bottom:0.6em;" /></a>
<div style="display:flex;flex-wrap:wrap;justify-content:center;gap:0.6em;margin-bottom:1em;">
<a href="https://ukisai.com"><strong>Website</strong></a> •
<a href="https://ukisai.com/products/swift"><strong>Learn more</strong></a> •
<a href="#quantizations"><strong>Quantizations</strong></a> •
<a href="#evaluation"><strong>Evaluation</strong></a> •
<a href="#how-to-use"><strong>How to use</strong></a> •
<a href="#license-and-attribution"><strong>License</strong></a>
</div>
</div>
Swift Bonsai 2 GGUF
Swift Bonsai 2 is UkisAI's reasoning-efficient derivative of Prism ML's Ternary Bonsai 2 27B.
It uses 39.8% fewer thinking tokens while scoring 0.19% higher than the base.
Both the 1-bit and 2-bit quantizations are available in this repository; each file is a plain Bonsai 2 pack.
Quantizations
<div style="max-width:100%;overflow-x:auto;">
<table>
<thead>
<tr>
<th>Quantization</th>
<th>File</th>
<th style="text-align: right;">Download size</th>
</tr>
</thead>
<tbody>
<tr>
<td>1-bit / PTQ1_0</td>
<td><a href="https://huggingface.co/ukisai/Swift-Bonsai-2-GGUF/resolve/main/Swift-Bonsai-2-PTQ1_0.gguf">Swift-Bonsai-2-PTQ1_0.gguf</a></td>
<td style="text-align: right;">5.947 GB</td>
</tr>
<tr>
<td>2-bit / PQ2_0</td>
<td><a href="https://huggingface.co/ukisai/Swift-Bonsai-2-GGUF/resolve/main/Swift-Bonsai-2-PQ2_0.gguf">Swift-Bonsai-2-PQ2_0.gguf</a></td>
<td style="text-align: right;">7.206 GB</td>
</tr>
</tbody>
</table>
</div>
Each file is a complete model with the Swift correction merged into the ternary weights: no adapter file, patch, or extra flag is needed, and the files are the same size as the base Bonsai 2 packs. PTQ1_0 remains the earlier Swift release; PQ2_0 now contains updated merged weights.
Training approach
We built Swift by identifying reasoning-marker tokens that, in our analysis, trigger overthinking in the model's reasoning rollouts.
We then fine-tuned the model by penalizing usage of those tokens while it reasons.
Swift produces shorter reasoning traces while keeping accuracy in line with the base model.
Evaluation
GPQA and C-Eval compare base Ternary Bonsai 2 27B with the historical Swift runtime correction under the saved protocols. IFBench and AIME compare base PQ2_0 with the current Swift PQ2_0 file. See the note below. Scores are the percentage of correct responses across all scored repetitions: three complete GPQA-Diamond runs and five runs each of C-Eval, IFBench, and AIME 2025. Token reductions are relative to the base model.
<style>
.bonsai-table { display:table !important; width:100% !important; box-sizing:border-box; table-layout:fixed; border-collapse:separate; border-spacing:0; overflow:hidden; border:1px solid #27344A; border-radius:20px; background:#0D111B; font-family:-apple-system,BlinkMacSystemFont,"Segoe UI",Roboto,sans-serif; font-size:14px; color:#BFBDBD; }
.bonsai-table th { padding:13px 8px; text-align:center; font-weight:700; color:#AEB5C7; background:#0D111B; border-right:1px solid #27344A; border-bottom:1px solid #27344A; }
.bonsai-table td { padding:14px 8px; text-align:center; color:#BFBDBD; background:#0D111B; border-right:1px solid #27344A; border-bottom:1px solid #27344A; vertical-align:middle; overflow-wrap:break-word; }
.bonsai-table tr > :last-child { border-right:0; }
.bonsai-table tbody tr:last-child td { border-bottom:0; }
.bonsai-table .benchmark-heading { color:#B7BDCD; background:#0D111B; border-bottom:3px solid #7D45B5; }
.bonsai-table .score-heading { color:#F0C5FF; background:#52239E; border-bottom:3px solid #7D45B5; }
.bonsai-table .tokens-heading, .bonsai-table .median-heading { color:#D4E8FF; background:#304FC2; border-bottom:3px solid #5687E6; }
.bonsai-table .change-heading { color:#61B9FF; background:#101D2D; border-bottom:3px solid #5687E6; }
.bonsai-table .change { color:#61B9FF; background:#101D2D; font-weight:700; }
.bonsai-table .gain { color:#DDA8FF; background:#171127; font-weight:700; }
.bonsai-table .model { padding-left:18px; text-align:left; color:#FFFFFF; font-weight:600; }
.bonsai-table strong { color:#FFFFFF; }
.bonsai-table .section { padding:12px 18px; text-align:left; color:#B489FF; background:#2A2541; font-weight:700; letter-spacing:.08em; text-transform:uppercase; border-top:1px solid #3A3159; border-bottom:1px solid #3A3159; }
.bonsai-table .bonsai { background:#171127; }
.bonsai-table thead tr:nth-child(2) .bonsai { color:#D3A0FF; }
.bonsai-table .detail { color:#8C94A8; font-size:12px; font-weight:500; }
.bonsai-table.bonsai-compact { display:table !important; width:100% !important; table-layout:fixed !important; }
@media (max-width: 640px) {
.bonsai-table { display:block !important; width:100% !important; max-width:100%; overflow-x:auto !important; -webkit-overflow-scrolling:touch; table-layout:auto !important; }
.bonsai-table th, .bonsai-table td { min-width:104px; }
.bonsai-table th:first-child, .bonsai-table td:first-child { min-width:170px; }
}
</style>
<table class="bonsai-table">
<thead>
<tr>
<th class="benchmark-heading" rowspan="2" style="width:28%;text-align:left;padding-left:18px;vertical-align:bottom;">Benchmark</th>
<th class="score-heading" colspan="2">Score</th>
<th class="tokens-heading" colspan="3">Mean tokens</th>
<th class="median-heading">Median tokens</th>
</tr>
<tr><th>Base</th><th class="bonsai">Swift</th><th>Base</th><th class="bonsai">Swift</th><th class="change-heading">Reduction</th><th class="change-heading">Reduction</th></tr>
</thead>
<tbody>
<tr><td class="section" colspan="7">General reasoning</td></tr>
<tr><td class="model">GPQA-Diamond</td><td>84.18%</td><td class="bonsai"><strong>84.34%</strong></td><td>20,066</td><td class="bonsai"><strong>16,245</strong></td><td class="change">↓ 19.0%</td><td class="change">↓ 39.8%</td></tr>
<tr><td class="model">C-Eval</td><td>81.37%</td><td class="bonsai"><strong>81.47%</strong></td><td>2,123</td><td class="bonsai"><strong>1,732</strong></td><td class="change">↓ 18.4%</td><td class="change">↓ 7.1%</td></tr>
<tr><td class="model">IFBench</td><td>82.13%</td><td class="bonsai"><strong>82.60%</strong></td><td>8,670</td><td class="bonsai"><strong>8,519</strong></td><td class="change">↓ 1.7%</td><td class="change">↑ 0.4%</td></tr>
<tr><td class="section" colspan="7">Mathematics</td></tr>
<tr><td class="model">AIME 2025</td><td>92.00%</td><td class="bonsai"><strong>93.33%</strong></td><td><strong>22,158</strong></td><td class="bonsai">23,101</td><td class="change">↑ 4.3%</td><td class="change">↑ 5.1%</td></tr>
</tbody>
</table>
Token statistics measure thinking tokens, except AIME, where they measure the full completion. IFBench uses the official loose scorer on the full 1,500 responses. An up arrow means the current 2-bit model used more tokens. Token changes are not a direct measurement of latency or cost changes.
<details>
<summary><strong>Benchmark methodology and reproduction settings</strong></summary>
<p><strong>Sampling:</strong> temperature 1.0, top-p 0.95, top-k 20, min-p 0, repetition penalty 1, presence penalty 0, and no additional inference-time logit penalty.</p>
<div style="max-width:100%;overflow-x:auto;"><table>
<thead>
<tr>
<th>Benchmark</th>
<th style="text-align: right;">Questions</th>
<th style="text-align: right;">Repetitions</th>
<th style="text-align: right;">Scored responses</th>
<th style="text-align: right;">Output cap</th>
</tr>
</thead>
<tbody>
<tr>
<td>GPQA-Diamond</td>
<td style="text-align: right;">198</td>
<td style="text-align: right;">3</td>
<td style="text-align: right;">594</td>
<td style="text-align: right;">81,920</td>
</tr>
<tr>
<td>C-Eval validation, 5-shot</td>
<td style="text-align: right;">1,346</td>
<td style="text-align: right;">5</td>
<td style="text-align: right;">6,730</td>
<td style="text-align: right;">16,384</td>
</tr>
<tr>
<td>IFBench</td>
<td style="text-align: right;">300</td>
<td style="text-align: right;">5</td>
<td style="text-align: right;">1,500</td>
<td style="text-align: right;">81,920</td>
</tr>
<tr>
<td>AIME 2025</td>
<td style="text-align: right;">30</td>
<td style="text-align: right;">5</td>
<td style="text-align: right;">150</td>
<td style="text-align: right;">81,920</td>
</tr>
</tbody>
</table></div>
<ul>
<li><strong>C-Eval</strong> covers the complete validation split, not the hidden test set. One duplicated prompt is cached per seed; all question IDs are scored and weighted separately.</li>
<li><strong>AIME scoring:</strong> the current base and Swift PQ2_0 files used identical corrected prompts (30 questions × 5 seeds, output cap 81,920) and the same archived Math-Verify 0.9.0 scorer, which checks the final response and falls back to reasoning. That scorer gives <strong>92.00% base / 93.33% Swift</strong>. Final-answer-only scoring gives <strong>90.67% / 90.67%</strong>. The AIME token columns report full completion tokens.</li>
<li><strong>IFBench</strong> reports official loose scoring for the current base and Swift PQ2_0 files (300 questions × 5 seeds). The current run completed all 1,500 responses without request errors; prompt and sampler settings were spot-checked against the base run.</li>
<li><strong>Model selection:</strong> GPQA informed the choice of training settings and is not an untouched holdout. Named benchmark exclusions and exact-match checks were performed; a comprehensive near-overlap audit was not completed.</li>
<li>GPQA and C-Eval are the original measured base/Swift runtime results. IFBench and AIME were refreshed on the current PQ2_0 file against base PQ2_0. PTQ1_0 remains the earlier merged release. Reproducibility can vary with runtime builds and hardware.</li>
<li>Small score differences are not established improvements. Exact aggregates, truncation counts, and confidence intervals where computed are available in <a href="https://huggingface.co/ukisai/Swift-Bonsai-2-GGUF/blob/main/benchmark_results.json"><code>benchmark_results.json</code></a>.</li>
</ul>
</details>
Research status and limitations
This is an experimental research release. In our internal evaluations and practical testing, benchmark scores did not consistently translate into reliable general-purpose behavior. Instruction following, tool use, and open-ended coding or agent tasks remain uneven. These observations concern the specific models, runtimes, and tests we used; they do not establish a general conclusion about ternary models.
We see ternary quantization as a promising direction for making larger models more accessible. Further advances in training, low-bit adaptation, and inference support may make the approach increasingly useful. This release shares a research result and its current limitations, rather than presenting a production-ready model.
How to use
Download and run
These files use Prism's PQ2_0 and PTQ1_0 ternary packs and run on the PrismML-Eng/llama.cpp fork (tested at revision 1a07bfa5f). Stock llama.cpp does not support these tensor types. No patch, adapter file, or extra flag is needed.
Linux prerequisites: Git, CMake, a C++ compiler, and the NVIDIA CUDA toolkit.
git clone https://github.com/PrismML-Eng/llama.cpp && cd llama.cpp
cmake -B build -DGGML_CUDA=ON && cmake --build build --config Release -j
hf download ukisai/Swift-Bonsai-2-GGUF Swift-Bonsai-2-PQ2_0.gguf --local-dir .
./build/bin/llama-server -m Swift-Bonsai-2-PQ2_0.gguf --host 127.0.0.1 --port 8080 \
-ngl 99 -c 32768 --flash-attn on --jinja \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --repeat-penalty 1.0
To run 1-bit / PTQ1_0 instead, download Swift-Bonsai-2-PTQ1_0.gguf and pass it to -m.
The command above serves a 32,768-token context on 127.0.0.1:8080. Use -c 98304 for the longer evaluation caps if memory permits. Runtime memory includes model weights and context/state caches; download size is not total VRAM usage.
OpenAI-compatible API
curl http://127.0.0.1:8080/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "bonsai-2-swift",
"messages": [{"role": "user", "content": "Explain the difference between correlation and causation."}],
"temperature": 1.0,
"top_p": 0.95,
"top_k": 20,
"min_p": 0.0,
"max_tokens": 4096,
"chat_template_kwargs": {"enable_thinking": true, "reasoning_effort": "xhigh"}
}'
The launcher flags above are the intended model settings. Custom integrations should preserve them.
License and attribution
Created using Bonsai by Prism ML. This is an independent UkisAI release, not an official Prism ML release or endorsement.
The model is distributed under Apache-2.0. The upstream notice and attributions are retained.
Citation
@misc{swift-bonsai-2-gguf,
title = {Swift Bonsai 2 GGUF},
author = {UkisAI},
year = {2026},
url = {https://huggingface.co/ukisai/Swift-Bonsai-2-GGUF}
}Run ukisai/Swift-Bonsai-2-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models