GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

ukisai/Swift-Bonsai-2-GGUF overview

license: apache 2.0 license link: https://huggingface.co/ukisai/Swift Bonsai 2 GGUF/blob/main/LICENSE library name: llama.cpp pipeline tag: text generation bas…

llama.cppggufternary1-bit2-bitpq2reasoningswiftexperimentaltext-generationbase_model:prism-ml/Ternary-Bonsai-2-27B-ggufbase_model:finetune:prism-ml/Ternary-Bonsai-2-27B-gguflicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~5.54 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
7
Likes
14
Pipeline
text-generation
Author

Repository Files & Downloads

2 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Swift-Bonsai-2-PQ2_0.ggufGGUFGGUF6.71 GBDownload
Swift-Bonsai-2-PTQ1_0.ggufGGUFGGUF5.54 GBDownload

Model Details

Model IDukisai/Swift-Bonsai-2-GGUF
Authorukisai
Pipelinetext-generation
Licenseapache-2.0
Base modelprism-ml/Ternary-Bonsai-2-27B-gguf
Last modified2026-09-24T16:16:58.000Z

Model README

---

license: apache-2.0

license_link: https://huggingface.co/ukisai/Swift-Bonsai-2-GGUF/blob/main/LICENSE

library_name: llama.cpp

pipeline_tag: text-generation

base_model:

  • prism-ml/Ternary-Bonsai-2-27B-gguf

base_model_relation: finetune

tags:

  • gguf
  • ternary
  • 1-bit
  • 2-bit
  • pq2
  • reasoning
  • swift
  • experimental

---

<div align="center">

<a href="https://ukisai.com"><img src="ukisai-banner.png" alt="UkisAI" style="width:100%;max-width:100%;height:auto;display:block;margin-bottom:0.6em;" /></a>

<div style="display:flex;flex-wrap:wrap;justify-content:center;gap:0.6em;margin-bottom:1em;">

<a href="https://ukisai.com"><strong>Website</strong></a> &nbsp;&bull;&nbsp;

<a href="https://ukisai.com/products/swift"><strong>Learn more</strong></a> &nbsp;&bull;&nbsp;

<a href="#quantizations"><strong>Quantizations</strong></a> &nbsp;&bull;&nbsp;

<a href="#evaluation"><strong>Evaluation</strong></a> &nbsp;&bull;&nbsp;

<a href="#how-to-use"><strong>How to use</strong></a> &nbsp;&bull;&nbsp;

<a href="#license-and-attribution"><strong>License</strong></a>

</div>

</div>

Swift Bonsai 2 GGUF

Swift Bonsai 2 is UkisAI's reasoning-efficient derivative of Prism ML's Ternary Bonsai 2 27B.

It uses 39.8% fewer thinking tokens while scoring 0.19% higher than the base.

Both the 1-bit and 2-bit quantizations are available in this repository; each file is a plain Bonsai 2 pack.

Quantizations

<div style="max-width:100%;overflow-x:auto;">

<table>

<thead>

<tr>

<th>Quantization</th>

<th>File</th>

<th style="text-align: right;">Download size</th>

</tr>

</thead>

<tbody>

<tr>

<td>1-bit / PTQ1_0</td>

<td><a href="https://huggingface.co/ukisai/Swift-Bonsai-2-GGUF/resolve/main/Swift-Bonsai-2-PTQ1_0.gguf">Swift-Bonsai-2-PTQ1_0.gguf</a></td>

<td style="text-align: right;">5.947 GB</td>

</tr>

<tr>

<td>2-bit / PQ2_0</td>

<td><a href="https://huggingface.co/ukisai/Swift-Bonsai-2-GGUF/resolve/main/Swift-Bonsai-2-PQ2_0.gguf">Swift-Bonsai-2-PQ2_0.gguf</a></td>

<td style="text-align: right;">7.206 GB</td>

</tr>

</tbody>

</table>

</div>

Each file is a complete model with the Swift correction merged into the ternary weights: no adapter file, patch, or extra flag is needed, and the files are the same size as the base Bonsai 2 packs. PTQ1_0 remains the earlier Swift release; PQ2_0 now contains updated merged weights.

Training approach

We built Swift by identifying reasoning-marker tokens that, in our analysis, trigger overthinking in the model's reasoning rollouts.

We then fine-tuned the model by penalizing usage of those tokens while it reasons.

Swift produces shorter reasoning traces while keeping accuracy in line with the base model.

Evaluation

GPQA and C-Eval compare base Ternary Bonsai 2 27B with the historical Swift runtime correction under the saved protocols. IFBench and AIME compare base PQ2_0 with the current Swift PQ2_0 file. See the note below. Scores are the percentage of correct responses across all scored repetitions: three complete GPQA-Diamond runs and five runs each of C-Eval, IFBench, and AIME 2025. Token reductions are relative to the base model.

<style>

.bonsai-table { display:table !important; width:100% !important; box-sizing:border-box; table-layout:fixed; border-collapse:separate; border-spacing:0; overflow:hidden; border:1px solid #27344A; border-radius:20px; background:#0D111B; font-family:-apple-system,BlinkMacSystemFont,"Segoe UI",Roboto,sans-serif; font-size:14px; color:#BFBDBD; }

.bonsai-table th { padding:13px 8px; text-align:center; font-weight:700; color:#AEB5C7; background:#0D111B; border-right:1px solid #27344A; border-bottom:1px solid #27344A; }

.bonsai-table td { padding:14px 8px; text-align:center; color:#BFBDBD; background:#0D111B; border-right:1px solid #27344A; border-bottom:1px solid #27344A; vertical-align:middle; overflow-wrap:break-word; }

.bonsai-table tr > :last-child { border-right:0; }

.bonsai-table tbody tr:last-child td { border-bottom:0; }

.bonsai-table .benchmark-heading { color:#B7BDCD; background:#0D111B; border-bottom:3px solid #7D45B5; }

.bonsai-table .score-heading { color:#F0C5FF; background:#52239E; border-bottom:3px solid #7D45B5; }

.bonsai-table .tokens-heading, .bonsai-table .median-heading { color:#D4E8FF; background:#304FC2; border-bottom:3px solid #5687E6; }

.bonsai-table .change-heading { color:#61B9FF; background:#101D2D; border-bottom:3px solid #5687E6; }

.bonsai-table .change { color:#61B9FF; background:#101D2D; font-weight:700; }

.bonsai-table .gain { color:#DDA8FF; background:#171127; font-weight:700; }

.bonsai-table .model { padding-left:18px; text-align:left; color:#FFFFFF; font-weight:600; }

.bonsai-table strong { color:#FFFFFF; }

.bonsai-table .section { padding:12px 18px; text-align:left; color:#B489FF; background:#2A2541; font-weight:700; letter-spacing:.08em; text-transform:uppercase; border-top:1px solid #3A3159; border-bottom:1px solid #3A3159; }

.bonsai-table .bonsai { background:#171127; }

.bonsai-table thead tr:nth-child(2) .bonsai { color:#D3A0FF; }

.bonsai-table .detail { color:#8C94A8; font-size:12px; font-weight:500; }

.bonsai-table.bonsai-compact { display:table !important; width:100% !important; table-layout:fixed !important; }

@media (max-width: 640px) {

.bonsai-table { display:block !important; width:100% !important; max-width:100%; overflow-x:auto !important; -webkit-overflow-scrolling:touch; table-layout:auto !important; }

.bonsai-table th, .bonsai-table td { min-width:104px; }

.bonsai-table th:first-child, .bonsai-table td:first-child { min-width:170px; }

}

</style>

<table class="bonsai-table">

<thead>

<tr>

<th class="benchmark-heading" rowspan="2" style="width:28%;text-align:left;padding-left:18px;vertical-align:bottom;">Benchmark</th>

<th class="score-heading" colspan="2">Score</th>

<th class="tokens-heading" colspan="3">Mean tokens</th>

<th class="median-heading">Median tokens</th>

</tr>

<tr><th>Base</th><th class="bonsai">Swift</th><th>Base</th><th class="bonsai">Swift</th><th class="change-heading">Reduction</th><th class="change-heading">Reduction</th></tr>

</thead>

<tbody>

<tr><td class="section" colspan="7">General reasoning</td></tr>

<tr><td class="model">GPQA-Diamond</td><td>84.18%</td><td class="bonsai"><strong>84.34%</strong></td><td>20,066</td><td class="bonsai"><strong>16,245</strong></td><td class="change">&#8595; 19.0%</td><td class="change">&#8595; 39.8%</td></tr>

<tr><td class="model">C-Eval</td><td>81.37%</td><td class="bonsai"><strong>81.47%</strong></td><td>2,123</td><td class="bonsai"><strong>1,732</strong></td><td class="change">&#8595; 18.4%</td><td class="change">&#8595; 7.1%</td></tr>

<tr><td class="model">IFBench</td><td>82.13%</td><td class="bonsai"><strong>82.60%</strong></td><td>8,670</td><td class="bonsai"><strong>8,519</strong></td><td class="change">&#8595; 1.7%</td><td class="change">&#8593; 0.4%</td></tr>

<tr><td class="section" colspan="7">Mathematics</td></tr>

<tr><td class="model">AIME 2025</td><td>92.00%</td><td class="bonsai"><strong>93.33%</strong></td><td><strong>22,158</strong></td><td class="bonsai">23,101</td><td class="change">&#8593; 4.3%</td><td class="change">&#8593; 5.1%</td></tr>

</tbody>

</table>

Token statistics measure thinking tokens, except AIME, where they measure the full completion. IFBench uses the official loose scorer on the full 1,500 responses. An up arrow means the current 2-bit model used more tokens. Token changes are not a direct measurement of latency or cost changes.

<details>

<summary><strong>Benchmark methodology and reproduction settings</strong></summary>

<p><strong>Sampling:</strong> temperature 1.0, top-p 0.95, top-k 20, min-p 0, repetition penalty 1, presence penalty 0, and no additional inference-time logit penalty.</p>

<div style="max-width:100%;overflow-x:auto;"><table>

<thead>

<tr>

<th>Benchmark</th>

<th style="text-align: right;">Questions</th>

<th style="text-align: right;">Repetitions</th>

<th style="text-align: right;">Scored responses</th>

<th style="text-align: right;">Output cap</th>

</tr>

</thead>

<tbody>

<tr>

<td>GPQA-Diamond</td>

<td style="text-align: right;">198</td>

<td style="text-align: right;">3</td>

<td style="text-align: right;">594</td>

<td style="text-align: right;">81,920</td>

</tr>

<tr>

<td>C-Eval validation, 5-shot</td>

<td style="text-align: right;">1,346</td>

<td style="text-align: right;">5</td>

<td style="text-align: right;">6,730</td>

<td style="text-align: right;">16,384</td>

</tr>

<tr>

<td>IFBench</td>

<td style="text-align: right;">300</td>

<td style="text-align: right;">5</td>

<td style="text-align: right;">1,500</td>

<td style="text-align: right;">81,920</td>

</tr>

<tr>

<td>AIME 2025</td>

<td style="text-align: right;">30</td>

<td style="text-align: right;">5</td>

<td style="text-align: right;">150</td>

<td style="text-align: right;">81,920</td>

</tr>

</tbody>

</table></div>

<ul>

<li><strong>C-Eval</strong> covers the complete validation split, not the hidden test set. One duplicated prompt is cached per seed; all question IDs are scored and weighted separately.</li>

<li><strong>AIME scoring:</strong> the current base and Swift PQ2_0 files used identical corrected prompts (30 questions &times; 5 seeds, output cap 81,920) and the same archived Math-Verify 0.9.0 scorer, which checks the final response and falls back to reasoning. That scorer gives <strong>92.00% base / 93.33% Swift</strong>. Final-answer-only scoring gives <strong>90.67% / 90.67%</strong>. The AIME token columns report full completion tokens.</li>

<li><strong>IFBench</strong> reports official loose scoring for the current base and Swift PQ2_0 files (300 questions &times; 5 seeds). The current run completed all 1,500 responses without request errors; prompt and sampler settings were spot-checked against the base run.</li>

<li><strong>Model selection:</strong> GPQA informed the choice of training settings and is not an untouched holdout. Named benchmark exclusions and exact-match checks were performed; a comprehensive near-overlap audit was not completed.</li>

<li>GPQA and C-Eval are the original measured base/Swift runtime results. IFBench and AIME were refreshed on the current PQ2_0 file against base PQ2_0. PTQ1_0 remains the earlier merged release. Reproducibility can vary with runtime builds and hardware.</li>

<li>Small score differences are not established improvements. Exact aggregates, truncation counts, and confidence intervals where computed are available in <a href="https://huggingface.co/ukisai/Swift-Bonsai-2-GGUF/blob/main/benchmark_results.json"><code>benchmark_results.json</code></a>.</li>

</ul>

</details>

Research status and limitations

This is an experimental research release. In our internal evaluations and practical testing, benchmark scores did not consistently translate into reliable general-purpose behavior. Instruction following, tool use, and open-ended coding or agent tasks remain uneven. These observations concern the specific models, runtimes, and tests we used; they do not establish a general conclusion about ternary models.

We see ternary quantization as a promising direction for making larger models more accessible. Further advances in training, low-bit adaptation, and inference support may make the approach increasingly useful. This release shares a research result and its current limitations, rather than presenting a production-ready model.

How to use

Download and run

These files use Prism's PQ2_0 and PTQ1_0 ternary packs and run on the PrismML-Eng/llama.cpp fork (tested at revision 1a07bfa5f). Stock llama.cpp does not support these tensor types. No patch, adapter file, or extra flag is needed.

Linux prerequisites: Git, CMake, a C++ compiler, and the NVIDIA CUDA toolkit.

git clone https://github.com/PrismML-Eng/llama.cpp && cd llama.cpp
cmake -B build -DGGML_CUDA=ON && cmake --build build --config Release -j
hf download ukisai/Swift-Bonsai-2-GGUF Swift-Bonsai-2-PQ2_0.gguf --local-dir .
./build/bin/llama-server -m Swift-Bonsai-2-PQ2_0.gguf --host 127.0.0.1 --port 8080 \
  -ngl 99 -c 32768 --flash-attn on --jinja \
  --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --repeat-penalty 1.0

To run 1-bit / PTQ1_0 instead, download Swift-Bonsai-2-PTQ1_0.gguf and pass it to -m.

The command above serves a 32,768-token context on 127.0.0.1:8080. Use -c 98304 for the longer evaluation caps if memory permits. Runtime memory includes model weights and context/state caches; download size is not total VRAM usage.

OpenAI-compatible API

curl http://127.0.0.1:8080/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "bonsai-2-swift",
    "messages": [{"role": "user", "content": "Explain the difference between correlation and causation."}],
    "temperature": 1.0,
    "top_p": 0.95,
    "top_k": 20,
    "min_p": 0.0,
    "max_tokens": 4096,
    "chat_template_kwargs": {"enable_thinking": true, "reasoning_effort": "xhigh"}
  }'

The launcher flags above are the intended model settings. Custom integrations should preserve them.

License and attribution

Created using Bonsai by Prism ML. This is an independent UkisAI release, not an official Prism ML release or endorsement.

The model is distributed under Apache-2.0. The upstream notice and attributions are retained.

Citation

@misc{swift-bonsai-2-gguf,
  title  = {Swift Bonsai 2 GGUF},
  author = {UkisAI},
  year   = {2026},
  url    = {https://huggingface.co/ukisai/Swift-Bonsai-2-GGUF}
}

Run ukisai/Swift-Bonsai-2-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models