GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

ukisai/Swift-1.5-Qwen3.8-27B-GGUF overview

<div align="center" <a href="https://ukisai.com" <img src="ukisai banner.png" alt="UkisAI" style="width:100%;max width:100%;height:auto;display:block;margin bo…

ggufllama.cppqwen3_8reasoningefficient-thinkingtoken-efficientpost-trainingterminal-benchimage-text-to-textbase_model:ukisai/Swift-1.5-Qwen3.8-27bbase_model:quantized:ukisai/Swift-1.5-Qwen3.8-27blicense:otherendpoints_compatibleregion:usimatrixconversational

Runs locally from ~884.6 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
90,036
Likes
133
Pipeline
image-text-to-text
Author

Repository Files & Downloads

23 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Swift-1.5-Qwen3.8-27B-IQ2_M.ggufGGUFIQ2_M9.80 GBDownload
Swift-1.5-Qwen3.8-27B-IQ2_S.ggufGGUFIQ2_S9.02 GBDownload
Swift-1.5-Qwen3.8-27B-IQ2_XS.ggufGGUFIQ2_XS8.46 GBDownload
Swift-1.5-Qwen3.8-27B-IQ2_XXS.ggufGGUFIQ2_XXS8.27 GBDownload
Swift-1.5-Qwen3.8-27B-IQ3_M.ggufGGUFIQ3_M13.84 GBDownload
Swift-1.5-Qwen3.8-27B-IQ3_XS.ggufGGUFIQ3_XS11.92 GBDownload
Swift-1.5-Qwen3.8-27B-IQ3_XXS.ggufGGUFIQ3_XXS11.47 GBDownload
Swift-1.5-Qwen3.8-27B-IQ4_NL.ggufGGUFIQ4_NL16.24 GBDownload
Swift-1.5-Qwen3.8-27B-IQ4_XS.ggufGGUFIQ4_XS14.41 GBDownload
Swift-1.5-Qwen3.8-27B-Q2_K.ggufGGUFQ2_K10.08 GBDownload
Swift-1.5-Qwen3.8-27B-Q3_K_L.ggufGGUFQ3_K_L13.15 GBDownload
Swift-1.5-Qwen3.8-27B-Q3_K_M.ggufGGUFQ3_K_M12.48 GBDownload
Swift-1.5-Qwen3.8-27B-Q3_K_S.ggufGGUFQ3_K_S11.86 GBDownload
Swift-1.5-Qwen3.8-27B-Q4_K_L.ggufGGUFQ4_K_L17.53 GBDownload
Swift-1.5-Qwen3.8-27B-Q4_K_M.ggufGGUFQ4_K_M16.24 GBDownload
Swift-1.5-Qwen3.8-27B-Q4_K_S.ggufGGUFQ4_K_S15.24 GBDownload
Swift-1.5-Qwen3.8-27B-Q5_K_M.ggufGGUFQ5_K_M19.49 GBDownload
Swift-1.5-Qwen3.8-27B-Q5_K_S.ggufGGUFQ5_K_S18.22 GBDownload
Swift-1.5-Qwen3.8-27B-Q6_K.ggufGGUFQ6_K22.22 GBDownload
Swift-1.5-Qwen3.8-27B-Q6_K_L.ggufGGUFQ6_K_L23.24 GBDownload
Swift-1.5-Qwen3.8-27B-Q6_K_S.ggufGGUFQ6_K_S21.29 GBDownload
Swift-1.5-Qwen3.8-27B-Q8_0.ggufGGUFQ8_027.05 GBDownload
mmproj-Swift-1.5-Qwen3.8-27B-F16.ggufGGUFF16884.6 MBDownload

Model Details

Model IDukisai/Swift-1.5-Qwen3.8-27B-GGUF
Authorukisai
Pipelineimage-text-to-text
Licenseother
Base modelukisai/Swift-1.5-Qwen3.8-27b
Last modified2026-10-01T12:53:24.000Z

Model README

---

license: other

license_name: swift-open-license-1.0

license_link: https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b/blob/main/LICENSE

library_name: gguf

pipeline_tag: image-text-to-text

tags:

  • gguf
  • llama.cpp
  • qwen3_8
  • reasoning
  • efficient-thinking
  • token-efficient
  • post-training
  • terminal-bench

base_model: ukisai/Swift-1.5-Qwen3.8-27b

base_model_relation: quantized

---

<div align="center">

<a href="https://ukisai.com"><img src="ukisai-banner.png" alt="UkisAI" style="width:100%;max-width:100%;height:auto;display:block;margin-bottom:0.6em;" /></a>

<div style="display:flex;justify-content:center;gap:0.6em;margin-bottom:1em;">

<a href="https://ukisai.com"><strong>Website</strong></a> &nbsp;&bull;&nbsp;

<a href="https://ukisai.com/products/swift"><strong>Learn more</strong></a> &nbsp;&bull;&nbsp;

<a href="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27B-GGUF"><strong>GGUF</strong></a> &nbsp;&bull;&nbsp;

<a href="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF"><strong>GSQ-RCO GGUF</strong></a> &nbsp;&bull;&nbsp;

<a href="#evaluation"><strong>Evaluation</strong></a> &nbsp;&bull;&nbsp;

<a href="#license-and-access"><strong>Enterprise licensing</strong></a>

</div>

</div>

Swift 1.5 Qwen3.8-27B

GGUF quantizations. Derived directly from Swift 1.5 with llama.cpp. Run a chosen quantization tier with a current llama.cpp-compatible runtime such as llama-server.

Swift 1.5 Qwen3.8-27B is UkisAI's reasoning-efficient derivative of Qwen3.8-27B.

It uses 58.5% fewer thinking tokens while scoring 0.35% higher than the base, for a 9.18× speed-up on several tasks.

Swift 1.5 is a direct upgrade from Swift 1.0, our model with 350k+ downloads, delivering stronger overall performance than both base and Swift 1.0 in various tasks, especially coding and agentic, while using fewer thinking tokens. We accomplished that by scaling up the post-training (RL and OPD) from the previous version.

Demo

We gave base Qwen3.8-27B and Swift 1.5 27B the same prompt:

> create a 3d little planet globe where I (player can walk around) and it has all these biomes to explore, the globe doesn't have to be too big, but still fun to go around. It's about a boy scout who is camping and goes around exploring.

<video src="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27B-GGUF/resolve/main/swift-1.5-planet-demo.mp4" controls autoplay muted loop playsinline style="width:100%;height:auto;border-radius:12px;"></video>

Try the game yourself here: https://ukisai.com/swift-games/27b

Base Qwen3.8-27B took 104.6 minutes to build its game. Swift 1.5 took 11.39 minutes.

Training approach

We made Swift efficient by figuring out which tokens were linked to pathological overthinking and penalizing them without "attacking" the reasoning length directly then regained the accuracy with RL and OPD, leading to "compressed" token usage while maintaining accuracy.

Swift 1.5 was made from Swift 1.0, on whom we scaled up the post-training methods that previously improved Swift1.0 model performance, this time with the main

focus on long-horizon, agentic, and coding tasks, as seen in the LiveCodeBench and Terminal Bench 2.1 improvements. Our training data is viewable here: https://huggingface.co/datasets/ukisai/Qwen3.8-27B-multi-turn-agent-sft albeit it is not used out of the box, but rather re-sampled, turned into proper RL environments etc.

Evaluation

The external results below compare Qwen3.8-27B, the foundation base model,

and Swift 1.5. Both models use

the same saved evaluation protocols, and all scores are reported as final aggregate

percentages.

<style>

.swift15-table { width:100%; table-layout:fixed; border-collapse:separate; border-spacing:0; overflow:hidden; border:1px solid #27344A; border-radius:20px; background:#0D111B; font-family:-apple-system,BlinkMacSystemFont,"Segoe UI",Roboto,sans-serif; font-size:14px; color:#BFBDBD; }

.swift15-table th { padding:13px 8px; text-align:center; font-weight:700; color:#AEB5C7; background:#0D111B; border-right:1px solid #27344A; border-bottom:1px solid #27344A; }

.swift15-table td { padding:14px 8px; text-align:center; color:#BFBDBD; background:#0D111B; border-right:1px solid #27344A; border-bottom:1px solid #27344A; vertical-align:middle; overflow-wrap:break-word; }

.swift15-table tr > :last-child { border-right:0; }

.swift15-table tbody tr:last-child td { border-bottom:0; }

.swift15-table .benchmark-heading { color:#B7BDCD; background:#0D111B; border-bottom:3px solid #7D45B5; }

.swift15-table .score-heading { color:#F0C5FF; background:#52239E; border-bottom:3px solid #7D45B5; }

.swift15-table .tokens-heading, .swift15-table .median-heading { color:#D4E8FF; background:#304FC2; border-bottom:3px solid #5687E6; }

.swift15-table .change-heading { color:#61B9FF; background:#101D2D; border-bottom:3px solid #5687E6; }

.swift15-table .change { color:#61B9FF; background:#101D2D; font-weight:700; }

.swift15-table .gain { color:#DDA8FF; background:#171127; font-weight:700; }

.swift15-table .model { padding-left:18px; text-align:left; color:#FFFFFF; font-weight:600; }

.swift15-table strong { color:#FFFFFF; }

.swift15-table .section { padding:12px 18px; text-align:left; color:#B489FF; background:#2A2541; font-weight:700; letter-spacing:.08em; text-transform:uppercase; border-top:1px solid #3A3159; border-bottom:1px solid #3A3159; }

.swift15-table .swift15 { background:#171127; }

.swift15-table thead tr:nth-child(2) .swift15 { color:#D3A0FF; }

.swift15-table .detail { color:#8C94A8; font-size:12px; font-weight:500; }

.swift15-table.swift15-compact { display:table !important; width:100% !important; table-layout:fixed !important; }

@media (max-width: 640px) {

.swift15-table { display:block !important; width:100% !important; max-width:100%; overflow-x:auto !important; -webkit-overflow-scrolling:touch; table-layout:auto !important; }

.swift15-table th, .swift15-table td { min-width:104px; }

.swift15-table th:first-child, .swift15-table td:first-child { min-width:170px; }

}

</style>

<table class="swift15-table">

<thead>

<tr>

<th class="benchmark-heading" rowspan="2" style="width:32%;text-align:left;padding-left:18px;vertical-align:bottom;">Benchmark</th>

<th class="score-heading" colspan="2">Final score</th>

<th class="tokens-heading" colspan="3">Mean tokens</th>

<th class="median-heading">Median tokens</th>

</tr>

<tr>

<th>Qwen3.8</th>

<th class="swift15">Swift 1.5</th>

<th>Qwen3.8</th>

<th class="swift15">Swift 1.5</th>

<th class="change-heading">Reduction</th>

<th class="change-heading">Reduction</th>

</tr>

</thead>

<tbody>

<tr><td class="section" colspan="7">General reasoning</td></tr>

<tr><td class="model">GPQA-Diamond</td><td>88.28%</td><td class="swift15"><strong>88.59%</strong></td><td>15,014</td><td class="swift15">8,717</td><td class="change">&#8595; 41.9%</td><td class="change">&#8595; 58.5%</td></tr>

<tr><td class="model">C-Eval</td><td>90.00%</td><td class="swift15"><strong>90.92%</strong></td><td>1,492</td><td class="swift15">819</td><td class="change">&#8595; 45.1%</td><td class="change">&#8595; 16.9%</td></tr>

<tr><td class="model">IFBench</td><td><strong>73.53%</strong></td><td class="swift15">72.07%</td><td>8,052</td><td class="swift15">4,955</td><td class="change">&#8595; 38.5%</td><td class="change">&#8595; 47.3%</td></tr>

<tr><td class="model">ERQA</td><td><strong>67.45%</strong></td><td class="swift15">65.40%</td><td>4,137</td><td class="swift15">1,906</td><td class="change">&#8595; 53.9%</td><td class="change">&#8595; 56.2%</td></tr>

<tr><td class="section" colspan="7">Mathematics</td></tr>

<tr><td class="model">AIME 2026</td><td><strong>98.67%</strong></td><td class="swift15">96.00%</td><td>22,014</td><td class="swift15">13,203</td><td class="change">&#8595; 40.0%</td><td class="change">&#8595; 48.5%</td></tr>

<tr><td class="model">HMMT November 2025</td><td><strong>99.33%</strong></td><td class="swift15">97.33%</td><td>22,032</td><td class="swift15">14,957</td><td class="change">&#8595; 32.1%</td><td class="change">&#8595; 47.8%</td></tr>

<tr><td class="section" colspan="7">Coding</td></tr>

<tr><td class="model">LiveCodeBench v6</td><td>76.76%</td><td class="swift15"><strong>81.71%</strong></td><td>11,184</td><td class="swift15">8,448</td><td class="change">&#8595; 24.5%</td><td class="change">&#8595; 46.3%</td></tr>

<tr><td class="section" colspan="7">Agent tasks</td></tr>

<tr><td class="model">Terminal-Bench 2.1</td><td>69.21%</td><td class="swift15"><strong>72.13%</strong></td><td>52,265</td><td class="swift15">43,733</td><td class="change">&#8595; 16.3%</td><td class="change">&#8595; 0.1%</td></tr>

</tbody>

</table>

Scores are final five-repeat aggregates under matched evaluation protocols. Mean-token

columns report reasoning tokens per trial; Terminal-Bench sums reasoning across agent calls.

<details>

<summary><strong>Benchmark methodology and reproduction settings</strong></summary>

<p style="font-size:13px;line-height:1.5;margin:8px 0;"><strong>Serving:</strong> BF16 · vLLM 0.27.1 · Qwen3 parser · context 262,144 · thinking xhigh.<br>

<strong>Sampling:</strong> temperature 1.0 · top_p 0.95 · top_k 20 · min_p 0 · presence_penalty 0 · repetition_penalty 1.<br>

<strong>Benchmarks:</strong> averages over five seeds (0–4) per model; five trials per task for Terminal-Bench, base and Swift 1.5 served at context 131,072 on the same Harbor build.</p>

<table style="display:table;width:100%;border-collapse:collapse;font-size:13px;line-height:1.3;margin:8px 0;">

<thead><tr><th style="padding:4px 8px;text-align:left;">Benchmark</th><th style="padding:4px 8px;text-align:right;">Output cap</th></tr></thead>

<tbody>

<tr><td style="padding:3px 8px;">GPQA-Diamond</td><td style="padding:3px 8px;text-align:right;">100,000</td></tr>

<tr><td style="padding:3px 8px;">C-Eval</td><td style="padding:3px 8px;text-align:right;">16,384</td></tr>

<tr><td style="padding:3px 8px;">IFBench</td><td style="padding:3px 8px;text-align:right;">81,920</td></tr>

<tr><td style="padding:3px 8px;">ERQA</td><td style="padding:3px 8px;text-align:right;">100,000</td></tr>

<tr><td style="padding:3px 8px;">AIME 2026</td><td style="padding:3px 8px;text-align:right;">250,000</td></tr>

<tr><td style="padding:3px 8px;">HMMT November 2025</td><td style="padding:3px 8px;text-align:right;">250,000</td></tr>

<tr><td style="padding:3px 8px;">LiveCodeBench v6</td><td style="padding:3px 8px;text-align:right;">32,768</td></tr>

<tr><td style="padding:3px 8px;">Terminal-Bench 2.1</td><td style="padding:3px 8px;text-align:right;">Agent/task limits</td></tr>

</tbody>

</table>

</details>

Efficiency across reasoning efforts

Qwen3.8's reasoning_effort setting lets users choose how much the model thinks.

For Swift 1.5 to be useful across these settings, it needs to reduce thinking while

keeping accuracy close to the base. We therefore tested xhigh, medium, and low:

thinking-token savings persist at every level.

<table class="swift15-table swift15-compact">

<thead>

<tr>

<th class="benchmark-heading" style="width:34%;text-align:left;padding-left:18px;">Reasoning effort</th>

<th class="score-heading" style="width:22%;">Qwen3.8</th>

<th class="score-heading" style="width:22%;">Swift 1.5</th>

<th class="tokens-heading" style="width:22%;">Mean thinking reduction</th>

</tr>

</thead>

<tbody>

<tr><td class="model">Xhigh</td><td>88.28%</td><td class="swift15"><strong>88.59%</strong></td><td class="change">&#8595; 41.9%</td></tr>

<tr><td class="model">Medium</td><td><strong>84.14%</strong></td><td class="swift15">82.22%</td><td class="change">&#8595; 24.8%</td></tr>

<tr><td class="model">Low</td><td>84.04%</td><td class="swift15"><strong>84.85%</strong></td><td class="change">&#8595; 28.7%</td></tr>

</tbody>

</table>

At low, Swift 1.5 scores above the base while using about 29% fewer thinking tokens.

Quantized Swift 1.5 models

| Format | Repository | Runtime |

| --- | --- | --- |

| GGUF | Swift-1.5-Qwen3.8-27B-GGUF | llama.cpp |

| GSQ-RCO GGUF (compact 2–3 bit) | Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF | llama.cpp |

| AWQ INT4 (W4A16) | Swift-1.5-Qwen3.8-27b-W4A16-AWQ | vLLM (compressed-tensors) |

| AutoRound INT4 (W4A16) | Swift-1.5-Qwen3.8-27b-W4A16-AutoRound | vLLM (auto-round) |

| AWQ + GPTQ INT4 (W4A16) | Swift-1.5-Qwen3.8-27b-INT4 | vLLM (compressed-tensors) |

| NVFP4 | Swift-1.5-Qwen3.8-27b-NVFP4 | NVIDIA Blackwell |

| AMD Quark FP8 (W8A8) | Swift-1.5-Qwen3.8-27b-Quark-FP8-dynamic-AMD | AMD Quark |

| MLX 5-bit | Swift-1.5-5bit-MLX | Apple MLX |

| MLX 4-bit | Swift-1.5-4bit-MLX | Apple MLX |

| MLX 3-bit (text only) | Swift-1.5-3bit-MLX-TextOnly | Apple MLX |

These results evaluate the merged Swift 1.5 checkpoint and three INT4 exports on

GPQA-Diamond (198 questions), IFBench (300 prompts), and AIME 2026 (30 problems).

Each model completed the full datasets with **one sample per prompt, seed 0, and

zero request errors**. This is a single-seed evaluation, separate from the

five-repeat BF16 release results above.

The Qwen-base columns use the saved seed/sample 0 runs.

Quantization recipes and serving settings differ from the new Swift 1.5 runs,

so these are reference comparisons rather than a controlled measurement of

the Swift adaptation. Token reductions below are recomputed from those same

reference samples.

<table class="swift15-table" style="display:table;width:100%;table-layout:fixed;">

<thead><tr>

<th class="benchmark-heading" style="width:32%;text-align:left;padding-left:18px;">Benchmark / Swift 1.5 quantization</th>

<th class="score-heading" style="width:16%;">Qwen base<br>accuracy</th>

<th class="score-heading" style="width:16%;">Swift 1.5 quant<br>accuracy</th>

<th class="tokens-heading" style="width:18%;">Mean token reduction</th>

<th class="median-heading" style="width:18%;">Median token reduction</th>

</tr></thead>

<tbody>

<tr><td class="model">GPQA-Diamond<br><span class="detail"><a href="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AWQ">AWQ INT4</a></span></td><td>86.36%</td><td class="swift15">88.38%</td><td class="change">&#8595; 51.5%</td><td class="change">&#8595; 64.4%</td></tr>

<tr><td class="model">GPQA-Diamond<br><span class="detail"><a href="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AutoRound">AutoRound INT4</a></span></td><td>86.36%</td><td class="swift15">89.39%</td><td class="change">&#8595; 50.5%</td><td class="change">&#8595; 57.8%</td></tr>

<tr><td class="model">GPQA-Diamond<br><span class="detail"><a href="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b-INT4">AWQ + GPTQ INT4</a></span></td><td>86.36%</td><td class="swift15">90.91%</td><td class="change">&#8595; 45.8%</td><td class="change">&#8595; 64.4%</td></tr>

<tr><td class="model">IFBench<br><span class="detail"><a href="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AWQ">AWQ INT4</a></span></td><td>72.00%</td><td class="swift15">72.00%</td><td class="change">&#8595; 36.9%</td><td class="change">&#8595; 49.3%</td></tr>

<tr><td class="model">IFBench<br><span class="detail"><a href="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AutoRound">AutoRound INT4</a></span></td><td>72.00%</td><td class="swift15">69.33%</td><td class="change">&#8595; 29.3%</td><td class="change">&#8595; 39.2%</td></tr>

<tr><td class="model">IFBench<br><span class="detail"><a href="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b-INT4">AWQ + GPTQ INT4</a></span></td><td>72.00%</td><td class="swift15">70.00%</td><td class="change">&#8595; 31.8%</td><td class="change">&#8595; 52.7%</td></tr>

<tr><td class="model">AIME 2026<br><span class="detail"><a href="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AWQ">AWQ INT4</a></span></td><td>70.00%</td><td class="swift15">86.67%</td><td class="change">&#8595; 29.2%</td><td class="change">&#8595; 36.2%</td></tr>

<tr><td class="model">AIME 2026<br><span class="detail"><a href="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AutoRound">AutoRound INT4</a></span></td><td>76.67%</td><td class="swift15">83.33%</td><td class="change">&#8595; 17.7%</td><td class="change">&#8595; 32.4%</td></tr>

<tr><td class="model">AIME 2026<br><span class="detail"><a href="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b-INT4">AWQ + GPTQ INT4</a></span></td><td>76.67%</td><td class="swift15">83.33%</td><td class="change">&#8595; 22.4%</td><td class="change">&#8595; 34.0%</td></tr>

</tbody>

</table>

AIME scoring: truncated responses count as incorrect for both columns.

The AMD Quark INT4 and FP8 exports have separate sanity evaluations; completed

results on these three reasoning benchmarks are not available for them.

<details>

<summary><strong>Quantized evaluation settings and BF16 reference</strong></summary>

Serving: vLLM 0.29.0, tensor parallelism 1, eager execution, BF16 activations,

context 131,072, template-default thinking without an effort override. The

AWQ + GPTQ export uses FP8 KV cache; BF16, AWQ, and AutoRound use auto KV dtype.

Sampling: temperature 1, top-p 0.95, top-k 20, min-p 0, presence penalty 0,

repetition penalty 1, seed 0. Output caps: GPQA 100,000, IFBench 81,920,

AIME 32,768. IFBench uses official strict prompt-level scoring.

GPQA token counts cover re-tokenized reasoning; IFBench and AIME count the

full generated response. Statistics include all responses, including truncations; medians

use the midpoint of the two central values when the sample count is even.

Saved Qwen references: W4A16 for GPQA and IFBench; Qwen AWQ for the

AWQ AIME row; Qwen W4A16 for the AutoRound and AWQ + GPTQ AIME rows. The

latter is a W4A16 reference for AutoRound, not an AutoRound base run.

The new runs do not reproduce the original software stack.

The fresh Swift 1.5 BF16 reference and all quantized exports scored as follows

under this single-seed protocol:

| Model | GPQA-Diamond | IFBench strict | AIME 2026 |

| --- | ---: | ---: | ---: |

| Swift 1.5 BF16 | 91.41% | 72.00% | 86.67% |

| AWQ INT4 | 88.38% | 72.00% | 86.67% |

| AutoRound INT4 | 89.39% | 69.33% | 83.33% |

| AWQ + GPTQ INT4 | 90.91% | 70.00% | 83.33% |

Truncation counts are recorded in the linked evaluation data.

These single-seed results do not establish

quality parity or replace the broader multi-seed evaluation.

Verified counts, token statistics, settings, and evidence hashes.

</details>

GGUF quantizations

<table class="swift15-table" style="display:table;width:100%;table-layout:fixed;">

<thead><tr>

<th class="benchmark-heading" style="width:20%;text-align:left;padding-left:18px;white-space:normal;">File</th>

<th class="score-heading" style="width:12%;white-space:normal;">Size</th>

<th class="score-heading" style="width:17%;white-space:normal;">KLD wikitext @512</th>

<th class="tokens-heading" style="width:17%;white-space:normal;">KLD wikitext @32k</th>

<th class="tokens-heading" style="width:17%;white-space:normal;">99% KLD @32k</th>

<th class="median-heading" style="width:17%;white-space:normal;">Top-p @32k</th>

</tr></thead>

<tbody>

<tr><td class="model">Q8_0</td><td>29.0 GB</td><td class="swift15">0.0008</td><td class="change">0.0006</td><td class="change">0.005</td><td>98.85%</td></tr>

<tr><td class="model">Q6_K_L</td><td>25.0 GB</td><td>0.0015</td><td class="change">0.0014</td><td class="change">0.010</td><td>98.16%</td></tr>

<tr><td class="model">Q6_K</td><td>23.9 GB</td><td>0.0018</td><td class="change">0.0016</td><td class="change">0.014</td><td>98.30%</td></tr>

<tr><td class="model">Q6_K_S</td><td>22.9 GB</td><td>0.0020</td><td class="change">0.0016</td><td class="change">0.014</td><td>98.24%</td></tr>

<tr><td class="model">Q5_K_M</td><td>20.9 GB</td><td>0.0052</td><td class="change">0.0061</td><td class="change">0.050</td><td>96.92%</td></tr>

<tr><td class="model">Q5_K_S</td><td>19.6 GB</td><td>0.0060</td><td class="change">0.0069</td><td class="change">0.058</td><td>96.91%</td></tr>

<tr><td class="model">Q4_K_L</td><td>18.8 GB</td><td>0.0106</td><td class="change">0.0103</td><td class="change">0.105</td><td>95.79%</td></tr>

<tr><td class="model">Q4_K_M</td><td>17.4 GB</td><td>0.0137</td><td class="change">0.0134</td><td class="change">0.163</td><td>95.03%</td></tr>

<tr><td class="model">IQ4_NL</td><td>17.4 GB</td><td>0.0152</td><td class="change">0.0140</td><td class="change">0.175</td><td>95.39%</td></tr>

<tr><td class="model">Q4_K_S</td><td>16.4 GB</td><td>0.0164</td><td class="change">0.0154</td><td class="change">0.175</td><td>94.84%</td></tr>

<tr><td class="model">IQ4_XS</td><td>15.5 GB</td><td>0.0179</td><td class="change">0.0173</td><td class="change">0.187</td><td>94.96%</td></tr>

<tr><td class="model">IQ3_M</td><td>14.9 GB</td><td>0.0410</td><td class="change">0.0380</td><td class="change">0.409</td><td>91.83%</td></tr>

<tr><td class="model">Q3_K_L</td><td>14.1 GB</td><td>0.0442</td><td class="change">0.0410</td><td class="change">0.415</td><td>91.63%</td></tr>

<tr><td class="model">Q3_K_M</td><td>13.4 GB</td><td>0.0570</td><td class="change">0.0562</td><td class="change">0.614</td><td>90.30%</td></tr>

<tr><td class="model">IQ3_XS</td><td>12.8 GB</td><td>0.0583</td><td class="change">0.0885</td><td class="change">1.130</td><td>88.54%</td></tr>

<tr><td class="model">Q3_K_S</td><td>12.7 GB</td><td>0.0648</td><td class="change">0.0658</td><td class="change">0.712</td><td>89.53%</td></tr>

<tr><td class="model">IQ3_XXS</td><td>12.3 GB</td><td>0.0742</td><td class="change">0.0844</td><td class="change">0.996</td><td>88.70%</td></tr>

<tr><td class="model">Q2_K</td><td>10.8 GB</td><td>0.1655</td><td class="change">0.1546</td><td class="change">1.728</td><td>84.00%</td></tr>

<tr><td class="model">IQ2_M</td><td>10.5 GB</td><td>0.1523</td><td class="change">0.1493</td><td class="change">1.568</td><td>84.17%</td></tr>

<tr><td class="model">IQ2_S</td><td>9.7 GB</td><td>0.2095</td><td class="change">0.2589</td><td class="change">3.024</td><td>80.58%</td></tr>

<tr><td class="model">IQ2_XS</td><td>9.1 GB</td><td>0.2433</td><td class="change">0.2622</td><td class="change">2.951</td><td>79.99%</td></tr>

<tr><td class="model">IQ2_XXS</td><td>8.9 GB</td><td>0.2866</td><td class="change">0.2769</td><td class="change">2.902</td><td>78.48%</td></tr>

</tbody>

</table>

Mean KL divergence against the Swift 1.5 BF16 source, lower is better. wikitext @512 is wikitext-2

test, 100 windows of 512 tokens. wikitext @32k is wikitext-2 train, 16 windows of 32,768 tokens, scoring

only the last 512 tokens of each window, so every scored token sees at least 32k tokens of context.

99% KLD is the 99th-percentile divergence on the same 32k run. Top-p is top-token agreement with BF16

on the 32k run.

Long context costs very little on this release: for every tier from Q8_0 through Q4_K_S, the 32k mean

is within 10% of the 512-token value (Q5_K_M and Q5_K_S rise about 15%), and Q4_K_M keeps the same

top token as BF16 on 95% of positions at 32k. The pick below follows the 99th-percentile tail at 32k:

0.163 for Q4_K_M, 0.050 for Q5_K_M, 0.014 for Q6_K, 0.005 for Q8_0.

<table class="swift15-table" style="display:table;width:100%;table-layout:fixed;">

<thead><tr>

<th class="benchmark-heading" style="width:55%;text-align:left;padding-left:18px;white-space:normal;">Use case</th>

<th class="score-heading" style="width:45%;white-space:normal;">Pick</th>

</tr></thead>

<tbody>

<tr><td class="model">24 GB cards, everyday use</td><td class="swift15">Q4_K_M</td></tr>

<tr><td class="model">Long agentic runs, strict tool-call formatting</td><td class="swift15">Q6_K or higher</td></tr>

<tr><td class="model">Maximum fidelity</td><td class="swift15">Q8_0</td></tr>

</tbody>

</table>

<details>

<summary><strong>Recipe</strong></summary>

All 22 tiers were built with llama.cpp commit 6f41ac5 from a BF16 conversion of the published Swift 1.5

safetensors. They reuse the importance matrix and the per-tensor type layouts (--tensor-type-file) that

bartowski computed for Swift 1.0 with his

quantization-config; Swift 1.5 has the same

architecture and tensor shapes. Every file was checked against BF16 on the harness above.

</details>

License and access

Swift 1.5 is a derivative of Qwen3.8-27B

(Copyright 2026 Alibaba Cloud, Apache License 2.0). UkisAI's contribution, including the adapted

weights, is licensed under the Swift Open License v1.0. See NOTICE for the change notice and attribution details.

Personal, research, educational, evaluation, and commercial use are free for individuals

and organizations with gross annual revenue, including affiliates, of up to US$1,000,000.

Above that threshold, commercial use requires a separate Swift Enterprise License.

Contact UkisAI for terms.

Nothing in the Swift Open License limits rights in Qwen3.8-27B itself under Apache 2.0.

Citation

~~~bibtex

@misc{swift-1.5-qwen3.8-27b,

title = {Swift 1.5 Qwen3.8-27B},

author = {UkisAI},

year = {2026},

url = {https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b}

}

~~~

Acknowledgements

We acknowledge the NVIDIA Innovation Lab,

Amazon Web Services, and

Google Cloud for providing compute credits and

infrastructure support for Swift's development, training, and evaluation.

Run ukisai/Swift-1.5-Qwen3.8-27B-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models