GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF overview

<div align="center" <a href="https://ukisai.com" <img src="ukisai banner.png" alt="UkisAI" style="width:100%;height:auto;" / </a <p <a href="https://ukisai.com…

ggufllama.cppqwen3_8moegsqrcoreasoningefficient-thinkingtoken-efficientpost-trainingimage-text-to-textarxiv:2604.18556arxiv:2605.00649base_model:ukisai/Swift1.5-Qwen3.8-Flash-Nextbase_model:quantized:ukisai/Swift1.5-Qwen3.8-Flash-Nextlicense:otherendpoints_compatibleregion:usimatrixconversational

Runs locally from ~553.2 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
222,601
Likes
137
Pipeline
image-text-to-text
Author

Repository Files & Downloads

10 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ2_XS-00001-of-00002.ggufGGUFIQ2_XS37.06 GBDownload
Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ2_XS-00002-of-00002.ggufGGUFIQ2_XS26.42 GBDownload
Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.ggufGGUFIQ3_S41.82 GBDownload
Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00002-of-00002.ggufGGUFIQ3_S36.18 GBDownload
Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS-00001-of-00002.ggufGGUFIQ3_XXS37.05 GBDownload
Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS-00002-of-00002.ggufGGUFIQ3_XXS33.70 GBDownload
Swift-Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-00001-of-00002.ggufGGUFQ2_037.07 GBDownload
Swift-Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-00002-of-00002.ggufGGUFQ2_024.91 GBDownload
imatrix-swiftfn-v1mix.ggufGGUFGGUF553.2 MBDownload
mmproj-Swift-Qwen3.8-Flash-Next-BF16.ggufGGUFBF16865.5 MBDownload

Model Details

Model IDukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF
Authorukisai
Pipelineimage-text-to-text
Licenseother
Base modelukisai/Swift-Qwen3.8-Flash-Next
Last modified2026-10-07T19:36:20.000Z

Model README

---

license: other

license_name: swift-open-license-1.0

license_link: https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF/blob/main/LICENSE

library_name: gguf

pipeline_tag: image-text-to-text

base_model: ukisai/Swift-Qwen3.8-Flash-Next

base_model_relation: quantized

tags:

  • gguf
  • llama.cpp
  • qwen3_8
  • moe
  • gsq
  • rco
  • reasoning
  • efficient-thinking
  • token-efficient
  • post-training

---

<div align="center">

<a href="https://ukisai.com"><img src="ukisai-banner.png" alt="UkisAI" style="width:100%;height:auto;" /></a>

<p><a href="https://ukisai.com"><b>Website</b></a> &bull; <a href="https://ukisai.com/products/swift"><b>Learn more</b></a> &bull; <a href="https://huggingface.co/ukisai/Swift-Qwen3.8-Flash-Next"><b>BF16 model</b></a> &bull; <a href="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-Flash-Next-GGUF"><b>Standard GGUFs</b></a> &bull; <a href="#evaluation"><b>Evaluation</b></a> &bull; <a href="#license-and-access"><b>Enterprise licensing</b></a></p>

</div>

Swift 1.5 Qwen3.8-Flash-Next · GSQ-RCO

Mixed-precision GGUF quantizations of Swift Flash Next, with Swift-specific GSQ refinement and reused ISTA GSQ-RCO per-tensor allocation profiles.

Swift Flash Next is UkisAI's reasoning-efficient derivative of Qwen3.8-Flash-Next. Its post-training targets shorter reasoning traces and coding, agentic and long-horizon tasks. See the original model card for model-level benchmarks and training details. Those benchmarks are separate from the quantization measurements below.

Swift 1.5 Flash-Next uses 63.4% fewer thinking tokens, with a 1.8x speed up while keeping the accuracy loss <1% vs base on xhigh.

Available quantizations

Each tier contains two GGUF shards. Download both files into the same directory and load shard 1; llama.cpp locates shard 2 automatically. Sizes are decimal GB and exclude runtime context/cache memory. Tier names describe mixed-precision allocation profiles.

| Tier | Combined GGUF size | Shards | Development KLD ↓ |

| --- | ---: | --- | ---: |

| IQ3_S | 83.74 GB | 1 · 2 | 0.144674 |

| IQ3_XXS | 75.97 GB | 1 · 2 | 0.240139 |

| IQ2_XS | 68.15 GB | 1 · 2 | 0.341275 |

| Q2_0 (experimental) | 66.55 GB | 1 · 2 | 0.424350 |

A BF16 vision projector is included separately (0.91 GB). The evaluation below covers text inference; it does not measure vision-task accuracy.

Exact model and shard identities are recorded in release-manifest.json and SHA256SUMS.

Original evaluated, unsplit GGUF files can also be reconstructed byte for byte using the exact-source recovery files and script. These small files preserve headers and padding; normal inference needs only the two model shards. All four reconstructed source hashes were verified before release.

Evaluation

KLD measures divergence from the corresponding BF16 model's next-token distribution; lower is better. Swift quants are measured against Swift Flash Next BF16. Development measurements use 100 chunks at a 512-token context, and informed refinement.

Reporting prose, code and math sets use 100 chunks each; German, French, Spanish and Chinese use 25 chunks each, all at context 512. The seven original reporting sets were reused. IQ3_S, IQ2_XS and Q2_0 additionally have results on a preregistered fresh English C4 shard; no matching fresh result is available for IQ3_XXS.

| Reporting text | IQ3_S | IQ3_XXS | IQ2_XS | Q2_0 (experimental) |

| --- | ---: | ---: | ---: | ---: |

| English prose | 0.062556 | 0.116077 | 0.188117 | 0.234242 |

| Fresh English sample | 0.067801 | — | 0.186271 | 0.236228 |

| CodeParrot code | 0.069016 | 0.118707 | 0.174294 | 0.235256 |

| GSM8K math text | 0.054960 | 0.086821 | 0.120991 | 0.149809 |

| German | 0.063634 | 0.109100 | 0.166814 | 0.219444 |

| French | 0.076760 | 0.133624 | 0.213379 | 0.300467 |

| Spanish | 0.042848 | 0.073681 | 0.119114 | 0.148757 |

| Chinese | 0.095915 | 0.174102 | 0.264923 | 0.385924 |

IQ2_XS is the standout. It has lower KLD than ISTA-DASLab's own GSQ-RCO IQ2_XS on seven of eight reporting sets, by 5–11% (math text −8.9%, Chinese −10.5%), and improves on its Swift starting quant in every domain, by 8–17%. IQ3_XXS improves on its Swift starting quant in six of seven domains, by 3–16%, and has lower KLD than ISTA's IQ3_XXS on math text (−7.8%) and Chinese (−5.1%); the other domains are within 1–4%.

> [!NOTE]

> Q2_0 is experimental. It improves on its Swift starting quant in five of seven domains, but has higher KLD than ISTA's Q2_0 on six of eight reporting sets. For a file of similar size, prefer IQ2_XS.

IQ3_S has the lowest KLD of the four tiers. On the development set its KLD is at parity with ISTA-DASLab's GSQ-RCO IQ3_S (0.144674 vs 0.146683, −1.4%). On the eight reporting sets it has lower KLD than ISTA's IQ3_S on 4 (Chinese −7.2%, math text −3.7%, Spanish −0.4%, code −0.1%) and higher KLD on 4 (French +8.8%, German +2.8%, fresh English C4 +0.6%, English prose +0.3%). It improves on its Swift starting quant in 8 of 8 domains, by 6–17%.

These comparisons use each model's own BF16 reference. They are not direct capability rankings, task-accuracy percentages or statistical-equivalence claims.

The full result table includes ISTA and starting-quant comparisons and reported error estimates. Evaluation metadata binds these results to the selected model identities and records limitations. Lexical overlap filtering does not prove semantic deduplication or absence of overfitting. These tests do not establish long-context quality.

Usage

Use a llama.cpp build supporting Qwen3.8-Flash-Next.

hf download ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF \
  --include "Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS-*.gguf" --local-dir .

llama-server \
  -m Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS-00001-of-00002.gguf \
  --jinja -fa on -ngl 99 -c 262144 \
  --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 \
  --presence-penalty 0.0 --repeat-penalty 1.0 --port 8000

Set context size and GPU offload to fit available memory. The example context setting is not a claim that these quantizations were evaluated at that length.

For image input, also download the projector and add --mmproj mmproj-Swift-Qwen3.8-Flash-Next-BF16.gguf to the server command:

hf download ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF mmproj-Swift-Qwen3.8-Flash-Next-BF16.gguf --local-dir .

Quantization procedure

  1. Reuse the corresponding ISTA GSQ-RCO allocation profile to construct a Swift starting quant.
  2. Apply Swift-specific refinement, including an attention/expert pass and a second expert pass, with native format checks.
  3. Freeze the evaluated model identity and package its tensors into two GGUF shards.
  4. Check that all 1,224 source tensors are present with no byte mismatches, and record shard hashes.

This release reuses allocation search results rather than claiming a new RCO search on Swift. For IQ3_S the second expert pass also learns IQ4_NL codes (the other tiers keep IQ4_NL codes fixed and learn only their scales), and its calibration set adds German and French Wikipedia and GitHub code (sources disjoint from the reporting texts). The recipe summary identifies the selected variants and records the segmented-pass RNG limitation for IQ2_XS and Q2_0. The Swift V1MIX importance matrix and its provenance are included.

Per-tensor allocation dumps are provided in tensor-allocation, tied to the unsplit model identities in the release manifest.

Methods and acknowledgements

GSQ and RCO were developed by the Deep Algorithms and Systems Lab at the Institute of Science and Technology Austria. This Swift adaptation is by UkisAI.

We acknowledge the Qwen team for the original model and ISTA-DASLab for the quantization methods and published allocations.

License and access

The Swift contribution is distributed under the Swift Open License v1.0. The original Qwen components retain the Qwen Community License 1.0. See NOTICE and the license texts for applicable terms.

Run ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models