GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

TheStageAI/Qwen3.5-9B-GGUF overview

<p align="center" <img src="./assets/thestage edge models header.png" width="100%" alt="TheStageAI Edge Models: the right model at every memory budget" </p <h1…

llama.cppggufquantizedmixed-precisionlocal-inferenceqwen3.5text-generationarxiv:2505.17595arxiv:2309.01885arxiv:2504.09629arxiv:2605.00649base_model:Qwen/Qwen3.5-9Bbase_model:quantized:Qwen/Qwen3.5-9Blicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~2.99 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation

Repository Files & Downloads

4 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3.5-9B-L-TS-Q8_0.ggufGGUFQ8_08.88 GBDownload
Qwen3.5-9B-M-TS-Q4_K_M.ggufGGUFQ4_K_M4.71 GBDownload
Qwen3.5-9B-S-TS-Q4_K_S.ggufGGUFQ4_K_S3.81 GBDownload
Qwen3.5-9B-XS-TS-Q3_K_S.ggufGGUFQ3_K_S2.99 GBDownload

Model Details

Model IDTheStageAI/Qwen3.5-9B-GGUF
AuthorTheStageAI
Pipelinetext-generation
Licenseapache-2.0
Base modelQwen/Qwen3.5-9B
Last modified2026-07-21T18:24:42.000Z

Model README

---

license: apache-2.0

base_model:

- Qwen/Qwen3.5-9B

base_model_relation: quantized

library_name: llama.cpp

pipeline_tag: text-generation

thumbnail: https://huggingface.co/TheStageAI/Qwen3.5-9B-GGUF/resolve/main/assets/thestage-edge-models-header.png

tags:

- gguf

- llama.cpp

- quantized

- mixed-precision

- local-inference

- qwen3.5

---

<p align="center">

<img src="./assets/thestage-edge-models-header.png" width="100%" alt="TheStageAI Edge Models: the right model at every memory budget">

</p>

<h1 align="center">Qwen3.5 9B</h1>

<p align="center">

Four GGUF checkpoints for llama.cpp, from 3.21 GB to 9.53 GB.

<br>

<strong>Start with M: 5.06 GB and ≈100% of the BF16 instruction-strict IFEval score.</strong>

</p>

<p align="center">

<a class="inline-block" href="https://github.com/TheStageAI/edge-lm"><img class="dark:hidden" src="./assets/cta-edge-lm-light.svg" width="146" height="42" alt="Explore edge-lm on GitHub"><img class="hidden dark:block" src="./assets/cta-edge-lm-dark.svg" width="146" height="42" alt="Explore edge-lm on GitHub"></a>&nbsp;

<a class="inline-block" href="https://docs.thestage.ai/"><img class="dark:hidden" src="./assets/cta-docs-light.svg" width="120" height="42" alt="Read TheStageAI documentation"><img class="hidden dark:block" src="./assets/cta-docs-dark.svg" width="120" height="42" alt="Read TheStageAI documentation"></a>&nbsp;

<a class="inline-block" href="https://app.thestage.ai/"><img class="dark:hidden" src="./assets/cta-platform-light.svg" width="146" height="42" alt="Open TheStageAI Platform"><img class="hidden dark:block" src="./assets/cta-platform-dark.svg" width="146" height="42" alt="Open TheStageAI Platform"></a>

</p>

Choose a checkpoint

| Tier | Size | Best for | File |

| --- | ---: | --- | --- |

| XS | 3.21 GB | Minimum footprint | Download |

| S | 4.09 GB | Compact | Download |

| M | 5.06 GB | Recommended | Download |

| L | 9.53 GB | High-precision Q8 | Download |

Other Qwen 3.5 sizes: <a href="https://huggingface.co/TheStageAI/Qwen3.5-0.8B-GGUF">0.8B</a> &nbsp;·&nbsp; <a href="https://huggingface.co/TheStageAI/Qwen3.5-2B-GGUF">2B</a> &nbsp;·&nbsp; <a href="https://huggingface.co/TheStageAI/Qwen3.5-4B-GGUF">4B</a>

Quickstart

llama-cli \
  --hf-repo TheStageAI/Qwen3.5-9B-GGUF \
  --hf-file Qwen3.5-9B-M-TS-Q4_K_M.gguf

Why we recommend M

At 5.06 GB, M matches the BF16 instruction-strict IFEval score within evaluation variation. On MMLU-Pro, it matches the BF16 score within evaluation variation. It uses 47% less disk than L.

| Tier | IFEval P / I (%) | MMLU-Pro (%) |

| --- | ---: | ---: |

| BF16 reference | 83.18 / 88.37 | 82.39 |

| L | 82.99 / 87.89 | 82.40 |

| M | 83.36 / 88.49 | 82.16 |

| S | 82.81 / 87.65 | 79.70 |

| XS | 79.30 / 85.13 | — |

P / I means prompt-strict / instruction-strict. IFEval uses deterministic non-thinking decoding; MMLU-Pro uses sampled long-form reasoning. Only complete model-level scores are reported. A dash means not reported.

> Reasoning: XS is intended for non-thinking use. Run it with --reasoning off. Choose S, M, or L for long-form reasoning.

<details>

<summary><b>Evaluation protocol</b></summary>

  • IFEval: 541 prompts, native chat template, enable_thinking=false, temperature 0.
  • MMLU-Pro: 12,032 questions, vLLM, 0-shot, native chat template qwen_mc_json_v1, enable_thinking=true; temperature=1, top_p=0.95, top_k=20, min_p=0, presence_penalty=1.5, frequency_penalty=0, repetition_penalty=1, seed=42; max_model_len=40960, max_new_tokens=32768; dataset revision b189ec765aa7ed75c8acfea42df31fdae71f97be.
  • The headline BF16 comparison uses instruction-strict IFEval; the raw scores are shown in the table.

In matched long-form reasoning diagnostics, XS produced longer trajectories and reached the 32,768-token output limit more often than S. The XS MMLU-Pro cell is marked for that reason.

</details>

How the checkpoints are built

All four checkpoints use the same production PTQ pipeline. The precision map is the only tier-specific part.

1. Fit native GGUF codes

Calibration activations produce a curvature objective weighted by true Fisher information for each quantized projection. We adapt the scale and minimum initialization from NeUQI to that objective, then solve integer codes on the target GGUF grid with a guarded cyclic coordinate-descent solver inspired by QuantEase. A final K-quant pass tunes stored scales and minima while keeping packed codes fixed.

2. Reconstruct the deployed trajectory

Layers are calibrated in execution order against activations from the already-quantized prefix. A dense reference path measures accumulated drift. Quantization Error Propagation (QEP) adds that drift to the next reconstruction target, so later layers optimize for the inputs they receive at inference time.

3. Allocate the byte budget

XS and S can choose Q2_K through Q8_0 for each quantizable group. The optimizer trades changes in the teacher distribution against the actual encoded byte cost, including scale and minimum metadata. ANNA provides constrained configuration search. RCO provides an exact-budget route (code). For this model, RCO selected both the XS and S precision maps. M and L keep the decoder qtypes at Q4_K_M and Q8_0, respectively, and use the same reconstruction and scale-tuning stages.

After schedule selection, PTQ runs again from the source weights. Each layer is then calibrated with the final upstream precision choices.

4. Align the full model

A short affine distillation pass tunes native FP16 scales and minima while qtypes, packed codes, dense weights, and tensor layouts stay fixed. The loss matches the teacher's next-token distribution without changing file size or runtime layout.

We load-test the shipping GGUF and evaluate it on a held-out set of 3,072 sequences with next-token KL. release-manifest.json records its SHA-256, downstream evaluation IDs, and tensor metadata. Final recommendations use complete-model benchmarks.

<details>

<summary><b>File details</b></summary>

| Tier | Hub selector | GGUF file type | Whole-file BPW |

| --- | --- | --- | ---: |

| XS | Q3_K_S | MOSTLY_Q2_K | 2.865 |

| S | Q4_K_S | MOSTLY_Q2_K | 3.659 |

| M | Q4_K_M | MOSTLY_Q4_K_M | 4.521 |

| L | Q8_0 | MOSTLY_Q8_0 | 8.518 |

The Hub selector controls sidebar grouping and download discovery. For XS and S it approximates the whole-file size class; release-manifest.json contains the exact tensor mix. M and L keep their decoder qtypes at Q4_K_M and Q8_0 within the same production PTQ pipeline.

Runtime memory also includes KV cache and buffers, which grow with context length.

</details>

TheStageAI deployment stack

These GGUF files target llama.cpp-compatible runtimes. edge-lm runs compressed MLX models on Macs and iPhones. ANNA searches compression configurations under size or compute constraints. The TheStageAI Platform and documentation cover compression, compilation, and serving workflows.

For a specific device, latency target, or memory budget, talk to our team.

Reproducibility

Citation

If you use this checkpoint, cite this release and follow the upstream model's citation guidance:

@misc{thestageai2026qwen3p59bgguf,
  author       = {{TheStageAI}},
  title        = {Qwen3.5 9B: TheStageAI GGUF Release},
  year         = {2026},
  month        = {jul},
  howpublished = {Hugging Face model release},
  url          = {https://huggingface.co/TheStageAI/Qwen3.5-9B-GGUF},
  note         = {XS, S, M, and L deployment tiers},
}

<details>

<summary><b>References</b></summary>

</details>

License

The checkpoint weights use the upstream model's Apache-2.0 license. llama.cpp and other runtime software keep their own licenses.

Run TheStageAI/Qwen3.5-9B-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models