WineryLabs/Winery-Qwen3.5-9B-GrandCru-GGUF overview
<p align="center" <a href="https://huggingface.co/WineryLabs" <img src="https://winery api zandy.zocomputer.io/card/m/9b.svg" alt="Winery Qwen3.5 9B GrandCru" …
Runs locally from ~8.87 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Winery-Qwen3.5-9B-GrandCru-Q8_0.gguf | GGUF | Q8_0 | 8.87 GB | Download |
Model Details
| Model ID | WineryLabs/Winery-Qwen3.5-9B-GrandCru-GGUF |
|---|---|
| Author | WineryLabs |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | Qwen/Qwen3.5-9B,XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B,Jackrong/Qwen3.5-9B-Claude-4.6-Opus-Reasoning-Distilled-v2,DavidAU/Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSORED |
| Last modified | 2026-10-03T09:05:11.000Z |
Model README
---
license: apache-2.0
base_model:
- Qwen/Qwen3.5-9B
- XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
- Jackrong/Qwen3.5-9B-Claude-4.6-Opus-Reasoning-Distilled-v2
- DavidAU/Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSORED
base_model_relation: merge
library_name: gguf
pipeline_tag: text-generation
tags: [gguf, merge, qwen3.5, winery, llama.cpp, model-soup]
---
<p align="center"><a href="https://huggingface.co/WineryLabs"><img src="https://winery-api-zandy.zocomputer.io/card/m/9b.svg" alt="Winery Qwen3.5-9B-GrandCru" width="100%"/></a></p>
🍷 Winery Qwen3.5 9B · Grand Cru
A score-weighted model soup of the three strongest Qwen3.5-9B fine-tunes out of the 10 we benchmarked, blended with the Winery fusion compiler. It's 9.5 GB at Q8_0.
Recipe
Linear merge. Each donor's weight is its own benchmark average:
| Donor | Weight (its avg) |
|---|---|
| XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B | 84.9 |
| Jackrong/Qwen3.5-9B-Claude-4.6-Opus-Reasoning-Distilled-v2 | 83.4 |
| DavidAU/Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSORED | 83.2 |
name Winery 9B Grand Cru
out q8_0
default linear 84.9,83.4,83.2
Scores (local quick bench, 0-shot chat, thinking off)
MMLU 200 / ARC-Challenge 150 / GSM8K 60 questions, run through llama-server.
| Model | MMLU | ARC-C | GSM8K | Avg |
|---|---|---|---|---|
| Winery 9B Grand Cru (this) | 70.0 | 96.7 | 90.0 | 85.6 |
| Best donor (MiMo-V2.6 Distill) | 69.5 | 95.3 | 90.0 | 84.9 |
| Qwen/Qwen3.5-9B (stock) | 69.5 | 94.7 | 66.7 | 77.0 |
These are small-sample numbers from our own harness, so treat them as directional only. They aren't leaderboard results.
Use
Works in llama.cpp (recent builds with Qwen3.5 support), Jan, LM Studio and the Winery app. Tested in Jan.
llama-cli -m Winery-Qwen3.5-9B-GrandCru-Q8_0.gguf -cnv
⚠️ One donor is an uncensored/abliterated tune, so this model refuses less than stock Qwen. You're responsible for how you use it.
Run WineryLabs/Winery-Qwen3.5-9B-GrandCru-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models