WineryLabs/Winery-Qwen3.5-2B-GrandCru-GGUF overview
<p align="center" <a href="https://huggingface.co/WineryLabs" <img src="https://winery api zandy.zocomputer.io/card/m/2b.svg" alt="Winery Qwen3.5 2B GrandCru" …
Runs locally from ~1.87 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Winery-Qwen3.5-2B-GrandCru-Q8_0.gguf | GGUF | Q8_0 | 1.87 GB | Download |
Model Details
| Model ID | WineryLabs/Winery-Qwen3.5-2B-GrandCru-GGUF |
|---|---|
| Author | WineryLabs |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | Qwen/Qwen3.5-2B,DavidAU/Qwen3.5-2B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKING,DavidAU/Qwen3.5-2B-Polaris-HighIQ-Thinking-Compact,Goekdeniz-Guelmez/Josiefied-Qwen3.5-2B-gabliterated-v1 |
| Last modified | 2026-10-03T09:05:14.000Z |
Model README
---
license: apache-2.0
base_model:
- Qwen/Qwen3.5-2B
- DavidAU/Qwen3.5-2B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKING
- DavidAU/Qwen3.5-2B-Polaris-HighIQ-Thinking-Compact
- Goekdeniz-Guelmez/Josiefied-Qwen3.5-2B-gabliterated-v1
base_model_relation: merge
library_name: gguf
pipeline_tag: text-generation
tags: [gguf, merge, qwen3.5, winery, llama.cpp, model-soup]
---
<p align="center"><a href="https://huggingface.co/WineryLabs"><img src="https://winery-api-zandy.zocomputer.io/card/m/2b.svg" alt="Winery Qwen3.5-2B-GrandCru" width="100%"/></a></p>
🍷 Winery Qwen3.5 2B · Grand Cru
A score-weighted model soup of the three strongest Qwen3.5-2B fine-tunes we tested, blended with the Winery fusion compiler. It's 1.88B params and 2.0 GB at Q8_0, so it runs well on phones and laptops.
Recipe
Linear merge. Each donor's weight is its own benchmark average:
| Donor | Weight (its avg) |
|---|---|
| DavidAU/Qwen3.5-2B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKING | 70.3 |
| DavidAU/Qwen3.5-2B-Polaris-HighIQ-Thinking-Compact | 68.3 |
| Goekdeniz-Guelmez/Josiefied-Qwen3.5-2B-gabliterated-v1 | 67.2 |
name Winery Grand Cru
out q8_0
default linear 70.3,68.3,67.2
Scores (local quick bench, 0-shot chat, thinking off)
MMLU 200 / ARC-Challenge 150 / GSM8K 60 questions, run through llama-server.
| Model | MMLU | ARC-C | GSM8K | Avg |
|---|---|---|---|---|
| Winery Grand Cru (this) | 51.5 | 83.3 | 76.7 | 70.5 |
| Best donor (DavidAU HERETIC) | 54.5 | 81.3 | 75.0 | 70.3 |
| Qwen/Qwen3.5-2B (stock) | 51.5 | 78.7 | 58.3 | 62.8 |
These are small-sample numbers from our own harness, so treat them as directional only. They aren't leaderboard results.
Use
Works in llama.cpp (recent builds with Qwen3.5 support), Jan, LM Studio and the Winery app.
llama-cli -m Winery-Qwen3.5-2B-GrandCru-Q8_0.gguf -cnv
⚠️ Two of the donors are abliterated/uncensored, so this model refuses less than stock Qwen. You're responsible for how you use it.
Run WineryLabs/Winery-Qwen3.5-2B-GrandCru-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models