GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF overview

<p align="center" <a href="https://huggingface.co/WineryLabs" <img src="https://winery api zandy.zocomputer.io/card/m/cuvee.svg" alt="Winery Qwen3.8 Flash Next…

ggufmergetask-arithmeticmoeqwen3.8flash-nextstratawineryabliteratedimatrixtext-generationbase_model:Qwen/Qwen3.8-Flash-Nextbase_model:merge:Qwen/Qwen3.8-Flash-Nextbase_model:huihui-ai/Huihui-Qwen3.8-Flash-Next-abliteratedbase_model:merge:huihui-ai/Huihui-Qwen3.8-Flash-Next-abliteratedbase_model:ukisai/Swift1.5-Qwen3.8-Flash-Nextbase_model:merge:ukisai/Swift1.5-Qwen3.8-Flash-Nextlicense:otherendpoints_compatibleregion:usconversational

Runs locally from ~4.09 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
1
Pipeline
text-generation

Repository Files & Downloads

9 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
GRAND/Winery-Cuvee-GRAND-00001-of-00003.ggufGGUFGGUF41.55 GBDownload
GRAND/Winery-Cuvee-GRAND-00002-of-00003.ggufGGUFGGUF41.69 GBDownload
GRAND/Winery-Cuvee-GRAND-00003-of-00003.ggufGGUFGGUF18.98 GBDownload
POCKET/Winery-Cuvee-POCKET-00001-of-00002.ggufGGUFGGUF41.87 GBDownload
POCKET/Winery-Cuvee-POCKET-00002-of-00002.ggufGGUFGGUF4.09 GBDownload
RESERVE/Winery-Cuvee-RESERVE-00001-of-00002.ggufGGUFGGUF41.87 GBDownload
RESERVE/Winery-Cuvee-RESERVE-00002-of-00002.ggufGGUFGGUF38.96 GBDownload
SLIM/Winery-Cuvee-SLIM-00001-of-00002.ggufGGUFGGUF41.85 GBDownload
SLIM/Winery-Cuvee-SLIM-00002-of-00002.ggufGGUFGGUF13.95 GBDownload

Model Details

Model IDWineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF
AuthorWineryLabs
Pipelinetext-generation
Licenseother
Base modelQwen/Qwen3.8-Flash-Next,ukisai/Swift1.5-Qwen3.8-Flash-Next,huihui-ai/Huihui-Qwen3.8-Flash-Next-abliterated
Last modified2026-10-04T05:25:08.000Z

Model README

---

license: other

license_name: qwen-community-1.0-and-swift-open-1.0

license_link: LICENSE

base_model:

  • Qwen/Qwen3.8-Flash-Next
  • ukisai/Swift1.5-Qwen3.8-Flash-Next
  • huihui-ai/Huihui-Qwen3.8-Flash-Next-abliterated

base_model_relation: merge

library_name: gguf

pipeline_tag: text-generation

tags: [gguf, merge, task-arithmetic, moe, qwen3.8, flash-next, strata, winery, abliterated, imatrix]

---

<p align="center"><a href="https://huggingface.co/WineryLabs"><img src="https://winery-api-zandy.zocomputer.io/card/m/cuvee.svg" alt="Winery Qwen3.8-Flash-Next Cuvée" width="100%"/></a></p>

🍷 Winery Qwen3.8-Flash-Next · Cuvée

One blend of Qwen3.8-Flash-Next (125B MoE, 512 experts, 10 active) in three bottles, sized so it runs on a gaming PC

with Winery Strata. That's a fork of

Strata with a wine-red glass UI, where Cuvée is the first model in the installer.

  • Short thinking from Swift 1.5 (UkisAI)
  • Refusals poured out by Huihui's abliteration
  • Base everywhere else: routed experts' gate/up, routers, hyper-connections, PLE n-gram table

🍷 Visit the Cuvée Space for the expert-cellar map and a bottle picker for your RAM.

The bottles

| Folder | For | Experts | Expert types | Size |

|---|---|---|---|---:|

| POCKET/ | 8 GB GPU · 24-32 GB RAM | 256 per layer | IQ2_XS gate/up · Q2_0 down, lean dense (Q4_K attn) | ~48 GB |

| SLIM/ | 32 GB RAM | 256 per layer | IQ3_XXS gate/up · IQ4_NL down | ~60 GB |

| RESERVE/ | 64 GB RAM | all 512 | IQ3_XXS gate/up · IQ4_NL down | ~86 GB |

| GRAND/ | 96-128 GB+ RAM | all 512 | Q4_K gate/up · Q5_1 down | ~113 GB |

All three keep the dense weights in Q6_K/Q8_0 and the 28.8 GB PLE table in IQ4_NL (Strata reads it from the SSD).

They were quantised with Unsloth's imatrix for Flash-Next.

Slim drops half the experts, the same move as ISTA-DASLab's Coder. The difference is that it keeps the 256

experts routed most often on general text (Unsloth's calibration routing counts, 72-91% of routed tokens per layer,

77% on average), not the ones code uses. Expect it to be a little weaker than Reserve, mostly on rare topics.

How it was made

Comparing the fine-tunes with the base tensor by tensor showed they touch different parts of the model:

| Tensors | Swift 1.5 | Huihui |

|---|:---:|:---:|

| attention q/k/v · DeltaNet qkv/z · shared-expert gate/up | ✓ | |

| attention/DeltaNet output · shared-expert down | ✓ | ✓ |

| routed experts' down | | ✓ |

| everything else | | |

So Cuvée = base + (Swift − base) + (Huihui − base) (task arithmetic, FP32 per tensor). Every tensor was streamed

straight from the three BF16 checkpoints with HTTP range requests, merged and written to a Q8_0 GGUF. The Q8_0 was then

quantised once to each bottle. The vision tower and MTP head are the original model's (Strata fetches them itself).

Run it

With Winery Strata (NVIDIA RTX 20-50 or a recent AMD card; 8 GB cards such as the RTX 4060 laptop GPU use POCKET, which setup picks for them by itself):

git clone https://huggingface.co/WineryLabs/Winery-Strata && cd Winery-Strata
START-HERE.bat --setup --family winery --model POCKET      # Windows, 8 GB GPU; SLIM / RESERVE / GRAND for more
START-HERE.bat --setup --family winery --model SLIM        # Windows; RESERVE / GRAND for more RAM
./setup.sh --setup --family winery --model RESERVE          # Linux

With llama.cpp (a build with qwen4exp support), e.g. llama-server -m RESERVE/Winery-Cuvee-RESERVE-00001-of-00002.gguf -c 32768 -ngl 99 --n-cpu-moe 48.

Not measured yet: speed on real hardware, and benchmark scores against the parent models. Sanity chats on CPU are in

TASTING.md.

Licence

A derivative of Qwen3.8-Flash-Next (Qwen Community License 1.0, LICENSE-QWEN) that contains UkisAI's

Swift Contribution (Swift Open License v1.0, LICENSE: free, including commercial use below US$1M yearly

revenue). Huihui's weights are under the Qwen licence. Attribution and our changes are in NOTICE.

Run WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models