Model Intelligence Sheet
ridd1er1/Qwen3.6-35B-A3B-DFlash-GGUF overview
Qwen 3.6 35B A3B DFlash GGUF llama.cpp quantizations of z lab DFlash draft model https://huggingface.co/z lab/Qwen3.6 35B A3B DFlash for Qwen 3.6 35B A3B https…
Runs locally from ~253.8 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
6 GGUF files detected
Direct downloads for local inference
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| qwen36-35b-a3b-dflash-IQ4_XS.gguf | GGUF | IQ4_XS | 253.8 MB | Download |
| qwen36-35b-a3b-dflash-Q4_K_M.gguf | GGUF | Q4_K_M | 278.2 MB | Download |
| qwen36-35b-a3b-dflash-Q5_K_M.gguf | GGUF | Q5_K_M | 328.2 MB | Download |
| qwen36-35b-a3b-dflash-Q6_K.gguf | GGUF | Q6_K | 381.4 MB | Download |
| qwen36-35b-a3b-dflash-Q8_0.gguf | GGUF | Q8_0 | 490.8 MB | Download |
| qwen36-35b-a3b-dflash-bf16.gguf | GGUF | BF16 | 914.6 MB | Download |
Model Details
Model README
---
base_model: z-lab/Qwen3.6-35B-A3B-DFlash
---
Qwen 3.6 35B A3B DFlash GGUF
llama.cpp quantizations of z-lab DFlash draft model for Qwen 3.6 35B A3B.
Use with BeeLlama.cpp — a llama.cpp fork with advanced DFlash support that enables using these draft models to their full potential.
Run ridd1er1/Qwen3.6-35B-A3B-DFlash-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models