GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Alittlehammmer/Qwen3.6-27B-DFlash-GGUF-llama.cpp overview

Qwen3.6 27B DFlash GGUF quantizations of z lab/Qwen3.6 27B DFlash https://huggingface.co/z lab/Qwen3.6 27B DFlash . Converted to BF16 using convert hf to gguf.…

ggufquantizedllama.cpptext-generationagentic-codingdflashbase_model:z-lab/Qwen3.6-27B-DFlashbase_model:quantized:z-lab/Qwen3.6-27B-DFlashlicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~985.2 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
1
Pipeline
text-generation

Repository Files & Downloads

5 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3.6-27B-DFlash-BF16.ggufGGUFBF163.23 GBDownload
Qwen3.6-27B-DFlash-Q4_K_M.ggufGGUFQ4_K_M985.2 MBDownload
Qwen3.6-27B-DFlash-Q5_K.ggufGGUFQ5_K1.14 GBDownload
Qwen3.6-27B-DFlash-Q6_K.ggufGGUFQ6_K1.33 GBDownload
Qwen3.6-27B-DFlash-Q8_0.ggufGGUFQ8_01.72 GBDownload

Model Details

Model IDAlittlehammmer/Qwen3.6-27B-DFlash-GGUF-llama.cpp
AuthorAlittlehammmer
Pipelinetext-generation
Licenseapache-2.0
Base modelz-lab/Qwen3.6-27B-DFlash
Last modified2026-06-29T18:40:52.000Z

Model README

---

base_model:

- z-lab/Qwen3.6-27B-DFlash

base_model_relation: quantized

quantized_by: Alittlehammmer

license: apache-2.0

license_link: https://huggingface.co/z-lab/Qwen3.6-27B-DFlash/blob/main/LICENSE

pipeline_tag: text-generation

library_name: gguf

tags:

- gguf

- quantized

- llama.cpp

- text-generation

- agentic-coding

- dflash

---

Qwen3.6-27B-DFlash

GGUF quantizations of z-lab/Qwen3.6-27B-DFlash.

Converted to BF16 using convert_hf_to_gguf.py, then quantized using llama-quantize from llama.cpp.

Available quants

| Quant | Bits | Size | Notes |

| ------ | ----- | ------- | -------------------------------- |

| Q4_K_M | 4 | ~1.03 GB | Average quality |

| Q5_K | 5 | ~1.22 GB | High quality |

| Q6_K | 6 | ~1.43 GB | Very high quality |

| Q8_0 | 8 | ~1.84 GB | Highest quality, near lossless, Recommended |

| BF16 | 16 | ~3.47 GB | Full precision, reference file |

Usage

Use in conjunction with existing Qwen3.6 Quants, example config if using llama-server:

[Qwen3.6-27B-Q8_0-DFlash]
sm = layer
model = /mnt/gguf/Qwen3.6-27B/Qwen3.6-27B-Q8_0.gguf
model-draft = /mnt/gguf/Qwen3.6-27B/Qwen3.6-27B-DFlash-Q8_0.gguf
spec-type = draft-dflash
spec-draft-n-max = 6 

(Note: For some reason I cannot get sm = tensor to work, it crashes on launch, pretty sure this is an issue in llama.cpp)

Original model

See the original model card

for details on capabilities, benchmarks, and license.

Run Alittlehammmer/Qwen3.6-27B-DFlash-GGUF-llama.cpp with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models