Model Intelligence Sheet
symrex/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF-dequantized-oQ8e-mtp overview
Qwen3.6 27B Fable Fusion 711 Uncensored Heretic NM DAU NEO MAX MTP GGUF dequantized oQ8e mtp This model was quantized using oQ https://github.com/jundot/omlx o…
Repository Files & Downloads
0 GGUF files detected
Direct downloads for local inference
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Browse files on Hugging Face | ||||
Model Details
Model README
---
library_name: mlx
tags:
- mlx
- oq
- quantized
base_model:
- >-
DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF
---
Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF-dequantized-oQ8e-mtp
This model was quantized using oQ (oMLX v0.5.7) mixed-precision quantization.
Quantization details
- Model type: qwen3_5
- Bits: 8
- Group size: 64
- Format: MLX safetensors
Performance Benchmark
- Run on: Apple Mac Studio M4 Max 128GB
- https://omlx.ai/benchmarks/wse53h85
oMLX - LLM inference, optimized for your Mac
https://github.com/jundot/omlx
Benchmark Model: Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF-dequantized-oQ8e-mtp
Engine: Auto
Context: Code (Python)
================================================================================
Single Request Results
--------------------------------------------------------------------------------
Test TTFT(ms) TPOT(ms) pp TPS tg TPS E2E(s) Throughput Peak Mem
pp1024/tg128 4127.9 61.57 248.1 tok/s 16.4 tok/s 11.960 96.3 tok/s 28.34 GB
pp4096/tg128 16366.9 62.22 250.3 tok/s 16.2 tok/s 24.279 174.0 tok/s 29.93 GB
pp8192/tg128 33175.4 63.21 246.9 tok/s 15.9 tok/s 41.217 201.9 tok/s 30.84 GB
pp16384/tg128 68451.5 64.77 239.4 tok/s 15.6 tok/s 76.691 215.3 tok/s 32.67 GB
pp32768/tg128 145258.7 67.43 225.6 tok/s 14.9 tok/s 153.836 213.8 tok/s 36.32 GB
pp65536/tg128 326820.5 73.41 200.5 tok/s 13.7 tok/s 336.158 195.3 tok/s 43.64 GB
pp131072/tg128 812699.7 84.11 161.3 tok/s 12.0 tok/s 823.399 159.3 tok/s 58.61 GB
Continuous Batching
pp1024 / tg128
--------------------------------------------------------------------------------
Batch tg TPS Speedup pp TPS pp TPS/req TTFT(ms) E2E(s)
1x 16.4 tok/s 1.00x 248.1 tok/s 248.1 tok/s 4127.9 11.960
2x 31.7 tok/s 1.93x 191.1 tok/s 95.5 tok/s 10717.1 18.802
4x 60.0 tok/s 3.66x 242.7 tok/s 60.7 tok/s 16691.4 25.414
8x 75.9 tok/s 4.63x 241.7 tok/s 30.2 tok/s 33204.9 47.391
Intelligence Benchmark Comparison
Intelligence Benchmark Comparison
--- Detail ---
Model: Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF-dequantized-oQ8e-mtp
Benchmark Accuracy Correct Total Time(s) Think
--------------------------------------------------------------
MMLU 91.5% 915 1000 29025.1 Yes
TRUTHFULQA 90.0% 735 817 27397 Yes
HUMANEVAL 92.1% 151 164 10308.3 YesRun symrex/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF-dequantized-oQ8e-mtp with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models