Brian6145/Qwen3.6-27B-Claude-Opus-DeepSeek-Distilled-Imatrix-MTP-GGUF overview
๐ง Opus DeepSeek Distilled Q4M A distilled Qwen3.6 27B GGUF optimized for local agentic reasoning, tool use, and long chain task execution. BenchLocal https://โฆ
Runs locally from ~10.12 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Qwen3.6-27B-Claude-Opus-DeepSeek-Distilled-Imatrix-MTP-GGUF-IQ4_XS.gguf | GGUF | IQ4_XS | 14.26 GB | Download |
| Qwen3.6-27B-Claude-Opus-DeepSeek-Distilled-Imatrix-MTP-GGUF-NVFP4.gguf | GGUF | GGUF | 18.34 GB | Download |
| Qwen3.6-27B-Claude-Opus-DeepSeek-Distilled-Imatrix-MTP-GGUF-Q2_K.gguf | GGUF | Q2_K | 10.12 GB | Download |
| Qwen3.6-27B-Claude-Opus-DeepSeek-Distilled-Imatrix-MTP-GGUF-Q3_K_M.gguf | GGUF | Q3_K_M | 12.57 GB | Download |
| Qwen3.6-27B-Claude-Opus-DeepSeek-Distilled-Imatrix-MTP-GGUF-Q4_K_M.gguf | GGUF | Q4_K_M | 15.66 GB | Download |
| Qwen3.6-27B-Claude-Opus-DeepSeek-Distilled-Imatrix-MTP-GGUF-Q5_K_M.gguf | GGUF | Q5_K_M | 18.19 GB | Download |
| Qwen3.6-27B-Claude-Opus-DeepSeek-Distilled-Imatrix-MTP-GGUF-Q6_K.gguf | GGUF | Q6_K | 20.89 GB | Download |
| Qwen3.6-27B-Claude-Opus-DeepSeek-Distilled-Imatrix-MTP-GGUF-Q8_0.gguf | GGUF | Q8_0 | 27.05 GB | Download |
Model Details
| Model ID | Brian6145/Qwen3.6-27B-Claude-Opus-DeepSeek-Distilled-Imatrix-MTP-GGUF |
|---|---|
| Author | Brian6145 |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | Qwen/Qwen3.6-27B |
| Last modified | 2026-07-14T04:38:58.000Z |
Model README
---
language:
- en
- zh
base_model:
- Qwen/Qwen3.6-27B
tags:
- gguf
- llama.cpp
- qwen
- qwen3
- qwen3.6
- mtp
- speculative-decoding
- quantized
- long-context
- chinese
- distillation
- agent
- react
pipeline_tag: text-generation
license: apache-2.0
library_name: gguf
inference:
parameters:
temperature: 0.6
top_p: 0.95
model-index:
- name: Opus-DeepSeek-Distilled-Q4M
results:
- task:
type: text-generation
dataset:
type: benchlocal
name: BenchLocal 6-pack
metrics:
- type: benchlocal-score
value: 86.5
name: 6-pack Total
- task:
type: question-answering
dataset:
type: gpqa
name: GPQA-Diamond-198
metrics:
- type: accuracy
value: 83.84
name: Accuracy
- task:
type: question-answering
dataset:
type: mmlu
name: MMLU-500 (5-shot)
metrics:
- type: accuracy
value: 91.80
name: Accuracy
---
๐ง Opus-DeepSeek-Distilled-Q4M
> A distilled Qwen3.6-27B GGUF optimized for local agentic reasoning, tool use, and long-chain task execution.





---
โ ๏ธ Sampling Parameters
> Please ensure temperature = 0.6 and top_p = 0.95 when using this model.
>
> This model was trained and validated at these specific parameters. Both too high and too low temperatures cause problems:
>
> ๐ฅ Too high (> 0.6) โ Output becomes divergent as token probabilities flatten. This causes malformed tool calls, function name hallucinations, unstable parameter generation, and uncontrollable agent behavior.
>
> ๐ง Too low (< 0.3) โ The model almost always picks the highest-probability token. In tool-calling scenarios this manifests as: repeating the same failed tool call instead of trying alternatives, shortened thinking traces leading to insufficient reasoning depth, and reduced robustness to edge cases.
>
> This is an important recommendation based on extensive real-world testing. Please verify these parameters in your inference framework.
---
๐ข Highlights
| Area | Score | vs Qwen3.6-27B q4_k_m |
|------|-------|----------------------|
| BenchLocal 6-pack ๐ | 86.5 | +8.3 |
| GPQA-Diamond-198 ๐ฌ | 83.84% | +10.14% |
| BugFind-15 ๐ | 80 | +20 |
| ToolCall-15 ๐ง | 97 | +4 |
| InstructFollow-15 ๐ | 94 | +17 |
| StructOutput-15 ๐ | 88 | +11 |
| MMLU-500 (5-shot) ๐ | 91.80% | ~tied (+0.2%) |
| DataExtract-15 ๐ | 81 | -2 |
> ๐ Output speed: ~60 tok/s on A100 40GB ยท ~100 tok/s on RTX PRO 6000 (q4_k_m + mtp=3)
---
๐ฅ Why This Model?
The original Qwen3.6-27B has solid foundational capabilities, but its agent behavior falls short โ prone to infinite loops when thinking, lacks structured agent design, and has room to improve in math reasoning.
This variant tackles all three through targeted distillation:
- โ Infinite loops โ Eliminated. ReAct-style reasoning-action orchestration fixes the root cause.
- โ Agent behavior โ Structured. Distilled Claude Opus's systematic thinking and organization.
- โ Math reasoning โ Strengthened. Absorbed capabilities from strong math/logic models.
The result is a local agent that doesn't just score high on benchmarks โ it works reliably in real engineering tasks like BugFind, where it now competes with GLM5.2.
> โ ๏ธ Note: Side-by-side GLM5.2 comparisons were evaluated using Opus 4.8 as a judge. Opus-as-judge has inherent biases โ results are indicative, not definitive.
---
๐ฏ Design Philosophy
Core insight: Qwen3.6-27B doesn't lack capability โ it lacks good agent behavior. That makes it worth iterating on.
We followed the ReAct paradigm (Yao et al., ICLR 2023) โ unifying reasoning and action into an alternating, constrained, executable loop โ rather than just making the model "think longer" or "call tools better."
| Teacher Model | Capability Distilled |
|---------------|---------------------|
| Claude Opus ๐ฏ | Systematic thinking, structured organization, concise reasoning |
| DeepSeek ๐งญ | Stable agent behavior, tool orchestration, task closure |
| Math/Logic models โ | Mathematical reasoning, logical deduction |
---
๐ Performance
BenchLocal 6-pack
| Pack | q4_k_m (ours) | Qwen/Qwen3.6-27B q4_k_m | Delta |
|------|:-:|:-:|:-:|
| BugFind-15 ๐ | 80 | 60 | +20 |
| ToolCall-15 ๐ง | 97 | 93 | +4 |
| DataExtract-15 ๐ | 81 | 83 | -2 |
| InstructFollow-15 ๐ | 94 | 77 | +17 |
| ReasonMath-15 โ | 79 | 79 | 0 |
| StructOutput-15 ๐ | 88 | 77 | +11 |
| Total ๐ | 86.5 | 78.2 | +8.3 |
Extended Evals
| Benchmark | Ours | Baseline | Notes |
|-----------|:----:|:--------:|-------|
| GPQA-Diamond-198 ๐ฌ | 83.84% | 73.7% | +10.14%, all 198 graded locally |
| MMLU-500 (5-shot) ๐ | 91.80% | 91.6% | Approximately tied |
---
๐ ๏ธ Usage
Recommended Stack
๐งฉ OpenCode + LM Studio
๐ Temperature: 0.6 ยท Top-p: 0.95
โก q4_k_m + mtp=3
Quick Start (llama.cpp)
# Download the GGUF
huggingface-cli download your-org/Opus-DeepSeek-Distilled-Q4M \
opus-deepseek-distilled-q4m-q4_k_m.gguf --local-dir ./models
# Run with llama.cpp
./llama-cli -m ./models/opus-deepseek-distilled-q4m-q4_k_m.gguf \
--temp 0.6 --top-p 0.95 \
-p "Your prompt here"
Via LM Studio
- Load the GGUF file in LM Studio
- Set backend to llama.cpp
- Enable MTP (set depth=3) under inference options
- Set
temperature = 0.6,top_p = 0.95 - Start the local API server
- Connect via OpenCode or any OpenAI-compatible client
> ๐ก Pro tip: For coding tasks, the temp 0.6 / top_p 0.95 combo delivers the best balance of creativity and correctness.
---
โ ๏ธ Known Limitations
Evaluation Methodology
- Opus-as-judge biases: GLM5.2 comparisons are judge-evaluated, not absolute rankings
- MMLU: Single-run 5-shot result; variations in shot selection may cause fluctuation
Capability Boundaries
- DataExtract: Scores 81 vs 83 baseline โ extraction tasks may have slight regression from distillation
- Closed-source teachers: Risks include inherited biases and TOS compliance โ assess for your use case
- 27B scale ceiling: May still hit capacity limits on extremely complex long-chain reasoning
Deployment Notes
- MTP=3: Boosts throughput but adds VRAM overhead โ disable or reduce on <24GB hardware
- Quantization: Only
q4_k_mtested;q5_k_mmay improve accuracy at higher VRAM cost - Imatrix: Uses importance-matrix quantization, not standard k-quant โ better parameter preservation at low bit widths
---
๐ Citation
@misc{opus-deepseek-distilled-q4m,
title = {Opus-DeepSeek-Distilled-Q4M: A Distilled Agentic GGUF for Local Deployment},
author = {Yin, Brian and BenchLocal Contributors},
year = {2026},
url = {https://github.com/brianyin/BenchLocal}
}
---
๐ Acknowledgements
- Qwen team โ excellent foundational model
- Unsloth โ efficient training infrastructure
- Merkyor โ identified ReAct as the key to solving agent infinite loops
- Community โ built on existing open-source exploration and practical experience
> It is because of this work that came before that we can continue pushing forward, arriving at today's more stable, more practical, and more complete agent โ and helping us get closer to the era of local agent AI.
Run Brian6145/Qwen3.6-27B-Claude-Opus-DeepSeek-Distilled-Imatrix-MTP-GGUF with guIDE
Download guIDE โ the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face ยท Compare models