GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE β†’
Model Intelligence Sheet

mfielding92/SmartCode-Fable-5-CoT-Reasoning-QVO-Qwen-3.6-27B-Distilled-GGUF overview

πŸš€ SmartCode ThinkingCap Fable 5 Distill Qwen3.6 27B 🧩 Finetuned from: bottlecapai/ThinkingCap Qwen3.6 27B https://huggingface.co/bottlecapai/ThinkingCap Qwen…

transformersggufquantizedunsloth-dynamicimatrixtext-generation-inferenceunslothqwen3_5fable 5fablecotreasoningsmartcodecodeqwen3_6token-efficientefficient-thinkingendataset:mfielding92/smartcode-fable-5-distill-cot-reasoning-1000xbase_model:bottlecapai/ThinkingCap-Qwen3.6-27Bbase_model:quantized:bottlecapai/ThinkingCap-Qwen3.6-27Blicense:apache-2.0endpoints_compatibleregion:us

Runs locally from ~11.18 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
332
Likes
0
Pipeline
β€”

Repository Files & Downloads

20 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
SmartCode-Fable-5-CoT-Reasoning-QVO-Qwen-3.6-27B-Distilled-BF16.ggufGGUFBF1650.90 GBDownload
SmartCode-Fable-5-CoT-Reasoning-QVO-Qwen-3.6-27B-Distilled-IQ3_S.ggufGGUFIQ3_S11.74 GBDownload
SmartCode-Fable-5-CoT-Reasoning-QVO-Qwen-3.6-27B-Distilled-IQ4_NL.ggufGGUFIQ4_NL14.94 GBDownload
SmartCode-Fable-5-CoT-Reasoning-QVO-Qwen-3.6-27B-Distilled-IQ4_NL_XL.ggufGGUFIQ4_NL_XL14.94 GBDownload
SmartCode-Fable-5-CoT-Reasoning-QVO-Qwen-3.6-27B-Distilled-IQ4_XS.ggufGGUFIQ4_XS14.26 GBDownload
SmartCode-Fable-5-CoT-Reasoning-QVO-Qwen-3.6-27B-Distilled-MXFP4_MOE.ggufGGUFGGUF27.05 GBDownload
SmartCode-Fable-5-CoT-Reasoning-QVO-Qwen-3.6-27B-Distilled-Q3_K_M.ggufGGUFQ3_K_M12.57 GBDownload
SmartCode-Fable-5-CoT-Reasoning-QVO-Qwen-3.6-27B-Distilled-Q3_K_S.ggufGGUFQ3_K_S11.41 GBDownload
SmartCode-Fable-5-CoT-Reasoning-QVO-Qwen-3.6-27B-Distilled-Q4_K_M.ggufGGUFQ4_K_M15.66 GBDownload
SmartCode-Fable-5-CoT-Reasoning-QVO-Qwen-3.6-27B-Distilled-Q4_K_S.ggufGGUFQ4_K_S14.74 GBDownload
SmartCode-Fable-5-CoT-Reasoning-QVO-Qwen-3.6-27B-Distilled-Q5_K_M.ggufGGUFQ5_K_M18.19 GBDownload
SmartCode-Fable-5-CoT-Reasoning-QVO-Qwen-3.6-27B-Distilled-Q5_K_S.ggufGGUFQ5_K_S17.67 GBDownload
SmartCode-Fable-5-CoT-Reasoning-QVO-Qwen-3.6-27B-Distilled-Q8_0.ggufGGUFQ8_027.05 GBDownload
SmartCode-Fable-5-CoT-Reasoning-QVO-Qwen-3.6-27B-Distilled-UD-Q2_K_XL.ggufGGUFQ2_K_XL11.18 GBDownload
SmartCode-Fable-5-CoT-Reasoning-QVO-Qwen-3.6-27B-Distilled-UD-Q3_K_XL.ggufGGUFQ3_K_XL13.68 GBDownload
SmartCode-Fable-5-CoT-Reasoning-QVO-Qwen-3.6-27B-Distilled-UD-Q4_K_XL.ggufGGUFQ4_K_XL16.65 GBDownload
SmartCode-Fable-5-CoT-Reasoning-QVO-Qwen-3.6-27B-Distilled-UD-Q5_K_XL.ggufGGUFQ5_K_XL18.95 GBDownload
SmartCode-Fable-5-CoT-Reasoning-QVO-Qwen-3.6-27B-Distilled-UD-Q6_K.ggufGGUFQ6_K20.89 GBDownload
SmartCode-Fable-5-CoT-Reasoning-QVO-Qwen-3.6-27B-Distilled-UD-Q6_K_XL.ggufGGUFQ6_K_XL24.20 GBDownload
SmartCode-Fable-5-CoT-Reasoning-QVO-Qwen-3.6-27B-Distilled-UD-Q8_K_XL.ggufGGUFQ8_K_XL33.32 GBDownload

Model Details

Model IDmfielding92/SmartCode-Fable-5-CoT-Reasoning-QVO-Qwen-3.6-27B-Distilled-GGUF
Authormfielding92
Pipelineβ€”
Licenseapache-2.0
Base modelbottlecapai/ThinkingCap-Qwen3.6-27B
Last modified2026-07-26T18:25:56.000Z

Model README

---

base_model: bottlecapai/ThinkingCap-Qwen3.6-27B

datasets:

  • mfielding92/smartcode-fable-5-distill-cot-reasoning-1000x

library_name: transformers

tags:

  • gguf
  • quantized
  • unsloth-dynamic
  • imatrix
  • text-generation-inference
  • transformers
  • unsloth
  • qwen3_5
  • fable 5
  • fable
  • cot
  • reasoning
  • smartcode
  • code
  • qwen3_6
  • token-efficient
  • efficient-thinking

license: apache-2.0

language:

  • en

pretty_name: SmartCode ThinkingCap Fable 5 Distill 27B

---

πŸš€ SmartCode ThinkingCap Fable 5 Distill (Qwen3.6-27B)

🎩 Layer 1: The ThinkingCap Foundation

The base model isn't just another Qwen3.6-27B repack. ThinkingCap is the finetune that took a serious swing at reasoning bloat:

  • ⚑ ~50% fewer thinking tokens on average than stock Qwen3.6-27B.
  • πŸ”₯ Over 90% reduction in best-case scenarios. It stops re-deriving the obvious.
  • 🎯 Less than 1% average accuracy difference across benchmarks compared to the original base.

Translation: you keep the intelligence, you drop the rambling. No more 3,000 tokens of "wait, but actually, let me reconsider" before printing a two-line function.

🍜 Layer 2: The SmartCode Fable 5 Injection

On TOP of that efficiency, this finetune pours in 1,000 curated reasoning sequences distilled from Fable 5, and with them, a genuine coding personality:

  • πŸ—£οΈ It talks to itself as it codes. The model narrates its work in real time, catching mistakes before they land in the output.
  • πŸ—ΊοΈ It plans out loud before writing a line. What it's building, how, and why, the same discipline agentic coding tools enforce, baked directly into the weights.
  • πŸ“ˆ Adaptive Reasoning depth. Chain-of-Thought scales with task complexity. Easy problems get short chains. Hard problems get the full Fable 5 treatment.
  • 🧬 No complexity overload. Most distills fail by cramming frontier-sized reasoning into a model that can't carry it. This dataset was built to fit.

Translation: it doesn't just answer, it works the problem. Plan first, code second, accuracy up.

The two layers fit together well. ThinkingCap was already trained to think only as much as necessary, and Adaptive Reasoning data teaches exactly that behavior at the frontier level. The techniques reinforce each other instead of fighting. You get frontier-flavored coding reasoning that terminates when it's actually done thinking.

🎯 READ THIS: Recommended Sampler Settings

This model was tuned around these exact settings. Do not skip this section. Bad sampler settings are the #1 cause of "the distill is broken" reports.

| Setting | Value |

| --- | --- |

| Temperature | 0.9 |

| Top P | 0.95 |

| Top K | 60 |

| Min P | 0 |

| Repeat Penalty | NONE (1.0 / disabled) |

| Presence Penalty | 1.0 |

Recommended quants

| Quant | Approx size | Notes |

|---|---|---|

| UD-Q3_K_XL | smaller, still very usable | best for 16GB VRAM cards |

| UD-Q4_K_XL | best quality/size ratio for most users | recommended default |

| UD-Q5_K_XL | higher quality | best for 24GB VRAM cards |

| UD-Q2_K_XL | extreme compression | available, but not recommended |

Run with llama.cpp

./llama.cpp/llama-cli \
  --model SmartCode-Fable-5-CoT-Reasoning-QVO-Qwen-3.6-27B-Distilled-UD-Q4_K_XL.gguf \
  --temp 0.9 --top-p 0.95 --top-k 60 --min-p 0.0 \
  --ctx-size 262144 --n-gpu-layers 99

> QKVO finetune also available: this version adapts K weights as well, so addressing/retrieval behavior (which tokens get attended to and how strongly) shifts with the finetuning data, on top of Q, V, and O. Stricter coherence to the dataset, but risk of loosing base-model functionality.

>

> mfielding92/SmartCode-Fable-5-CoT-Reasoning-QKVO-Qwen-3.6-27B-Distilled-GGUF

<details>

<summary>βš™οΈ <strong>Technical Specs</strong></summary>

  • Architecture: qwen3_5 (27B parameters)
  • Base: bottlecapai/ThinkingCap-Qwen3.6-27B
  • Max sequence length: 8,192 tokens per training entry
  • Reasoning cap in training data: under 2,048 tokens, for density and balance
  • Training samples: 1,000 Fable 5 distilled CoT sequences
  • Trained 2x faster with Unsloth and Huggingface's TRL library

</details>

<details>

<summary>πŸ’‘ <strong>Additional Usage Notes</strong></summary>

Standard Qwen3.6 chat template applies. Keep the reasoning block enabled, since this model's entire value proposition lives in how efficiently it uses that block.

Because both the base finetune and this distill were tuned toward short, decisive reasoning, aggressive thinking-budget forcing is generally unnecessary and may hurt output quality on the hard tail of problems. Let it decide.

</details>

<details>

<summary>⚠️ <strong>Usage Policy</strong></summary>

The model weights are released under apache-2.0.

However, the SmartCode Fable 5 Distill dataset it was trained on is designated personal hobbyist and educational use only, with commercial use prohibited to remain in alignment with provider policies. If you care about staying clean on distillation-provenance grounds, treat this model the same way: personal and educational use.

</details>

---

<img src="https://raw.githubusercontent.com/unslothai/unsloth/main/images/unsloth%20made%20with%20love.png" width="200"/>

Run mfielding92/SmartCode-Fable-5-CoT-Reasoning-QVO-Qwen-3.6-27B-Distilled-GGUF with guIDE

Download guIDE β€” the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE β†’ Β· Browse 524k+ models Β· Compare models

Source: Hugging Face Β· Compare models