GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF overview

Parable Qwen3 4B Claude Fable 5 GGUF Parable banner.svg A 4B local coding model with agent instincts. Planning, tool habits and terminal reasoning distilled fr…

llama.cppggufqloraagenticcodingreasoningqwen3local-llmollamalm-studiotext-generationdataset:Glint-Research/Fable-5-tracesdataset:Roman1111111/gpt5.5-terminalbase_model:Qwen/Qwen3-4Bbase_model:quantized:Qwen/Qwen3-4Blicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~2.33 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
146,204
Likes
14
Pipeline
text-generation
Author

Repository Files & Downloads

5 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Parable-Qwen3-4B-Claude-Fable-5-GGUF-F16.ggufGGUFF167.50 GBDownload
Parable-Qwen3-4B-Claude-Fable-5-GGUF-Q4_K_M.ggufGGUFQ4_K_M2.33 GBDownload
Parable-Qwen3-4B-Claude-Fable-5-GGUF-Q5_K_M.ggufGGUFQ5_K_M2.69 GBDownload
Parable-Qwen3-4B-Claude-Fable-5-GGUF-Q6_K.ggufGGUFQ6_K3.08 GBDownload
Parable-Qwen3-4B-Claude-Fable-5-GGUF-Q8_0.ggufGGUFQ8_03.99 GBDownload

Model Details

Model IDAnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF
AuthorAnkitAI
Pipelinetext-generation
Licenseapache-2.0
Base modelQwen/Qwen3-4B
Last modified2026-08-11T13:00:55.000Z

Model README

---

license: apache-2.0

base_model: Qwen/Qwen3-4B

datasets:

- Glint-Research/Fable-5-traces

- Roman1111111/gpt5.5-terminal

pipeline_tag: text-generation

library_name: llama.cpp

tags:

- gguf

- qlora

- agentic

- coding

- reasoning

- qwen3

- local-llm

- ollama

- lm-studio

---

Parable-Qwen3-4B-Claude-Fable-5-GGUF

!Parable

A 4B local coding model with agent instincts. Planning, tool habits and

terminal reasoning distilled from real Claude Fable 5 agent sessions, not

synthetic Q&A. Runs on ~2.5 GB of RAM.

ollama run hf.co/AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF:Q4_K_M

v3.1 (2026-08-11)

Retrained on corpus v3.1: the v2 agent traces plus 1,807 execution-verified

solutions generated by the previous build and kept only where the code actually

ran against its tests. Two seeds souped, merged at the v2.1 scale.

The result matches or beats base Qwen3-4B on all four execution benchmarks,

where the previous build trailed it on three. If you pulled this model before

11 August 2026, re-pull.

Files

| File | Quant | Size | |

|---|---|---|---|

| Parable-Qwen3-4B-Claude-Fable-5-GGUF-Q4_K_M.gguf | Q4_K_M | 2.5 GB | recommended |

| Parable-Qwen3-4B-Claude-Fable-5-GGUF-Q5_K_M.gguf | Q5_K_M | 2.9 GB | |

| Parable-Qwen3-4B-Claude-Fable-5-GGUF-Q6_K.gguf | Q6_K | 3.3 GB | |

| Parable-Qwen3-4B-Claude-Fable-5-GGUF-Q8_0.gguf | Q8_0 | 4.3 GB | |

| Parable-Qwen3-4B-Claude-Fable-5-GGUF-F16.gguf | F16 | 8.1 GB | for re-quantizing |

What it is good at

  • It answers. Base Qwen3-4B spends its whole budget inside <think> on

34% of ordinary prompts and returns nothing. This model answers 34/34 on

the same suite, with 140x less reasoning text and no thinking-mode flag to

manage.

  • Agent-shaped reasoning. Trained on genuine multi-step agent sessions,

so plans, tool selection and terminal workflows come out structured

instead of improvised.

  • Small enough to keep open. Q4_K_M is 2.5 GB. Laptop, old GPU, modest

desktop — it runs, offline, with your code staying on your machine.

Evaluation

Measured on identical harnesses, greedy decoding, Q4_K_M builds, thinking

disabled on every row. Base and this model run through the same instrument in

the same session.

| | Base Qwen3-4B | This model (v3.1) |

|---|---|---|

| HumanEval | 73.2 | 74.4 |

| HumanEval+ | 68.3 | 68.3 |

| MBPP | 69.0 | 72.8 |

| MBPP+ | 59.8 | 63.8 |

| Held-out agent-trace loss | 2.155 | 1.446 |

Measured on the v2.1 build and carried forward (the training objective and

chat behaviour are unchanged):

| | Base Qwen3-4B | Parable |

|---|---|---|

| Prompts answered (34-prompt suite) | 27/34 | 34/34 |

| BFCL simple_python | 95.3 | 92.3 |

| BFCL multiple | 94.5 | 90.0 |

Choosing between this and the base

Take this model for local agent and coding work where you want

structured, reliable answers every time: it fits the agent-session

distribution far better and never silently returns empty.

Take the base model if your workload is maximum-accuracy function

calling in a tool-calling harness, where its few extra points matter more

than reasoning style.

Model details

  • Base: Qwen/Qwen3-4B (4B, Apache-2.0)
  • Method: QLoRA (nf4, r16, alpha 32) on all-linear targets, completion-only

loss masking, 30% general-instruction replay mix, seed-averaged weights,

merged at scale 0.6 (v2.1 recalibration)

  • Data: genuine Claude Fable 5 agent sessions + gpt5.5-terminal

transcripts, deduplicated and decontaminated against the reported benchmarks

Provenance & licensing

Fine-tuned from Qwen/Qwen3-4B (Apache-2.0). Training data:

Glint-Research/Fable-5-traces

(AGPL-3.0) and

Roman1111111/gpt5.5-terminal

(MIT). Because those traces originate from third-party assistants, the

providers' terms may apply to downstream training and distillation. If you

plan to build on this model commercially, confirm your use aligns with those

terms.

Citation

@misc{aglawe2026agenttrace,
  author    = {Aglawe, Ankit},
  title     = {Agent-Trace Fine-Tuning of Small Language Models under Constrained Compute},
  year      = {2026},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.21676407},
  url       = {https://doi.org/10.5281/zenodo.21676407}
}

Acknowledgements

The Qwen team for the base model; Glint-Research and Roman1111111 for the

trace datasets; empero-ai for the recipe this series iterates on.

Run AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models