AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF overview
<picture <source media=" prefers color scheme: dark " srcset="https://raw.githubusercontent.com/ankit aglawe/parable assets/main/parable header dark.png" <img ā¦
Runs locally from ~2.33 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Parable-Qwen3-4B-Claude-Fable-5-GGUF-F16.gguf | GGUF | F16 | 7.50 GB | Download |
| Parable-Qwen3-4B-Claude-Fable-5-GGUF-Q4_K_M.gguf | GGUF | Q4_K_M | 2.33 GB | Download |
| Parable-Qwen3-4B-Claude-Fable-5-GGUF-Q5_K_M.gguf | GGUF | Q5_K_M | 2.69 GB | Download |
| Parable-Qwen3-4B-Claude-Fable-5-GGUF-Q6_K.gguf | GGUF | Q6_K | 3.08 GB | Download |
| Parable-Qwen3-4B-Claude-Fable-5-GGUF-Q8_0.gguf | GGUF | Q8_0 | 3.99 GB | Download |
Model Details
| Model ID | AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF |
|---|---|
| Author | AnkitAI |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | Qwen/Qwen3-4B |
| Last modified | 2026-07-23T08:58:28.000Z |
Model README
---
base_model: Qwen/Qwen3-4B
base_model_relation: finetune
datasets:
- Glint-Research/Fable-5-traces
- Roman1111111/gpt5.5-terminal
license: apache-2.0
language:
- en
pipeline_tag: text-generation
library_name: gguf
tags:
- gguf
- qlora
- agentic
- agent
- coding
- tool-use
- function-calling
- terminal
- reasoning
- claude
- claude-fable-5
- distillation
- trace-training
- llama.cpp
- ollama
- lm-studio
- qwen3
---
<picture>
<source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/ankit-aglawe/parable-assets/main/parable_header_dark.png">
<img alt="Parable" src="https://raw.githubusercontent.com/ankit-aglawe/parable-assets/main/parable_header.png">
</picture>
šŖ¶ Parable-Qwen3-4B ā trained on genuine Claude Fable 5 agent traces
A small local model with agent instincts ā planning, tool use, and terminal habits distilled from real agent sessions, not synthetic Q&A. v2: every prompt gets an answer, eval-gated before publish.
> ~4 GB of RAM is all you need. Laptop, old GPU, modest desktop ā the Q4 build runs anywhere.
> One command and you have a private, offline reasoning model on your machine:
>
> ```bash
> ollama run hf.co/AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF:Q4_K_M
> ```
---
v2 is here (2026-07-22)
Every prompt now gets an answer. v1 inherited Qwen3's runaway thinking: on 21% of ordinary prompts it spent its entire token budget reasoning and returned nothing. v2 fixes that ā 34/34 prompts answered, answers roughly half as long, coding ability held.
Same repo, same links, same commands. Pull again and you have it. v1 remains in the version history below.
Full family. This 4B sits between its Parable siblings ā a 3B and two 8Bs. Browse the full collection.
---
Pick your size
| File | Size | Fits in | Notes |
|---|---|---|---|
| Q4_K_M | 2.5 GB | ~4 GB RAM/VRAM | ā Recommended ā best size/quality balance |
| Q5_K_M | 2.9 GB | ~4.5 GB | Higher quality |
| Q6_K | 3.3 GB | ~5 GB | Near-lossless |
| Q8_0 | 4.3 GB | ~6 GB | Maximum quality |
Full-precision safetensors (vLLM, transformers, further fine-tuning): Parable-Qwen3-4B-Claude-Fable-5
How to run it
Ollama (shorter command via the Parable namespace, or pull straight from this repo):
ollama run parable/qwen3-fable:4b
# or straight from this repo:
ollama run hf.co/AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF:Q4_K_M
llama.cpp:
llama-cli -m Parable-Qwen3-4B-Claude-Fable-5-GGUF-Q4_K_M.gguf --jinja \
-p "Write a bash one-liner to find the 10 largest files in a directory tree."
LM Studio: lms get parable/qwen3-fable, or search "parable" in-app (parable on LM Studio Hub).
Python (llama-cpp-python):
from llama_cpp import Llama
llm = Llama.from_pretrained(
repo_id="AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF",
filename="*Q4_K_M.gguf", n_ctx=8192,
)
out = llm.create_chat_completion(
messages=[{"role": "user", "content": "Write a Python function that retries an HTTP request with exponential backoff."}],
max_tokens=3000, temperature=0.3,
)
print(out["choices"][0]["message"]["content"])
Answers first
v2 responds directly ā no <think> preamble to strip, no reasoning budget to babysit. The planning instincts from the agent traces are still there; they just don't get narrated at you first.
Sampling: temperature 0.3ā0.7, max_tokens 800ā1500 is plenty for most prompts. (v1 needed ā„2500 and could still time out mid-thought. That's over.)
---
The numbers
Two builds, one harness, greedy decoding, HumanEval+ scored with EvalPlus on identical hardware:
| | Qwen3-4B base | Parable v1 | Parable v2 |
|---|---|---|---|
| Prompts answered (34-prompt suite) | ā | 27/34 | 34/34 |
| HumanEval | 76.8 | 72.0 | 73.2 |
| HumanEval+ | 71.3 | 68.3 | 67.1 |
| Held-out trace loss | 2.846 | ā | 1.876 (ā34%) |
The headline is the first row. On 7 of 34 ordinary prompts ā "write a retry function with exponential backoff", "write a bash script that watches a log file" ā v1 burned its whole budget thinking and returned an empty response. v2 returns 0 empty responses, using 140Ć less reasoning text. Coding scores are unchanged within noise, so this is a reliability fix, not a capability trade.
Against the base model, v2 trails by 3.6 HumanEval points ā the cost of specializing on agent traces, and about a third of the regression the same recipe produces at 3B. Trace loss drops 34% on held-out sessions: it learned the agent style without paying the usual forgetting tax.
<sub>Reproducibility: two training seeds land within 0.001 trace loss of each other (1.877 / 1.876); we ship the second and publish both in parable-v2-artifacts. Benchmarks run with thinking disabled on both base and fine-tune ā the base is a hybrid-thinking model and scores near zero in default mode because it never emits a final answer. Comparing against that would have flattered us.</sub>
Training data
- Glint-Research/Fable-5-traces ā 4.4k real Claude Fable 5 coding-agent session traces with
<think>reasoning and tool calls (AGPL-3.0) - Roman1111111/gpt5.5-terminal ā terminal-agent task solutions (MIT)
v2 trains on corpus v2 ā 10.6k curated traces (13Ć v1) with completion-only loss masking, a 30% general-instruction replay mix, and span-filtered truncation. Every example passed a quality gate (schema validation, secrets scrub, length filtering). QLoRA fine-tune (TRL), quantized with llama.cpp.
Good to know
- Answers directly rather than reasoning aloud. If you want visible chain-of-thought, prompt for it explicitly ("think step by step, show your work").
- Tuned hard toward agentic coding behavior; that focus trades some general-knowledge breadth, as with any specialized fine-tune in this class.
- Verify critical output. Small models over-commit to plausible specifics; treat generated commands and code as drafts to review.
- Inherits Qwen3-4B's base limitations and knowledge cutoff.
Base & license
Weights: Apache-2.0 (inherited from Qwen/Qwen3-4B). Training data: Fable-5-traces AGPL-3.0, gpt5.5-terminal MIT ā since those traces originate from third-party assistants, the providers' terms may apply to downstream training and distillation; confirm your use aligns with them before building on this model commercially.
Get Parable
| Platform | |
|---|---|
| Ollama | ollama run parable/qwen3-fable:4b Ā· parable namespace |
| Ollama (family flagship, best per size) | ollama run parable/fable |
| Hugging Face | full collection ā GGUF quants, full weights, eval reports |
| LM Studio | lms get parable/qwen3-fable Ā· parable on LM Studio Hub |
Acknowledgements
Glint-Research & Roman1111111 for the open trace data Ā· Qwen team for the base Ā· empero-ai whose Qwable recipe this release follows Ā· mlx-lm & llama.cpp
---
Four gigabytes of RAM. Real Fable 5 reasoning. Yours, offline, right now.
ollama run hf.co/AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF:Q4_K_M
Version history
- v2 (2026-07-22) ā rebuilt recipe: 10.6k-trace corpus v2, completion-only loss masking, 30% replay mix, span-filtered truncation. Fixes v1's empty-response failure mode (7/34 ā 0/34). Coding parity with v1, trace loss ā34% vs base.
- v1 (2026-07-11) ā initial release: QLoRA on 4.4k Fable-5 traces via mlx-lm.
More on the Parable models: ankitaglawe.com/parable
Run AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF with guIDE
Download guIDE ā the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face Ā· Compare models