AnkitAI/Parable-SmolLM3-3B-Claude-Fable-5-GGUF overview
Parable SmolLM3 3B Claude Fable 5 GGUF SmolLM3 3B banner.svg Part of the Parable series: small local LLMs fine tuned on genuine agent traces. This is HuggingFa…
Runs locally from ~1.78 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Parable-SmolLM3-3B-Claude-Fable-5-GGUF-F16.gguf | GGUF | F16 | 5.74 GB | Download |
| Parable-SmolLM3-3B-Claude-Fable-5-GGUF-Q4_K_M.gguf | GGUF | Q4_K_M | 1.78 GB | Download |
| Parable-SmolLM3-3B-Claude-Fable-5-GGUF-Q5_K_M.gguf | GGUF | Q5_K_M | 2.06 GB | Download |
| Parable-SmolLM3-3B-Claude-Fable-5-GGUF-Q6_K.gguf | GGUF | Q6_K | 2.36 GB | Download |
| Parable-SmolLM3-3B-Claude-Fable-5-GGUF-Q8_0.gguf | GGUF | Q8_0 | 3.05 GB | Download |
Model Details
| Model ID | AnkitAI/Parable-SmolLM3-3B-Claude-Fable-5-GGUF |
|---|---|
| Author | AnkitAI |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | HuggingFaceTB/SmolLM3-3B |
| Last modified | 2026-07-30T08:29:14.000Z |
Model README
---
license: apache-2.0
base_model: HuggingFaceTB/SmolLM3-3B
datasets:
- Glint-Research/Fable-5-traces
- Roman1111111/gpt5.5-terminal
pipeline_tag: text-generation
library_name: llama.cpp
tags:
- gguf
- qlora
- agentic
- coding
- reasoning
- smollm3
---
Parable-SmolLM3-3B-Claude-Fable-5-GGUF
Part of the Parable series: small local LLMs fine-tuned on genuine agent
traces. This is HuggingFaceTB/SmolLM3-3B tuned on real Claude Fable 5 agent
transcripts so its step-by-step reasoning voice carries into local use.
Full-precision weights: Parable-SmolLM3-3B-Claude-Fable-5
Files
| File | Quant | Size | |
|---|---|---|---|
| Parable-SmolLM3-3B-Claude-Fable-5-GGUF-Q4_K_M.gguf | Q4_K_M | 1.9 GB | recommended default |
| Parable-SmolLM3-3B-Claude-Fable-5-GGUF-Q5_K_M.gguf | Q5_K_M | 2.2 GB | |
| Parable-SmolLM3-3B-Claude-Fable-5-GGUF-Q6_K.gguf | Q6_K | 2.5 GB | |
| Parable-SmolLM3-3B-Claude-Fable-5-GGUF-Q8_0.gguf | Q8_0 | 3.3 GB | |
| Parable-SmolLM3-3B-Claude-Fable-5-GGUF-F16.gguf | F16 | 6.2 GB | for re-quantizing |
Usage
llama-cli -m Parable-SmolLM3-3B-Claude-Fable-5-GGUF-Q4_K_M.gguf \
-c 4096 -p "Write a python function that reverses a string." --temp 0.6
SmolLM3 is supported by current llama.cpp releases (brew, Ollama, LM Studio
builds included); no special build is required.
Output begins with a <think>...</think> reasoning block, then the answer.
If you are building on top of this model, parse and strip the think block
before showing text to end users. The chat template identifies the model as
"Parable, a coding assistant that reasons before it answers."
Model details
- Base: HuggingFaceTB/SmolLM3-3B (3B, Apache-2.0, 64k context)
- Method: MLX QLoRA on a 4-bit quantized base, 16 layers adapted,
6.7M trainable parameters (0.218%)
- Data: 4,076 training rows from real Claude Fable 5 agent-session
traces plus gpt5.5-terminal transcripts, prepared at a 4,096-token window
(268 over-length rows dropped; 226/226 rows held out for validation/test)
- Schedule: 1,200-iteration budget across a paused-and-resumed run;
best checkpoint selected on validation loss (iteration 200 of the final
segment, val 1.154)
Evaluation
| | Held-out trace test loss |
|---|---|
| SmolLM3-3B base | 1.889 |
| This model | 1.115 |
The tuned model fits the Fable-5 reasoning distribution 41% better by
held-out loss on a 226-row test split never seen in training. That is the
honest headline for what this fine-tune does; we do not claim general
benchmark gains.
This lane trains on trace data without a replay mix, so impact on general
coding benchmarks is unmeasured here. The series' technical report
(DOI: 10.5281/zenodo.21676407)
documents why that matters and what replay does about it.
Limitations
- Training ran on a 4-bit quantized base (16 GB M1 constraint); the F16
merge and higher quants cannot exceed 4-bit-base quality.
- Modest scale: one seed, loss-based evaluation, no external benchmark run
for this model yet.
- Not trained for: multi-file repo navigation, vision, non-English.
- Inherits SmolLM3-3B's knowledge cutoff. Treat generated commands as
drafts to review.
Quantization
Quantized with llama.cpp llama-quantize from the F16 merge.
Provenance & licensing
Fine-tuned from HuggingFaceTB/SmolLM3-3B (Apache-2.0). Training data:
(AGPL-3.0) and
(MIT). Because those traces originate from third-party assistants, the
providers' terms may apply to downstream training and distillation. If you
plan to build on this model commercially, confirm your use aligns with those
terms.
Citation
@misc{aglawe2026parable,
author = {Aglawe, Ankit},
title = {Agent-Trace Fine-Tuning of Small Language Models under Constrained Compute},
year = {2026},
doi = {10.5281/zenodo.21676407},
url = {https://doi.org/10.5281/zenodo.21676407}
}
Acknowledgements
The SmolLM3 team at Hugging Face for the base model; Glint-Research and
Roman1111111 for the trace datasets; empero-ai for the recipe this series
iterates on.
Run AnkitAI/Parable-SmolLM3-3B-Claude-Fable-5-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models