GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

AnkitAI/Parable-SmolLM3-3B-Claude-Fable-5-GGUF overview

Parable SmolLM3 3B Claude Fable 5 GGUF SmolLM3 3B banner.svg Part of the Parable series: small local LLMs fine tuned on genuine agent traces. This is HuggingFa…

llama.cppggufqloraagenticcodingreasoningsmollm3text-generationdataset:Glint-Research/Fable-5-tracesdataset:Roman1111111/gpt5.5-terminalbase_model:HuggingFaceTB/SmolLM3-3Bbase_model:quantized:HuggingFaceTB/SmolLM3-3Blicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~1.78 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

5 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Parable-SmolLM3-3B-Claude-Fable-5-GGUF-F16.ggufGGUFF165.74 GBDownload
Parable-SmolLM3-3B-Claude-Fable-5-GGUF-Q4_K_M.ggufGGUFQ4_K_M1.78 GBDownload
Parable-SmolLM3-3B-Claude-Fable-5-GGUF-Q5_K_M.ggufGGUFQ5_K_M2.06 GBDownload
Parable-SmolLM3-3B-Claude-Fable-5-GGUF-Q6_K.ggufGGUFQ6_K2.36 GBDownload
Parable-SmolLM3-3B-Claude-Fable-5-GGUF-Q8_0.ggufGGUFQ8_03.05 GBDownload

Model Details

Model IDAnkitAI/Parable-SmolLM3-3B-Claude-Fable-5-GGUF
AuthorAnkitAI
Pipelinetext-generation
Licenseapache-2.0
Base modelHuggingFaceTB/SmolLM3-3B
Last modified2026-07-30T08:29:14.000Z

Model README

---

license: apache-2.0

base_model: HuggingFaceTB/SmolLM3-3B

datasets:

- Glint-Research/Fable-5-traces

- Roman1111111/gpt5.5-terminal

pipeline_tag: text-generation

library_name: llama.cpp

tags:

- gguf

- qlora

- agentic

- coding

- reasoning

- smollm3

---

Parable-SmolLM3-3B-Claude-Fable-5-GGUF

!SmolLM3-3B

Part of the Parable series: small local LLMs fine-tuned on genuine agent

traces. This is HuggingFaceTB/SmolLM3-3B tuned on real Claude Fable 5 agent

transcripts so its step-by-step reasoning voice carries into local use.

Full-precision weights: Parable-SmolLM3-3B-Claude-Fable-5

Files

| File | Quant | Size | |

|---|---|---|---|

| Parable-SmolLM3-3B-Claude-Fable-5-GGUF-Q4_K_M.gguf | Q4_K_M | 1.9 GB | recommended default |

| Parable-SmolLM3-3B-Claude-Fable-5-GGUF-Q5_K_M.gguf | Q5_K_M | 2.2 GB | |

| Parable-SmolLM3-3B-Claude-Fable-5-GGUF-Q6_K.gguf | Q6_K | 2.5 GB | |

| Parable-SmolLM3-3B-Claude-Fable-5-GGUF-Q8_0.gguf | Q8_0 | 3.3 GB | |

| Parable-SmolLM3-3B-Claude-Fable-5-GGUF-F16.gguf | F16 | 6.2 GB | for re-quantizing |

Usage

llama-cli -m Parable-SmolLM3-3B-Claude-Fable-5-GGUF-Q4_K_M.gguf \
  -c 4096 -p "Write a python function that reverses a string." --temp 0.6

SmolLM3 is supported by current llama.cpp releases (brew, Ollama, LM Studio

builds included); no special build is required.

Output begins with a <think>...</think> reasoning block, then the answer.

If you are building on top of this model, parse and strip the think block

before showing text to end users. The chat template identifies the model as

"Parable, a coding assistant that reasons before it answers."

Model details

  • Base: HuggingFaceTB/SmolLM3-3B (3B, Apache-2.0, 64k context)
  • Method: MLX QLoRA on a 4-bit quantized base, 16 layers adapted,

6.7M trainable parameters (0.218%)

  • Data: 4,076 training rows from real Claude Fable 5 agent-session

traces plus gpt5.5-terminal transcripts, prepared at a 4,096-token window

(268 over-length rows dropped; 226/226 rows held out for validation/test)

  • Schedule: 1,200-iteration budget across a paused-and-resumed run;

best checkpoint selected on validation loss (iteration 200 of the final

segment, val 1.154)

Evaluation

| | Held-out trace test loss |

|---|---|

| SmolLM3-3B base | 1.889 |

| This model | 1.115 |

The tuned model fits the Fable-5 reasoning distribution 41% better by

held-out loss on a 226-row test split never seen in training. That is the

honest headline for what this fine-tune does; we do not claim general

benchmark gains.

This lane trains on trace data without a replay mix, so impact on general

coding benchmarks is unmeasured here. The series' technical report

(DOI: 10.5281/zenodo.21676407)

documents why that matters and what replay does about it.

Limitations

  • Training ran on a 4-bit quantized base (16 GB M1 constraint); the F16

merge and higher quants cannot exceed 4-bit-base quality.

  • Modest scale: one seed, loss-based evaluation, no external benchmark run

for this model yet.

  • Not trained for: multi-file repo navigation, vision, non-English.
  • Inherits SmolLM3-3B's knowledge cutoff. Treat generated commands as

drafts to review.

Quantization

Quantized with llama.cpp llama-quantize from the F16 merge.

Provenance & licensing

Fine-tuned from HuggingFaceTB/SmolLM3-3B (Apache-2.0). Training data:

Glint-Research/Fable-5-traces

(AGPL-3.0) and

Roman1111111/gpt5.5-terminal

(MIT). Because those traces originate from third-party assistants, the

providers' terms may apply to downstream training and distillation. If you

plan to build on this model commercially, confirm your use aligns with those

terms.

Citation

@misc{aglawe2026parable,
  author = {Aglawe, Ankit},
  title  = {Agent-Trace Fine-Tuning of Small Language Models under Constrained Compute},
  year   = {2026},
  doi    = {10.5281/zenodo.21676407},
  url    = {https://doi.org/10.5281/zenodo.21676407}
}

Acknowledgements

The SmolLM3 team at Hugging Face for the base model; Glint-Research and

Roman1111111 for the trace datasets; empero-ai for the recipe this series

iterates on.

Run AnkitAI/Parable-SmolLM3-3B-Claude-Fable-5-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models