GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

prithivMLmods/Smaug-Mini-GGUF overview

Smaug Mini GGUF Smaug Mini is Abacus.AI's agentic finetune of Qwen3.8 27B, trained via on policy reinforcement learning GRPO, LoRA merged into the language tru…

transformersgguftext-generation-inferencellama-cppagenticsmaugabacusaiimage-text-to-textenbase_model:abacusai/Smaug-Minibase_model:quantized:abacusai/Smaug-Minilicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~888.0 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
1
Pipeline
image-text-to-text

Repository Files & Downloads

4 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Smaug-Mini.BF16.ggufGGUFGGUF50.11 GBDownload
Smaug-Mini.Q4_K_M.ggufGGUFGGUF15.41 GBDownload
Smaug-Mini.Q5_K_M.ggufGGUFGGUF17.91 GBDownload
Smaug-Mini.mmproj-bf16.ggufGGUFBF16888.0 MBDownload

Model Details

Model IDprithivMLmods/Smaug-Mini-GGUF
AuthorprithivMLmods
Pipelineimage-text-to-text
Licenseapache-2.0
Base modelabacusai/Smaug-Mini
Last modified2026-09-22T03:05:57.000Z

Model README

---

license: apache-2.0

language:

  • en

base_model:

  • abacusai/Smaug-Mini

pipeline_tag: image-text-to-text

library_name: transformers

tags:

  • text-generation-inference
  • llama-cpp
  • agentic
  • smaug
  • abacusai

---

Smaug-Mini-GGUF

> Smaug-Mini is Abacus.AI's agentic finetune of Qwen3.8-27B, trained via on-policy reinforcement learning (GRPO, LoRA-merged into the language trunk only, vision tower left bitwise-identical to the base) over multi-turn, tool-using automation episodes with verified, outcome-based rewards, targeting more reliable end-to-end tool use and automation performance while holding general capabilities at parity with the base model. It retains Qwen3.8-27B's architecture, layout, 262,144-token context, and xhigh/medium/low reasoning-effort interface exactly, functioning as a drop-in replacement, while delivering notable agentic gains — +4.5 on AutomationBench (41.8 vs. 37.3), +17.1 on JobBench (50.5 vs. 33.4), +13.5 on NL2Repo-Bench (55.8 vs. 42.3), and +2.0 overall on LiveBench (76.9 vs. 75.3) — alongside modest improvements on reasoning benchmarks like HLE and IFBench, with GPQA-diamond and MMMU-Pro held essentially at parity. Its key behavioral shift is redistributing deliberation rather than adding it: episodes finish about three steps sooner at roughly unchanged total reasoning volume, and episodes that exhaust their step budget without completing drop from 3.4% to 1.0%. It's served via vLLM with a Qwen3 reasoning parser and tool-call parser (recommended sampling: temperature 1.0, top_p 0.95, reasoning effort xhigh), with the inherited MTP head left untrained against the updated trunk — so speculative decoding via MTP should stay disabled — and is released under Apache 2.0, inherited from Qwen3.8-27B.

Model Files

| File Name | Quant Type | File Size | File Link | Description |

|-----------|------------|-----------|-----------|-------------|

| Smaug-Mini.BF16.gguf | BF16 | 53.8 GB | Link | Full BF16 weights. Highest quality, largest file size. |

| Smaug-Mini.Q4_K_M.gguf | Q4_K_M | 16.5 GB | Link | Good quality, default size for most use cases, recommended. |

| Smaug-Mini.Q5_K_M.gguf | Q5_K_M | 19.2 GB | Link | High quality, recommended. |

| Smaug-Mini.mmproj-bf16.gguf | mmproj-bf16 | 931 MB | Link | Multimodal projection file in BF16 format. Used for vision/language models. |

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Run prithivMLmods/Smaug-Mini-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models