stamsam/Instella-Prometheus-GGUF overview
Instella Prometheus GGUF 🔥 Instella Prometheus prometheus.png The code fire, packed for local inference. This is the dedicated GGUF release of Instella Promet…
Runs locally from ~7.65 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | stamsam/Instella-Prometheus-GGUF |
|---|---|
| Author | stamsam |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | stamsam/Instella-Prometheus,amd/Instella-MoE-16B-A3B-SFT |
| Last modified | 2026-08-01T13:03:03.000Z |
Model README
---
license: apache-2.0
base_model:
- stamsam/Instella-Prometheus
- amd/Instella-MoE-16B-A3B-SFT
tags:
- gguf
- llama.cpp
- instella
- moe
- coding
- python
- text-generation
- quantized
pipeline_tag: text-generation
library_name: gguf
---
Instella-Prometheus GGUF 🔥
The code-fire, packed for local inference.
This is the dedicated GGUF release of Instella-Prometheus, a merged, standalone coding model built from AMD's Instella-MoE-16B-A3B-SFT. Prometheus is a 16B-total / ~2.8B-active MoE model, distilled for direct, code-first answers under a bare user prompt:
- No system prompt required
- No thinking-tag suppression or decoder bans
- No LoRA adapter or PEFT dependency
- Designed for llama.cpp and compatible GGUF runtimes
Choose your quantization
| Quantization | File | Approx. size | Use when |
|---|---|---:|---|
| Q8_0 | Instella-Prometheus-Q8_0.gguf | 16.9 GB | You want near-full-weight quality |
| Q4_K_M | Instella-Prometheus-Q4_K_M.gguf | 9.4 GB | You want the best quality/size balance |
| Q3_K_M | Instella-Prometheus-Q3_K_M.gguf | 8.2 GB | You need the lowest memory footprint of these three |
All files are in the repository root so the Hugging Face Hub can identify and display them as GGUF quantized variants.
llama.cpp
# Example: Q4_K_M
./llama-cli \
-m Instella-Prometheus-Q4_K_M.gguf \
-p "Write a Python function to merge overlapping intervals." \
-n 1024
For chat frontends, use the model's built-in GGUF metadata and chat template when supported by the runtime. The intended contract is simply: user task in, clean answer out.
What changed from base?
On a 12-task bare-user Python suite:
| Metric | Base Instella | Instella-Prometheus |
|---|---:|---:|
| Code blocks | 2 / 12 | 12 / 12 |
| Syntax valid | 2 / 12 | 12 / 12 |
| Functional passes | 2 / 12 | 10 / 12 |
| Think-tag leakage | 12 / 12 | 0 / 12 |
| Natural EOS | 2 / 12 | 12 / 12 |
The full Transformers release, training details, dataset provenance, and evaluation notes are available in the canonical model repository.
License
Apache-2.0. See the canonical repository for base-model and dataset acknowledgements.
The fire belongs to them. The torch is yours.
Run stamsam/Instella-Prometheus-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models