quimmedes/Agnes-3.0-Flash-XYZ-GGUF overview
Agnes 3.0 Flash Preview GGUF quantizations GGUF quantizations of Agnes AI/Agnes 3.0 Flash https://huggingface.co/Agnes AI/Agnes 3.0 Flash the open weight Previ…
Runs locally from ~12.12 GB disk (16 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Agnes-3.0-Flash-Q3-v4-XYZ.gguf | GGUF | Q3 | 12.12 GB | Download |
| Agnes-3.0-Flash-Q3.5-v4-XYZ.gguf | GGUF | Q3 | 15.37 GB | Download |
| Agnes-3.0-Flash-Q4-v4-XYZ.gguf | GGUF | Q4 | 16.57 GB | Download |
| Agnes-3.0-Flash-Q4.5-v4-XYZ.gguf | GGUF | Q4 | 19.59 GB | Download |
| Agnes-3.0-Flash-Q5-v4-XYZ.gguf | GGUF | Q5 | 20.15 GB | Download |
| Agnes-3.0-Flash-Q5.5-v4-XYZ.gguf | GGUF | Q5 | 20.59 GB | Download |
| Agnes-3.0-Flash-Q6-v4-XYZ.gguf | GGUF | Q6 | 24.32 GB | Download |
| Agnes-3.0-Flash-Q7-v4-XYZ.gguf | GGUF | Q7 | 24.60 GB | Download |
| Agnes-3.0-Flash-Q8-v4-XYZ.gguf | GGUF | Q8 | 27.32 GB | Download |
| Agnes-3.0-Flash-Q9-v4-XYZ.gguf | GGUF | Q9 | 32.28 GB | Download |
| extra/Agnes-3.0-Flash-Q4.5-XYZ.gguf | GGUF | Q4 | 17.37 GB | Download |
| extra/Agnes-3.0-Flash-Q4_K_M-control.gguf | GGUF | Q4_K_M | 19.07 GB | Download |
| extra/Agnes-3.0-Flash-Q5-XYZ.gguf | GGUF | Q5 | 19.52 GB | Download |
| extra/Agnes-3.0-Flash-Q8-XYZ.gguf | GGUF | Q8 | 31.21 GB | Download |
Model Details
| Model ID | quimmedes/Agnes-3.0-Flash-XYZ-GGUF |
|---|---|
| Author | quimmedes |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | Agnes-AI/Agnes-3.0-Flash |
| Last modified | 2026-09-12T21:02:50.000Z |
Model README
---
license: apache-2.0
base_model: Agnes-AI/Agnes-3.0-Flash
tags:
- gguf
- quantization
- hybrid-attention
- ssm
- long-context
- imatrix
- text-generation
pipeline_tag: text-generation
---
Agnes-3.0-Flash Preview - GGUF quantizations
GGUF quantizations of Agnes-AI/Agnes-3.0-Flash
(the open-weight Preview checkpoint, Apache-2.0), quantized with
imatrix-weighted llama-quantize.
> These files need a patched build. Agnes has an extra parallel SwiGLU branch in every layer
> (ffn_gate_par / ffn_up_par / ffn_down_par) that stock llama.cpp does not know about.
> The fork cafe-llama.cpp adds those tensors to the
> qwen35 architecture (plus the recurrent/full attention layer map and the delta-attention V-head
> reorder). Loading these GGUFs on an unpatched build will fail or mis-load.
Main ladder
Oficial llama.cpp is not supporting this yet, you have to use cafe-llama.cpp
https://github.com/quimmedes/cafe-llama.cpp
| file | size on disk | answer when served | uploaded |
|---|---|---|---|
| Agnes-3.0-Flash-Q3-v4-XYZ.gguf | 12.12 GiB | The capital of France is Paris. | yes |
| Agnes-3.0-Flash-Q3.5-v4-XYZ.gguf | 15.37 GiB | The capital of France is Paris. | yes |
| Agnes-3.0-Flash-Q4-v4-XYZ.gguf | 16.57 GiB | The capital of France is Paris. | yes |
| Agnes-3.0-Flash-Q4.5-v4-XYZ.gguf | 19.59 GiB | The capital of France is Paris. | yes |
| Agnes-3.0-Flash-Q5-v4-XYZ.gguf | 20.15 GiB | The capital of France is Paris. | yes |
| Agnes-3.0-Flash-Q5.5-v4-XYZ.gguf | 20.59 GiB | The capital of France is Paris. | yes |
| Agnes-3.0-Flash-Q6-v4-XYZ.gguf | 24.32 GiB | The capital of France is Paris. | yes |
| Agnes-3.0-Flash-Q7-v4-XYZ.gguf | 24.60 GiB | The capital of France is Paris. | yes |
| Agnes-3.0-Flash-Q8-v4-XYZ.gguf | 27.32 GiB | The capital of France is Paris. | yes |
| Agnes-3.0-Flash-Q9-v4-XYZ.gguf | 32.28 GiB | The capital of France is Paris. | yes |
Total: 212.9 GiB - all of it uploaded.
Other quantizations (extra/)
| file | size on disk | answer when served |
|---|---|---|
| extra/Agnes-3.0-Flash-Q3-XYZ.gguf | 18.00 GiB | The capital of France is Paris. |
| extra/Agnes-3.0-Flash-Q4.5-XYZ.gguf | 17.37 GiB | The capital of France is Paris. |
| extra/Agnes-3.0-Flash-Q5-XYZ.gguf | 19.52 GiB | The capital of France is Paris. |
| extra/Agnes-3.0-Flash-Q8-XYZ.gguf | 31.21 GiB | The capital of France is Paris. |
| extra/Agnes-3.0-Flash-Q4_K_M-control.gguf | 19.07 GiB | The capital of France is Paris. |
Total: 105.2 GiB. Same model and same runtime as the ladder above; different tensor-type
allocation (more precision kept in attention, more compression in the FFN).
How to run
# 24 GB card, long context, MTP speculative decoding
llama-server -m Agnes-3.0-Flash-Q4.5-v4-XYZ.gguf \
--host 127.0.0.1 --port 8080 -ngl 20 -nr -c 32768 -np 1 \
-ctk q8_0 -ctv q8_0 -fa on --spec-type draft-mtp
-nr(no repack) keeps weights memory-mapped, which matters at 33B.- On a 24 GB card
-nglaround 20 fits; a full offload only leaves room for a small context. - The MTP block is included (
nextn, 1 layer), so--spec-type draft-mtpworks. - Base card's recommended sampling:
temp 1.0,top_p 0.95,top_k 20.
What this checkpoint actually is (findings, with the method)
While converting this model I compared it against Qwen/Qwen3.8-27B, because both report the same
architecture (qwen35). The bytes say the two are the same blueprint. Method: sha256 over HTTPS for
the small files, and byte-comparison of tensor ranges (safetensors data_offsets header, Range:
requests) for the weights.
Byte-identical:
| item | result |
|---|---|
| vocab.json | sha256 ce99b4cb2983 == Qwen3.8-27B |
| merges.txt | sha256 a9d356d7bdf1 == Qwen3.8-27B |
| chat_template.jinja | identical (same tool-call format, same reasoning_effort levels) |
| generation_config.json | identical |
| architecture hyper-parameters | hidden 5120, 24 q / 4 kv heads, head_dim 256, FFN 17408, delta-rule 16 key / 48 value heads, conv kernel 4, partial_rotary_factor 0.25, mrope [11,11,10] interleaved, theta 1e7, vocab 248320, MTP 1 layer, swish output gate, tie_word_embeddings=false |
| vision tower | same config (only the model_type label differs: agnes_vision vs qwen3_5), same weights (compared attn.proj of block 0: 2,654,208 bytes equal), same image/video token ids |
| embed_tokens.weight | identical (248320 x 5120) |
| norms | input_layernorm, post_attention_layernorm, q_norm, k_norm, final norm identical |
| SSM constants | A_log, dt_bias, conv1d, ssm_norm identical |
Different:
| item | Agnes | Qwen3.8-27B |
|---|---|---|
| layers | 72 (54 delta-rule + 18 global, 3:1) | 64 (48 linear + 16 full) |
| parallel FFN | SwiGLU 2048 in every layer | absent |
| all projection matrices | retrained | - |
The parameter count agrees: 27B x 72/64 = 30.4B, plus the parallel branch
(3 x 2048 x 5120 x 72 ~ 2.3B) = ~32.6B, and this GGUF measures 32.63B. Nothing is unaccounted for.
So: Agnes-3.0-Flash Preview carries the Qwen3.8-27B blueprint - same tokenizer, same chat
template, same vision tower (identical weights), same embeddings/norms/SSM constants - with 8 extra
delta layers, a new parallel-FFN branch, and retrained projections. Both models are Apache-2.0, so
this is entirely legitimate; the reason to write it down is that people know what they are downloading.
How these were made
- Safetensors -> GGUF with the fork's converter. Changes needed for this model: register the
parallel-FFN tensors, record the delta-attention layers as recurrent (otherwise every layer looks
like full attention and attn_qkv fails), and map .delta_attn./.global_attn. onto the standard
qwen35 names.
- Imatrix built from 50 chunks of the wikitext-2 raw test split; the bf16 GGUF scores
PPL 6.5737 +/- 0.14183 with it.
- Quantized with
llama-quantize --imatrix, per-tensor types: attention and SSM tensors stay at
higher precision, the FFN carries the compression. The parallel FFN follows the FFN of its block.
Practical consequences
- It is dense, not a MoE. No experts: 0 of the 1521 checkpoint tensors contain "expert", and the
config has no MoE keys (intermediate_size 17408 plus a parallel 2048 SwiGLU). All ~33B parameters
are read for every token, so expect dense-model speed at equal quantization, not "3B active" speed.
The base card's benchmark table compares against models labelled "X active"; Agnes is not one.
- Long context is comparatively cheap. Only 18 of the 72 layers hold a KV cache that grows with
context (the 54 delta-rule layers keep a fixed-size recurrent state). At 128k with q8_0 KV that is
about 5 GB (18 x 2 x 4 KV heads x 256 dim x 131072 x 1.06 B), which is why a 24 GB card can hold
the Q4.5 weights and a long context.
- Vision is not in these files. The text tower and the MTP layer are; the projector is separate.
Because the vision tower weights are byte-identical to Qwen3.8-27B's, the
mmproj-Qwen3.8-27B-bf16.gguf from quimmedes/Qwen3.8-27B-XYZ is expected to work - but a full
image encode could not be run on this machine (the vision graph needed more RAM than was free), so
treat that as unverified.
- These are quantizations of a preview checkpoint; the production Agnes model has a different
configuration and its benchmarks do not apply here.
Credits
- Base model: Agnes-AI/Agnes-3.0-Flash (Apache-2.0).
- Runtime: quimmedes/cafe-llama.cpp - required for the
parallel-FFN tensors.
Run quimmedes/Agnes-3.0-Flash-XYZ-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models