gkraker04/Nanbeige4.2-3B-GGUF overview
license: apache 2.0 language: en zh library name: gguf pipeline tag: text generation tags: nanbeige gguf quantized 3b llama.cpp imatrix base model: Nanbeige/Na…
Runs locally from ~2.7 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Nanbeige4.2-3B-IQ2_M.gguf | GGUF | IQ2_M | 1.55 GB | Download |
| Nanbeige4.2-3B-IQ2_S.gguf | GGUF | IQ2_S | 1.48 GB | Download |
| Nanbeige4.2-3B-IQ2_XS.gguf | GGUF | IQ2_XS | 1.36 GB | Download |
| Nanbeige4.2-3B-IQ3_M.gguf | GGUF | IQ3_M | 1.94 GB | Download |
| Nanbeige4.2-3B-IQ3_S.gguf | GGUF | IQ3_S | 1.87 GB | Download |
| Nanbeige4.2-3B-IQ3_XS.gguf | GGUF | IQ3_XS | 1.80 GB | Download |
| Nanbeige4.2-3B-IQ3_XXS.gguf | GGUF | IQ3_XXS | 1.66 GB | Download |
| Nanbeige4.2-3B-IQ4_NL.gguf | GGUF | IQ4_NL | 2.32 GB | Download |
| Nanbeige4.2-3B-IQ4_XS.gguf | GGUF | IQ4_XS | 2.21 GB | Download |
| Nanbeige4.2-3B-Q2_K.gguf | GGUF | Q2_K | 1.64 GB | Download |
| Nanbeige4.2-3B-Q2_K_S.gguf | GGUF | Q2_K_S | 1.56 GB | Download |
| Nanbeige4.2-3B-Q3_K_L.gguf | GGUF | Q3_K_L | 2.15 GB | Download |
| Nanbeige4.2-3B-Q3_K_M.gguf | GGUF | Q3_K_M | 2.02 GB | Download |
| Nanbeige4.2-3B-Q3_K_S.gguf | GGUF | Q3_K_S | 1.86 GB | Download |
| Nanbeige4.2-3B-Q4_0.gguf | GGUF | Q4_0 | 2.32 GB | Download |
| Nanbeige4.2-3B-Q4_1.gguf | GGUF | Q4_1 | 2.52 GB | Download |
| Nanbeige4.2-3B-Q4_K_M.gguf | GGUF | Q4_K_M | 2.40 GB | Download |
| Nanbeige4.2-3B-Q4_K_S.gguf | GGUF | Q4_K_S | 2.33 GB | Download |
| Nanbeige4.2-3B-Q5_0.gguf | GGUF | Q5_0 | 2.75 GB | Download |
| Nanbeige4.2-3B-Q5_1.gguf | GGUF | Q5_1 | 2.95 GB | Download |
| Nanbeige4.2-3B-Q5_K_M.gguf | GGUF | Q5_K_M | 2.78 GB | Download |
| Nanbeige4.2-3B-Q5_K_S.gguf | GGUF | Q5_K_S | 2.74 GB | Download |
| Nanbeige4.2-3B-Q6_K.gguf | GGUF | Q6_K | 3.19 GB | Download |
| Nanbeige4.2-3B-Q8_0.gguf | GGUF | Q8_0 | 4.13 GB | Download |
| Nanbeige4.2-3B-bf16.gguf | GGUF | BF16 | 7.77 GB | Download |
| imatrix.gguf | GGUF | GGUF | 2.7 MB | Download |
Model Details
| Model ID | gkraker04/Nanbeige4.2-3B-GGUF |
|---|---|
| Author | gkraker04 |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | Nanbeige/Nanbeige4.2-3B |
| Last modified | 2026-08-10T13:48:07.000Z |
Model README
---
license: apache-2.0
language:
- en
- zh
library_name: gguf
pipeline_tag: text-generation
tags:
- nanbeige
- gguf
- quantized
- 3b
- llama.cpp
- imatrix
base_model: Nanbeige/Nanbeige4.2-3B
---
Nanbeige4.2-3B GGUF
GGUF quantizations of Nanbeige4.2-3B, a compact 3B agentic model built on a Looped Transformer architecture.
These quants are built for buun-llama-cpp, the fork that added Nanbeige architecture support. They'll work with any GGUF-compatible runtime that supports Nanbeige models.
Produced with importance-matrix-guided quantization (llama-quantize --imatrix) from the bf16 source, ordered highest to lowest bits-per-weight.
Which Quant Should I Use?
| Use Case | Recommended Quant | Size | Why |
|----------|------------------|------|-----|
| Maximum quality | Q8_0 or Q6_K | 3.2–4.1 GB | Virtually lossless |
| Best tradeoff | Q4_K_M | 2.4 GB | Sweet spot: near-lossless quality, tiny footprint |
| Memory constrained | Q3_K_M or IQ3_M | 1.9–2.0 GB | Still very capable |
| Edge/mobile | IQ2_M or Q2_K | 1.5–1.6 GB | Noticeable but functional |
Quantization Ladder
| Quant | Size | bpw |
|-------|------|-----|
| bf16 | 7.77 GB | 16.00 |
| Q8_0 | 4.13 GB | 8.00 |
| Q6_K | 3.19 GB | 6.00 |
| Q5_1 | 2.95 GB | 5.65 |
| Q5_K_M | 2.78 GB | 5.50 |
| Q5_K_S | 2.74 GB | 5.21 |
| Q5_0 | 2.75 GB | 5.21 |
| Q4_1 | 2.52 GB | 4.78 |
| Q4_K_M | 2.40 GB | 4.58 |
| Q4_K_S | 2.33 GB | 4.37 |
| IQ4_NL | 2.32 GB | 4.50 |
| Q4_0 | 2.32 GB | 4.34 |
| IQ4_XS | 2.21 GB | 4.25 |
| Q3_K_L | 2.15 GB | 4.03 |
| Q3_K_M | 2.02 GB | 3.74 |
| IQ3_M | 1.94 GB | 3.66 |
| Q3_K_S | 1.86 GB | 3.41 |
| IQ3_S | 1.87 GB | 3.44 |
| IQ3_XS | 1.80 GB | 3.30 |
| IQ3_XXS | 1.66 GB | 3.06 |
| Q2_K | 1.64 GB | 2.96 |
| Q2_K_S | 1.56 GB | 2.96 |
| IQ2_M | 1.55 GB | 2.70 |
| IQ2_S | 1.48 GB | 2.50 |
| IQ2_XS | 1.36 GB | 2.31 |
Quality Benchmarks
Quality measured with MMLU-Pro (128 random questions across 14 categories, 5-shot CoT) and MuSR (128 multi-step reasoning questions). CoT reasoning enabled server-side via --reasoning on. Recovery = accuracy relative to bf16 baseline.
_Results filling in as evaluation progresses._
| Quant | Size | MMLU-Pro | MuSR |
|-------|------|----------|------|
| bf16 | 7.77 GB | 64.8% | 35.2% |
| Q8_0 | 4.13 GB | 64.1% | 34.4% |
| Q6_K | 3.19 GB | 64.8% | 35.2% |
| Q5_1 | 2.95 GB | 64.8% | 34.4% |
| Q5_K_M | 2.78 GB | 64.1% | 38.3% |
| Q5_K_S | 2.74 GB | 69.5% | 31.2% |
| Q5_0 | 2.75 GB | 60.9% | 36.7% |
| Q4_1 | 2.52 GB | 67.2% | 33.6% |
| Q4_K_M | 2.40 GB | 62.5% | 39.1% |
| Q4_K_S | 2.33 GB | 64.1% | 39.8% |
| IQ4_NL | 2.32 GB | 61.7% | 33.6% |
| Q4_0 | 2.32 GB | 68.0% | 35.9% |
| IQ4_XS | 2.21 GB | 65.6% | 35.2% |
| Q3_K_L | 2.15 GB | 64.1% | 7.8% |
| Q3_K_M | 2.02 GB | 62.5% | 10.2% |
| IQ3_M | 1.94 GB | 60.9% | 25.0% |
| Q3_K_S | 1.86 GB | 57.8% | 11.7% |
| IQ3_S | 1.87 GB | 58.6% | 18.8% |
| IQ3_XS | 1.80 GB | 57.8% | 7.0% |
| IQ3_XXS | 1.66 GB | 50.8% | 19.5% |
| Q2_K | 1.64 GB | 49.2% | 32.0% |
| Q2_K_S | 1.56 GB | 39.1% | 3.9% |
| IQ2_M | 1.55 GB | 48.4% | 6.2% |
| IQ2_S | 1.48 GB | 43.0% | 4.7% |
| IQ2_XS | 1.36 GB | 38.3% | 3.9% |
Perplexity Reference
On the calibration set (107 chunks, 4K context):
| Quant | PPL | Δ vs bf16 |
|-------|-----|-----------|
| bf16 | 11.64 | — |
| Q8_0 | 11.71 | +0.07 |
| Q4_K_M | 11.68 | +0.04 |
| Q3_K_M | 12.57 | +0.93 |
| Q2_K | 13.89 | +2.24 |
Note: perplexity measures next-token prediction loss, not task accuracy. Q4_K_M is functionally near-lossless for this model.
Usage
llama.cpp Server
llama-server -m Nanbeige4.2-3B-Q4_K_M.gguf \
--flash-attn on \
--cache-type-k vbr \
--batch-size 64 --ubatch-size 64 \
--parallel 1 --port 8081
CLI
llama-cli -m Nanbeige4.2-3B-Q4_K_M.gguf \
-p "Write a Python function that..." -n 512
Recommended Settings
See the original model page for recommended usage settings.
References
License
Apache 2.0
Run gkraker04/Nanbeige4.2-3B-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models