GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

gkraker04/Nanbeige4.2-3B-GGUF overview

license: apache 2.0 language: en zh library name: gguf pipeline tag: text generation tags: nanbeige gguf quantized 3b llama.cpp imatrix base model: Nanbeige/Na…

ggufnanbeigequantized3bllama.cppimatrixtext-generationenzhbase_model:Nanbeige/Nanbeige4.2-3Bbase_model:quantized:Nanbeige/Nanbeige4.2-3Blicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~2.7 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
27,645
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

26 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Nanbeige4.2-3B-IQ2_M.ggufGGUFIQ2_M1.55 GBDownload
Nanbeige4.2-3B-IQ2_S.ggufGGUFIQ2_S1.48 GBDownload
Nanbeige4.2-3B-IQ2_XS.ggufGGUFIQ2_XS1.36 GBDownload
Nanbeige4.2-3B-IQ3_M.ggufGGUFIQ3_M1.94 GBDownload
Nanbeige4.2-3B-IQ3_S.ggufGGUFIQ3_S1.87 GBDownload
Nanbeige4.2-3B-IQ3_XS.ggufGGUFIQ3_XS1.80 GBDownload
Nanbeige4.2-3B-IQ3_XXS.ggufGGUFIQ3_XXS1.66 GBDownload
Nanbeige4.2-3B-IQ4_NL.ggufGGUFIQ4_NL2.32 GBDownload
Nanbeige4.2-3B-IQ4_XS.ggufGGUFIQ4_XS2.21 GBDownload
Nanbeige4.2-3B-Q2_K.ggufGGUFQ2_K1.64 GBDownload
Nanbeige4.2-3B-Q2_K_S.ggufGGUFQ2_K_S1.56 GBDownload
Nanbeige4.2-3B-Q3_K_L.ggufGGUFQ3_K_L2.15 GBDownload
Nanbeige4.2-3B-Q3_K_M.ggufGGUFQ3_K_M2.02 GBDownload
Nanbeige4.2-3B-Q3_K_S.ggufGGUFQ3_K_S1.86 GBDownload
Nanbeige4.2-3B-Q4_0.ggufGGUFQ4_02.32 GBDownload
Nanbeige4.2-3B-Q4_1.ggufGGUFQ4_12.52 GBDownload
Nanbeige4.2-3B-Q4_K_M.ggufGGUFQ4_K_M2.40 GBDownload
Nanbeige4.2-3B-Q4_K_S.ggufGGUFQ4_K_S2.33 GBDownload
Nanbeige4.2-3B-Q5_0.ggufGGUFQ5_02.75 GBDownload
Nanbeige4.2-3B-Q5_1.ggufGGUFQ5_12.95 GBDownload
Nanbeige4.2-3B-Q5_K_M.ggufGGUFQ5_K_M2.78 GBDownload
Nanbeige4.2-3B-Q5_K_S.ggufGGUFQ5_K_S2.74 GBDownload
Nanbeige4.2-3B-Q6_K.ggufGGUFQ6_K3.19 GBDownload
Nanbeige4.2-3B-Q8_0.ggufGGUFQ8_04.13 GBDownload
Nanbeige4.2-3B-bf16.ggufGGUFBF167.77 GBDownload
imatrix.ggufGGUFGGUF2.7 MBDownload

Model Details

Model IDgkraker04/Nanbeige4.2-3B-GGUF
Authorgkraker04
Pipelinetext-generation
Licenseapache-2.0
Base modelNanbeige/Nanbeige4.2-3B
Last modified2026-08-10T13:48:07.000Z

Model README

---

license: apache-2.0

language:

  • en
  • zh

library_name: gguf

pipeline_tag: text-generation

tags:

  • nanbeige
  • gguf
  • quantized
  • 3b
  • llama.cpp
  • imatrix

base_model: Nanbeige/Nanbeige4.2-3B

---

Nanbeige4.2-3B GGUF

GGUF quantizations of Nanbeige4.2-3B, a compact 3B agentic model built on a Looped Transformer architecture.

These quants are built for buun-llama-cpp, the fork that added Nanbeige architecture support. They'll work with any GGUF-compatible runtime that supports Nanbeige models.

Produced with importance-matrix-guided quantization (llama-quantize --imatrix) from the bf16 source, ordered highest to lowest bits-per-weight.

Which Quant Should I Use?

| Use Case | Recommended Quant | Size | Why |

|----------|------------------|------|-----|

| Maximum quality | Q8_0 or Q6_K | 3.2–4.1 GB | Virtually lossless |

| Best tradeoff | Q4_K_M | 2.4 GB | Sweet spot: near-lossless quality, tiny footprint |

| Memory constrained | Q3_K_M or IQ3_M | 1.9–2.0 GB | Still very capable |

| Edge/mobile | IQ2_M or Q2_K | 1.5–1.6 GB | Noticeable but functional |

Quantization Ladder

| Quant | Size | bpw |

|-------|------|-----|

| bf16 | 7.77 GB | 16.00 |

| Q8_0 | 4.13 GB | 8.00 |

| Q6_K | 3.19 GB | 6.00 |

| Q5_1 | 2.95 GB | 5.65 |

| Q5_K_M | 2.78 GB | 5.50 |

| Q5_K_S | 2.74 GB | 5.21 |

| Q5_0 | 2.75 GB | 5.21 |

| Q4_1 | 2.52 GB | 4.78 |

| Q4_K_M | 2.40 GB | 4.58 |

| Q4_K_S | 2.33 GB | 4.37 |

| IQ4_NL | 2.32 GB | 4.50 |

| Q4_0 | 2.32 GB | 4.34 |

| IQ4_XS | 2.21 GB | 4.25 |

| Q3_K_L | 2.15 GB | 4.03 |

| Q3_K_M | 2.02 GB | 3.74 |

| IQ3_M | 1.94 GB | 3.66 |

| Q3_K_S | 1.86 GB | 3.41 |

| IQ3_S | 1.87 GB | 3.44 |

| IQ3_XS | 1.80 GB | 3.30 |

| IQ3_XXS | 1.66 GB | 3.06 |

| Q2_K | 1.64 GB | 2.96 |

| Q2_K_S | 1.56 GB | 2.96 |

| IQ2_M | 1.55 GB | 2.70 |

| IQ2_S | 1.48 GB | 2.50 |

| IQ2_XS | 1.36 GB | 2.31 |

Quality Benchmarks

Quality measured with MMLU-Pro (128 random questions across 14 categories, 5-shot CoT) and MuSR (128 multi-step reasoning questions). CoT reasoning enabled server-side via --reasoning on. Recovery = accuracy relative to bf16 baseline.

!Accuracy Chart

_Results filling in as evaluation progresses._

| Quant | Size | MMLU-Pro | MuSR |

|-------|------|----------|------|

| bf16 | 7.77 GB | 64.8% | 35.2% |

| Q8_0 | 4.13 GB | 64.1% | 34.4% |

| Q6_K | 3.19 GB | 64.8% | 35.2% |

| Q5_1 | 2.95 GB | 64.8% | 34.4% |

| Q5_K_M | 2.78 GB | 64.1% | 38.3% |

| Q5_K_S | 2.74 GB | 69.5% | 31.2% |

| Q5_0 | 2.75 GB | 60.9% | 36.7% |

| Q4_1 | 2.52 GB | 67.2% | 33.6% |

| Q4_K_M | 2.40 GB | 62.5% | 39.1% |

| Q4_K_S | 2.33 GB | 64.1% | 39.8% |

| IQ4_NL | 2.32 GB | 61.7% | 33.6% |

| Q4_0 | 2.32 GB | 68.0% | 35.9% |

| IQ4_XS | 2.21 GB | 65.6% | 35.2% |

| Q3_K_L | 2.15 GB | 64.1% | 7.8% |

| Q3_K_M | 2.02 GB | 62.5% | 10.2% |

| IQ3_M | 1.94 GB | 60.9% | 25.0% |

| Q3_K_S | 1.86 GB | 57.8% | 11.7% |

| IQ3_S | 1.87 GB | 58.6% | 18.8% |

| IQ3_XS | 1.80 GB | 57.8% | 7.0% |

| IQ3_XXS | 1.66 GB | 50.8% | 19.5% |

| Q2_K | 1.64 GB | 49.2% | 32.0% |

| Q2_K_S | 1.56 GB | 39.1% | 3.9% |

| IQ2_M | 1.55 GB | 48.4% | 6.2% |

| IQ2_S | 1.48 GB | 43.0% | 4.7% |

| IQ2_XS | 1.36 GB | 38.3% | 3.9% |

Perplexity Reference

On the calibration set (107 chunks, 4K context):

| Quant | PPL | Δ vs bf16 |

|-------|-----|-----------|

| bf16 | 11.64 | — |

| Q8_0 | 11.71 | +0.07 |

| Q4_K_M | 11.68 | +0.04 |

| Q3_K_M | 12.57 | +0.93 |

| Q2_K | 13.89 | +2.24 |

Note: perplexity measures next-token prediction loss, not task accuracy. Q4_K_M is functionally near-lossless for this model.

Usage

llama.cpp Server

llama-server -m Nanbeige4.2-3B-Q4_K_M.gguf \
  --flash-attn on \
  --cache-type-k vbr \
  --batch-size 64 --ubatch-size 64 \
  --parallel 1 --port 8081

CLI

llama-cli -m Nanbeige4.2-3B-Q4_K_M.gguf \
  -p "Write a Python function that..." -n 512

Recommended Settings

See the original model page for recommended usage settings.

References

License

Apache 2.0

Run gkraker04/Nanbeige4.2-3B-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models