abenzerps/MiniCPM5-2B-GGUF overview
MiniCPM5 2B GGUF GGUF quantizations of OpenBMB/MiniCPM5 2B https://huggingface.co/openbmb/MiniCPM5 2B , a 2B dense Llama based model for local deployment, codi…
Runs locally from ~925.9 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| MiniCPM5-2B-IQ2_M.gguf | GGUF | IQ2_M | 925.9 MB | Download |
| MiniCPM5-2B-IQ3_M.gguf | GGUF | IQ3_M | 1.14 GB | Download |
| MiniCPM5-2B-IQ4_XS.gguf | GGUF | IQ4_XS | 1.33 GB | Download |
| MiniCPM5-2B-Q2_K.gguf | GGUF | Q2_K | 991.8 MB | Download |
| MiniCPM5-2B-Q3_K_M.gguf | GGUF | Q3_K_M | 1.20 GB | Download |
| MiniCPM5-2B-Q4_0.gguf | GGUF | Q4_0 | 1.39 GB | Download |
| MiniCPM5-2B-Q4_K_M.gguf | GGUF | Q4_K_M | 1.45 GB | Download |
| MiniCPM5-2B-Q4_K_S.gguf | GGUF | Q4_K_S | 1.40 GB | Download |
| MiniCPM5-2B-Q5_K_M.gguf | GGUF | Q5_K_M | 1.68 GB | Download |
| MiniCPM5-2B-Q6_K.gguf | GGUF | Q6_K | 1.93 GB | Download |
| MiniCPM5-2B-Q8_0.gguf | GGUF | Q8_0 | 2.50 GB | Download |
Model Details
| Model ID | abenzerps/MiniCPM5-2B-GGUF |
|---|---|
| Author | abenzerps |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | openbmb/MiniCPM5-2B |
| Last modified | 2026-09-07T23:51:55.000Z |
Model README
---
license: apache-2.0
language:
- en
- zh
base_model: openbmb/MiniCPM5-2B
base_model_relation: quantized
pipeline_tag: text-generation
library_name: gguf
tags:
- gguf
- llama.cpp
- minicpm5
- long-context
- 131k-context
- dense
- tool-calling
---
MiniCPM5-2B GGUF
GGUF quantizations of OpenBMB/MiniCPM5-2B, a 2B dense Llama-based model for local deployment, coding, reasoning, long-context work, and tool use. The source checkpoint supports a native context length of 131,072 tokens (128K).
Benchmarks
!MiniCPM5-2B evaluation results
Benchmark results reported by OpenBMB for the original MiniCPM5-2B checkpoint.
Capability comparison reported by OpenBMB for the original MiniCPM5-2B checkpoint.
GGUF files
| Quantization | File | Size |
| --- | --- | ---: |
| Q2_K | MiniCPM5-2B-Q2_K.gguf | 1.04 GB |
| Q3_K_M | MiniCPM5-2B-Q3_K_M.gguf | 1.29 GB |
| Q4_0 | MiniCPM5-2B-Q4_0.gguf | 1.49 GB |
| Q4_K_S | MiniCPM5-2B-Q4_K_S.gguf | 1.50 GB |
| Q4_K_M | MiniCPM5-2B-Q4_K_M.gguf | 1.56 GB |
| Q5_K_M | MiniCPM5-2B-Q5_K_M.gguf | 1.81 GB |
| Q6_K | MiniCPM5-2B-Q6_K.gguf | 2.07 GB |
| Q8_0 | MiniCPM5-2B-Q8_0.gguf | 2.68 GB |
| IQ2_M | MiniCPM5-2B-IQ2_M.gguf | 0.97 GB |
| IQ3_M | MiniCPM5-2B-IQ3_M.gguf | 1.23 GB |
| IQ4_XS | MiniCPM5-2B-IQ4_XS.gguf | 1.42 GB |
The model is text-only. No vision projector or MTP files are included. The IQ files use an importance matrix generated from WikiText-2 and are intended for recent llama.cpp builds. SHA-256 checksums are provided in SHA256SUMS.txt.
Chat template
The GGUF files embed the upstream chat template. chat_template.jinja is provided as an external copy for runtimes that require a separate template file.
Usage
Use a current llama.cpp build with MiniCPM5 support. The example below uses an 8K context; increase -c up to 131072 when sufficient memory is available.
llama-cli \
-m MiniCPM5-2B-Q4_K_M.gguf \
-c 8192 --jinja \
--temp 1.0 --top-p 0.95 \
-p "Explain why reproducible builds matter."
For an OpenAI-compatible server:
llama-server \
-m MiniCPM5-2B-Q4_K_M.gguf \
-c 8192 --jinja --host 0.0.0.0 --port 8080
Tool-call behavior depends on the serving runtime's parser and API integration; use the embedded template and verify tool calls in the target application.
Source
- Model: OpenBMB/MiniCPM5-2B
- Source revision:
3497c460c89e00520c3cfa2e73f49ab7647f1177 - Conversion: the original Q4_0–Q8_0 files use upstream llama.cpp commit
f114f91f9ed6792cf402437e3874adad98902744; the additional Q2_K, Q3_K_M, Q4_K_S, IQ2_M, IQ3_M, and IQ4_XS files use upstream commit67672dc5b76f8bc17785a19d3dc6d1463fc2902c - License: Apache-2.0
- Checksums: SHA256SUMS.txt
Run abenzerps/MiniCPM5-2B-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models