junwatu/Mellum2-12B-A2.5B-Instruct-GGUF overview
Mellum2 12B A2.5B Instruct GGUF This is a GGUF quantization of JetBrains/Mellum2 12B A2.5B Instruct https://huggingface.co/JetBrains/Mellum2 12B A2.5B Instruct…
Runs locally from ~7.52 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Mellum2-12B-A2.5B-Instruct-Q4_K_M.gguf | GGUF | Q4_K_M | 7.52 GB | Download |
Model Details
| Model ID | junwatu/Mellum2-12B-A2.5B-Instruct-GGUF |
|---|---|
| Author | junwatu |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | JetBrains/Mellum2-12B-A2.5B-Instruct |
| Last modified | 2026-07-22T01:13:55.000Z |
Model README
---
license: apache-2.0
base_model: JetBrains/Mellum2-12B-A2.5B-Instruct
pipeline_tag: text-generation
tags:
- gguf
- llama.cpp
- mellum2
- moe
- code
- quantized
pinned: true
---
Mellum2 12B A2.5B Instruct GGUF
This is a GGUF quantization of JetBrains/Mellum2-12B-A2.5B-Instruct.
Model
Mellum2 is a Mixture-of-Experts model from JetBrains.
Key details:
- Total parameters:
12B - Active parameters per token:
2.5B - Architecture: MoE
- Experts:
64 - Active experts per token:
8 - Context length:
131,072tokens - License: Apache 2.0
- Original model: JetBrains/Mellum2-12B-A2.5B-Instruct
- Model collection: JetBrains Mellum 2
Quantization
| Field | Value |
|---|---|
| File | Mellum2-12B-A2.5B-Instruct-Q4_K_M.gguf |
| Hugging Face file size | 8.1 GB |
The quantizer reported fallback quantization for 28 tensors. This happened because some Mellum2 expert tensors have width 896, which is not divisible by the block size required by some K-quant formats.
Practical meaning:
- The model is labeled
Q4_K_M. - Some tensors use fallback formats such as
q5_0orq8_0. - The final file is larger than a pure Q4 estimate.
Important Compatibility Warning
This GGUF requires a llama.cpp build with Mellum2 support.
This GGUF was converted and quantized with the Mellum2 PR branch below. If you use another llama.cpp build, verify that it includes Mellum2 support before loading the model.
Use the Mellum2 PR branch: Xarbirus/llama.cpp/tree/mellum2
Related upstream PR: ggml-org/llama.cpp#23966
Build a compatible llama.cpp:
git clone --branch mellum2 https://github.com/Xarbirus/llama.cpp
cd llama.cpp
cmake -B build
cmake --build build --config Release -j
Local Usage
Example:
./build/bin/llama-cli \
-m ./Mellum2-12B-A2.5B-Instruct-Q4_K_M.gguf \
-c 8192 \
-ngl 99 \
-p "Write a Python function that validates whether a string is a palindrome."
Runtime memory depends on context length, prompt size, backend, and machine memory. Adjust -c and -ngl for your hardware.
Links
- Base model: JetBrains/Mellum2-12B-A2.5B-Instruct
- Mellum2 collection: JetBrains Mellum 2
- Compatible
llama.cppbranch: Xarbirus/llama.cpp/tree/mellum2 - Upstream
llama.cppPR: ggml-org/llama.cpp#23966 llama.cppproject: ggml-org/llama.cpp- License: Apache License 2.0
License
This GGUF quantization follows the base model license: Apache 2.0
Base model: JetBrains/Mellum2-12B-A2.5B-Instruct
Check the original model card for the full license terms before redistribution or production use.
Run junwatu/Mellum2-12B-A2.5B-Instruct-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models