GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

junwatu/Mellum2-12B-A2.5B-Instruct-GGUF overview

Mellum2 12B A2.5B Instruct GGUF This is a GGUF quantization of JetBrains/Mellum2 12B A2.5B Instruct https://huggingface.co/JetBrains/Mellum2 12B A2.5B Instruct…

ggufllama.cppmellum2moecodequantizedtext-generationbase_model:JetBrains/Mellum2-12B-A2.5B-Instructbase_model:quantized:JetBrains/Mellum2-12B-A2.5B-Instructlicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~7.52 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
53
Likes
2
Pipeline
text-generation
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Mellum2-12B-A2.5B-Instruct-Q4_K_M.ggufGGUFQ4_K_M7.52 GBDownload

Model Details

Model IDjunwatu/Mellum2-12B-A2.5B-Instruct-GGUF
Authorjunwatu
Pipelinetext-generation
Licenseapache-2.0
Base modelJetBrains/Mellum2-12B-A2.5B-Instruct
Last modified2026-07-22T01:13:55.000Z

Model README

---

license: apache-2.0

base_model: JetBrains/Mellum2-12B-A2.5B-Instruct

pipeline_tag: text-generation

tags:

- gguf

- llama.cpp

- mellum2

- moe

- code

- quantized

pinned: true

---

Mellum2 12B A2.5B Instruct GGUF

This is a GGUF quantization of JetBrains/Mellum2-12B-A2.5B-Instruct.

Model

Mellum2 is a Mixture-of-Experts model from JetBrains.

Key details:

Quantization

| Field | Value |

|---|---|

| File | Mellum2-12B-A2.5B-Instruct-Q4_K_M.gguf |

| Hugging Face file size | 8.1 GB |

The quantizer reported fallback quantization for 28 tensors. This happened because some Mellum2 expert tensors have width 896, which is not divisible by the block size required by some K-quant formats.

Practical meaning:

  • The model is labeled Q4_K_M.
  • Some tensors use fallback formats such as q5_0 or q8_0.
  • The final file is larger than a pure Q4 estimate.

Important Compatibility Warning

This GGUF requires a llama.cpp build with Mellum2 support.

This GGUF was converted and quantized with the Mellum2 PR branch below. If you use another llama.cpp build, verify that it includes Mellum2 support before loading the model.

Use the Mellum2 PR branch: Xarbirus/llama.cpp/tree/mellum2

Related upstream PR: ggml-org/llama.cpp#23966

Build a compatible llama.cpp:

git clone --branch mellum2 https://github.com/Xarbirus/llama.cpp
cd llama.cpp
cmake -B build
cmake --build build --config Release -j

Local Usage

Example:

./build/bin/llama-cli \
  -m ./Mellum2-12B-A2.5B-Instruct-Q4_K_M.gguf \
  -c 8192 \
  -ngl 99 \
  -p "Write a Python function that validates whether a string is a palindrome."

Runtime memory depends on context length, prompt size, backend, and machine memory. Adjust -c and -ngl for your hardware.

Links

License

This GGUF quantization follows the base model license: Apache 2.0

Base model: JetBrains/Mellum2-12B-A2.5B-Instruct

Check the original model card for the full license terms before redistribution or production use.

Run junwatu/Mellum2-12B-A2.5B-Instruct-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models