GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

NANI-Nithin/Mellum2-12B-A2.5B-Instruct-GGUF overview

Mellum2 12B A2.5B Instruct GGUF This repository contains GGUF quantizations of JetBrains/Mellum2 12B A2.5B Instruct for efficient local inference with llama.cp…

ggufllama.cppollamalm-studioquantizationtext-generationenbase_model:JetBrains/Mellum2-12B-A2.5B-Instructbase_model:quantized:JetBrains/Mellum2-12B-A2.5B-Instructlicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~7.52 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation

Repository Files & Downloads

4 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Mellum2-Q4_K_M.ggufGGUFQ4_K_M7.52 GBDownload
Mellum2-Q5_K_M.ggufGGUFQ5_K_M8.58 GBDownload
Mellum2-Q6_K.ggufGGUFQ6_K10.13 GBDownload
Mellum2-Q8_0.ggufGGUFQ8_012.04 GBDownload

Model Details

Model IDNANI-Nithin/Mellum2-12B-A2.5B-Instruct-GGUF
AuthorNANI-Nithin
Pipelinetext-generation
Licenseapache-2.0
Base modelJetBrains/Mellum2-12B-A2.5B-Instruct
Last modified2026-08-04T21:03:51.000Z

Model README

---

license: apache-2.0

base_model: JetBrains/Mellum2-12B-A2.5B-Instruct

library_name: gguf

tags:

- gguf

- llama.cpp

- ollama

- lm-studio

- quantization

pipeline_tag: text-generation

language:

- en

---

Mellum2-12B-A2.5B-Instruct-GGUF

This repository contains GGUF quantizations of JetBrains/Mellum2-12B-A2.5B-Instruct for efficient local inference with llama.cpp, Ollama, LM Studio, and other GGUF-compatible runtimes.

Model details

  • Base model: JetBrains/Mellum2-12B-A2.5B-Instruct
  • Format: GGUF
  • Architecture: Mellum2
  • Task type: Instruction-following assistant
  • Context length: 131,072 tokens
  • License: Apache 2.0

Available quantizations

| File | Quantization | Size | Notes |

|---|---:|---:|---|

| Mellum2-12B-A2.5B-Instruct-Q4_K_M.gguf | Q4_K_M | 8.07 GB | Lower-memory deployment with good quality-to-size tradeoff |

| Mellum2-12B-A2.5B-Instruct-Q5_K_M.gguf | Q5_K_M | 9.21 GB | Higher-quality 5-bit deployment with a modest size increase |

Intended use

This model is intended for:

  • local chat and assistant workflows
  • coding assistance
  • tool-use experiments
  • CPU-friendly or memory-constrained inference
  • users who want a balance between quality and memory usage

Notes

These quantizations were produced with llama.cpp. During conversion, some tensors may require fallback quantization depending on the model architecture, which is expected and does not prevent successful inference.

Usage

llama.cpp

./llama-server -m Mellum2-12B-A2.5B-Instruct-Q5_K_M.gguf --jinja --port 8000

Ollama

Create a Modelfile pointing to the GGUF file:

FROM ./Mellum2-12B-A2.5B-Instruct-Q4_K_M.gguf

Then run:

ollama create mellum2-q4 -f Modelfile
ollama run mellum2-q4

License

Released under the same license as the base model: Apache 2.0.

Acknowledgments

  • Base model: JetBrains/Mellum2-12B-A2.5B-Instruct
  • Quantization performed with llama.cpp

Run NANI-Nithin/Mellum2-12B-A2.5B-Instruct-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models