NANI-Nithin/Mellum2-12B-A2.5B-Instruct-GGUF overview
Mellum2 12B A2.5B Instruct GGUF This repository contains GGUF quantizations of JetBrains/Mellum2 12B A2.5B Instruct for efficient local inference with llama.cp…
Runs locally from ~7.52 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | NANI-Nithin/Mellum2-12B-A2.5B-Instruct-GGUF |
|---|---|
| Author | NANI-Nithin |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | JetBrains/Mellum2-12B-A2.5B-Instruct |
| Last modified | 2026-08-04T21:03:51.000Z |
Model README
---
license: apache-2.0
base_model: JetBrains/Mellum2-12B-A2.5B-Instruct
library_name: gguf
tags:
- gguf
- llama.cpp
- ollama
- lm-studio
- quantization
pipeline_tag: text-generation
language:
- en
---
Mellum2-12B-A2.5B-Instruct-GGUF
This repository contains GGUF quantizations of JetBrains/Mellum2-12B-A2.5B-Instruct for efficient local inference with llama.cpp, Ollama, LM Studio, and other GGUF-compatible runtimes.
Model details
- Base model: JetBrains/Mellum2-12B-A2.5B-Instruct
- Format: GGUF
- Architecture: Mellum2
- Task type: Instruction-following assistant
- Context length: 131,072 tokens
- License: Apache 2.0
Available quantizations
| File | Quantization | Size | Notes |
|---|---:|---:|---|
| Mellum2-12B-A2.5B-Instruct-Q4_K_M.gguf | Q4_K_M | 8.07 GB | Lower-memory deployment with good quality-to-size tradeoff |
| Mellum2-12B-A2.5B-Instruct-Q5_K_M.gguf | Q5_K_M | 9.21 GB | Higher-quality 5-bit deployment with a modest size increase |
Intended use
This model is intended for:
- local chat and assistant workflows
- coding assistance
- tool-use experiments
- CPU-friendly or memory-constrained inference
- users who want a balance between quality and memory usage
Notes
These quantizations were produced with llama.cpp. During conversion, some tensors may require fallback quantization depending on the model architecture, which is expected and does not prevent successful inference.
Usage
llama.cpp
./llama-server -m Mellum2-12B-A2.5B-Instruct-Q5_K_M.gguf --jinja --port 8000
Ollama
Create a Modelfile pointing to the GGUF file:
FROM ./Mellum2-12B-A2.5B-Instruct-Q4_K_M.gguf
Then run:
ollama create mellum2-q4 -f Modelfile
ollama run mellum2-q4
License
Released under the same license as the base model: Apache 2.0.
Acknowledgments
- Base model:
JetBrains/Mellum2-12B-A2.5B-Instruct - Quantization performed with
llama.cpp
Run NANI-Nithin/Mellum2-12B-A2.5B-Instruct-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models