Serpen/Minimax-M3-MSA-GGUF overview
Model This repo contains specialized MoE quants for MiniMax M3 with the Indexer Tensors preserved at FP32. A BF16 MMPROJ file for image vision input has also b…
Runs locally from ~7.9 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| M3-mmproj.gguf | GGUF | GGUF | 1.62 GB | Download |
| Q4_K_M/MiniMax-M3-Q4_K_M-00001-of-00007.gguf | GGUF | Q4_K_M | 7.9 MB | Download |
| Q4_K_M/MiniMax-M3-Q4_K_M-00002-of-00007.gguf | GGUF | Q4_K_M | 46.10 GB | Download |
| Q4_K_M/MiniMax-M3-Q4_K_M-00003-of-00007.gguf | GGUF | Q4_K_M | 45.43 GB | Download |
| Q4_K_M/MiniMax-M3-Q4_K_M-00004-of-00007.gguf | GGUF | Q4_K_M | 45.55 GB | Download |
| Q4_K_M/MiniMax-M3-Q4_K_M-00005-of-00007.gguf | GGUF | Q4_K_M | 45.27 GB | Download |
| Q4_K_M/MiniMax-M3-Q4_K_M-00006-of-00007.gguf | GGUF | Q4_K_M | 45.43 GB | Download |
| Q4_K_M/MiniMax-M3-Q4_K_M-00007-of-00007.gguf | GGUF | Q4_K_M | 18.33 GB | Download |
| imatrix.gguf | GGUF | GGUF | 441.4 MB | Download |
Model Details
Model README
---
base_model:
- MiniMaxAI/MiniMax-M3
---
Model
This repo contains specialized MoE-quants for MiniMax-M3 with the Indexer Tensors preserved at FP32. A BF16 MMPROJ file for image vision input has also been provided.
The text model files should load on mainline as this PR got merged. The MMPROJ requires this one, which is a superset of the first PR.
| Quant | Size | Mixture | PPL | 1-(Mean PPL(Q)/PPL(base)) | KLD |
| :----- | :-------------------- | :------------------------ | :------------------ | :------------------------ | :------------------ |
| Q4_K_M | 246.11 GiB (4.96 BPW) | Q8_0 / Q4_K / Q4_K / Q5_K | 5.287961 ± 0.035451 | +1.9814% | 0.069890 ± 0.000978 |
Run Serpen/Minimax-M3-MSA-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models