GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

AesSedai/MiniMax-M3-GGUF overview

This repo contains specialized MoE quants for MiniMaxAI/MiniMax M3. The idea being that given the huge size of the FFN tensors compared to the rest of the tens…

ggufbase_model:MiniMaxAI/MiniMax-M3base_model:quantized:MiniMaxAI/MiniMax-M3endpoints_compatibleregion:usimatrixconversational

Runs locally from ~7.9 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
6
Pipeline
Author

Repository Files & Downloads

34 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
IQ2_S/MiniMax-M3-IQ2_S-00001-of-00004.ggufGGUFIQ2_S7.9 MBDownload
IQ2_S/MiniMax-M3-IQ2_S-00002-of-00004.ggufGGUFIQ2_S46.51 GBDownload
IQ2_S/MiniMax-M3-IQ2_S-00003-of-00004.ggufGGUFIQ2_S46.09 GBDownload
IQ2_S/MiniMax-M3-IQ2_S-00004-of-00004.ggufGGUFIQ2_S41.40 GBDownload
IQ3_S/MiniMax-M3-IQ3_S-00001-of-00005.ggufGGUFIQ3_S7.9 MBDownload
IQ3_S/MiniMax-M3-IQ3_S-00002-of-00005.ggufGGUFIQ3_S46.09 GBDownload
IQ3_S/MiniMax-M3-IQ3_S-00003-of-00005.ggufGGUFIQ3_S45.91 GBDownload
IQ3_S/MiniMax-M3-IQ3_S-00004-of-00005.ggufGGUFIQ3_S45.91 GBDownload
IQ3_S/MiniMax-M3-IQ3_S-00005-of-00005.ggufGGUFIQ3_S10.12 GBDownload
IQ4_XS/MiniMax-M3-IQ4_XS-00001-of-00006.ggufGGUFIQ4_XS7.9 MBDownload
IQ4_XS/MiniMax-M3-IQ4_XS-00002-of-00006.ggufGGUFIQ4_XS46.46 GBDownload
IQ4_XS/MiniMax-M3-IQ4_XS-00003-of-00006.ggufGGUFIQ4_XS45.80 GBDownload
IQ4_XS/MiniMax-M3-IQ4_XS-00004-of-00006.ggufGGUFIQ4_XS45.80 GBDownload
IQ4_XS/MiniMax-M3-IQ4_XS-00005-of-00006.ggufGGUFIQ4_XS45.80 GBDownload
IQ4_XS/MiniMax-M3-IQ4_XS-00006-of-00006.ggufGGUFIQ4_XS5.25 GBDownload
Q4_K_M/MiniMax-M3-Q4_K_M-00001-of-00007.ggufGGUFQ4_K_M7.9 MBDownload
Q4_K_M/MiniMax-M3-Q4_K_M-00002-of-00007.ggufGGUFQ4_K_M46.10 GBDownload
Q4_K_M/MiniMax-M3-Q4_K_M-00003-of-00007.ggufGGUFQ4_K_M45.43 GBDownload
Q4_K_M/MiniMax-M3-Q4_K_M-00004-of-00007.ggufGGUFQ4_K_M45.55 GBDownload
Q4_K_M/MiniMax-M3-Q4_K_M-00005-of-00007.ggufGGUFQ4_K_M45.27 GBDownload
Q4_K_M/MiniMax-M3-Q4_K_M-00006-of-00007.ggufGGUFQ4_K_M45.43 GBDownload
Q4_K_M/MiniMax-M3-Q4_K_M-00007-of-00007.ggufGGUFQ4_K_M18.33 GBDownload
Q5_K_M/MiniMax-M3-Q5_K_M-00001-of-00008.ggufGGUFQ5_K_M7.9 MBDownload
Q5_K_M/MiniMax-M3-Q5_K_M-00002-of-00008.ggufGGUFQ5_K_M46.34 GBDownload
Q5_K_M/MiniMax-M3-Q5_K_M-00003-of-00008.ggufGGUFQ5_K_M46.07 GBDownload
Q5_K_M/MiniMax-M3-Q5_K_M-00004-of-00008.ggufGGUFQ5_K_M46.07 GBDownload
Q5_K_M/MiniMax-M3-Q5_K_M-00005-of-00008.ggufGGUFQ5_K_M46.07 GBDownload
Q5_K_M/MiniMax-M3-Q5_K_M-00006-of-00008.ggufGGUFQ5_K_M46.07 GBDownload
Q5_K_M/MiniMax-M3-Q5_K_M-00007-of-00008.ggufGGUFQ5_K_M46.07 GBDownload
Q5_K_M/MiniMax-M3-Q5_K_M-00008-of-00008.ggufGGUFQ5_K_M18.51 GBDownload
mmproj-MiniMax-M3-BF16.ggufGGUFBF161.62 GBDownload
mmproj-MiniMax-M3-F16.ggufGGUFF161.61 GBDownload
mmproj-MiniMax-M3-F32.ggufGGUFF323.22 GBDownload
mmproj-MiniMax-M3-Q8_0.ggufGGUFQ8_0882.9 MBDownload

Model Details

Model IDAesSedai/MiniMax-M3-GGUF
AuthorAesSedai
Pipeline
License
Base modelMiniMaxAI/MiniMax-M3
Last modified2026-07-27T05:29:10.000Z

Model README

---

base_model:

  • MiniMaxAI/MiniMax-M3

---

This repo contains specialized MoE-quants for MiniMaxAI/MiniMax-M3. The idea being that given the huge size of the FFN tensors compared to the rest of the tensors in the model, it should be possible to achieve a better quality while keeping the overall size of the entire model smaller compared to a similar naive quantization. To that end, the quantization type default is kept in high quality and the FFN UP + FFN GATE tensors are quanted down along with the FFN DOWN tensors.

Additionally, I've provided a J-Space lens in the lens/ folder, trained on ~128 text prompts.

| Quant | Size | Mixture | PPL | 1-(Mean PPL(Q)/PPL(base)) | KLD |

| :----- | :-------------------- | :------------------------------- | :------------------ | :------------------------ | :------------------ |

| Q8_0 | 421.84 GiB (8.50 BPW) | Q8_0 | 5.202737 ± 0.034827 | +0.3357% | 0.030136 ± 0.000884 |

| Q5_K_M | 295.20 GiB (5.95 BPW) | Q8_0 / Q5_K / Q5_K / Q6_K | 5.218907 ± 0.034969 | +0.6475% | 0.041986 ± 0.001016 |

| Q4_K_M | 246.11 GiB (4.96 BPW) | Q8_0 / Q4_K / Q4_K / Q5_K | 5.282848 ± 0.035408 | +1.8807% | 0.068863 ± 0.000982 |

| IQ4_XS | 189.12 GiB (3.81 BPW) | Q6_K / IQ3_S / IQ3_S / IQ4_XS | 5.657391 ± 0.038606 | +9.1038% | 0.151112 ± 0.001437 |

| IQ3_S | 148.04 GiB (2.98 BPW) | Q6_K / IQ2_S / IQ2_S / IQ3_S | 6.417468 ± 0.044920 | +23.7620% | 0.321256 ± 0.002052 |

| IQ2_S | 134.01 GiB (2.70 BPW) | Q6_K / IQ2_XS / IQ2_XS / IQ3_XXS | 6.636852 ± 0.046125 | +27.9929% | 0.397892 ± 0.002336 |

!kld_graph

!ppl_graph

Run AesSedai/MiniMax-M3-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models