GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Sciguy429/MiniMax-M3-MSA-GGUF overview

NOTICE: This repo is being kept up for archival purposes for now. These quants were used during the development of the llama.cpp PR which added MSA support for…

ggufbase_model:MiniMaxAI/MiniMax-M3base_model:quantized:MiniMaxAI/MiniMax-M3endpoints_compatibleregion:usimatrixconversational

Runs locally from ~7.9 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
4,299
Likes
2
Pipeline
Author

Repository Files & Downloads

48 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
IQ2_M_IF16/MiniMax-M3-IQ2_M-IF16_MSA-00001-of-00006.ggufGGUFIQ2_M_IF167.9 MBDownload
IQ2_M_IF16/MiniMax-M3-IQ2_M-IF16_MSA-00002-of-00006.ggufGGUFIQ2_M_IF1629.60 GBDownload
IQ2_M_IF16/MiniMax-M3-IQ2_M-IF16_MSA-00003-of-00006.ggufGGUFIQ2_M_IF1629.68 GBDownload
IQ2_M_IF16/MiniMax-M3-IQ2_M-IF16_MSA-00004-of-00006.ggufGGUFIQ2_M_IF1629.72 GBDownload
IQ2_M_IF16/MiniMax-M3-IQ2_M-IF16_MSA-00005-of-00006.ggufGGUFIQ2_M_IF1629.68 GBDownload
IQ2_M_IF16/MiniMax-M3-IQ2_M-IF16_MSA-00006-of-00006.ggufGGUFIQ2_M_IF1610.37 GBDownload
Q6_K_IF32/MiniMax-M3-Q6_K-IF32_MSA-00001-of-00013.ggufGGUFQ6_K_IF327.9 MBDownload
Q6_K_IF32/MiniMax-M3-Q6_K-IF32_MSA-00002-of-00013.ggufGGUFQ6_K_IF3229.17 GBDownload
Q6_K_IF32/MiniMax-M3-Q6_K-IF32_MSA-00003-of-00013.ggufGGUFQ6_K_IF3228.40 GBDownload
Q6_K_IF32/MiniMax-M3-Q6_K-IF32_MSA-00004-of-00013.ggufGGUFQ6_K_IF3228.40 GBDownload
Q6_K_IF32/MiniMax-M3-Q6_K-IF32_MSA-00005-of-00013.ggufGGUFQ6_K_IF3228.40 GBDownload
Q6_K_IF32/MiniMax-M3-Q6_K-IF32_MSA-00006-of-00013.ggufGGUFQ6_K_IF3228.40 GBDownload
Q6_K_IF32/MiniMax-M3-Q6_K-IF32_MSA-00007-of-00013.ggufGGUFQ6_K_IF3228.40 GBDownload
Q6_K_IF32/MiniMax-M3-Q6_K-IF32_MSA-00008-of-00013.ggufGGUFQ6_K_IF3228.40 GBDownload
Q6_K_IF32/MiniMax-M3-Q6_K-IF32_MSA-00009-of-00013.ggufGGUFQ6_K_IF3228.40 GBDownload
Q6_K_IF32/MiniMax-M3-Q6_K-IF32_MSA-00010-of-00013.ggufGGUFQ6_K_IF3228.40 GBDownload
Q6_K_IF32/MiniMax-M3-Q6_K-IF32_MSA-00011-of-00013.ggufGGUFQ6_K_IF3228.40 GBDownload
Q6_K_IF32/MiniMax-M3-Q6_K-IF32_MSA-00012-of-00013.ggufGGUFQ6_K_IF3228.40 GBDownload
Q6_K_IF32/MiniMax-M3-Q6_K-IF32_MSA-00013-of-00013.ggufGGUFQ6_K_IF3213.23 GBDownload
Q8_0_IF32/MiniMax-M3-Q8_0-IF32_MSA-00001-of-00016.ggufGGUFQ8_0_IF327.9 MBDownload
Q8_0_IF32/MiniMax-M3-Q8_0-IF32_MSA-00002-of-00016.ggufGGUFQ8_0_IF3227.99 GBDownload
Q8_0_IF32/MiniMax-M3-Q8_0-IF32_MSA-00003-of-00016.ggufGGUFQ8_0_IF3229.41 GBDownload
Q8_0_IF32/MiniMax-M3-Q8_0-IF32_MSA-00004-of-00016.ggufGGUFQ8_0_IF3229.41 GBDownload
Q8_0_IF32/MiniMax-M3-Q8_0-IF32_MSA-00005-of-00016.ggufGGUFQ8_0_IF3229.41 GBDownload
Q8_0_IF32/MiniMax-M3-Q8_0-IF32_MSA-00006-of-00016.ggufGGUFQ8_0_IF3229.41 GBDownload
Q8_0_IF32/MiniMax-M3-Q8_0-IF32_MSA-00007-of-00016.ggufGGUFQ8_0_IF3229.41 GBDownload
Q8_0_IF32/MiniMax-M3-Q8_0-IF32_MSA-00008-of-00016.ggufGGUFQ8_0_IF3229.41 GBDownload
Q8_0_IF32/MiniMax-M3-Q8_0-IF32_MSA-00009-of-00016.ggufGGUFQ8_0_IF3229.41 GBDownload
Q8_0_IF32/MiniMax-M3-Q8_0-IF32_MSA-00010-of-00016.ggufGGUFQ8_0_IF3229.41 GBDownload
Q8_0_IF32/MiniMax-M3-Q8_0-IF32_MSA-00011-of-00016.ggufGGUFQ8_0_IF3229.41 GBDownload
Q8_0_IF32/MiniMax-M3-Q8_0-IF32_MSA-00012-of-00016.ggufGGUFQ8_0_IF3229.41 GBDownload
Q8_0_IF32/MiniMax-M3-Q8_0-IF32_MSA-00013-of-00016.ggufGGUFQ8_0_IF3229.41 GBDownload
Q8_0_IF32/MiniMax-M3-Q8_0-IF32_MSA-00014-of-00016.ggufGGUFQ8_0_IF3229.41 GBDownload
Q8_0_IF32/MiniMax-M3-Q8_0-IF32_MSA-00015-of-00016.ggufGGUFQ8_0_IF3229.41 GBDownload
Q8_0_IF32/MiniMax-M3-Q8_0-IF32_MSA-00016-of-00016.ggufGGUFQ8_0_IF3212.19 GBDownload
UD-IQ2_M_IF16/MiniMax-M3-UD-IQ2_M-IF16_MSA-00001-of-00006.ggufGGUFIQ2_M_IF167.9 MBDownload
UD-IQ2_M_IF16/MiniMax-M3-UD-IQ2_M-IF16_MSA-00002-of-00006.ggufGGUFIQ2_M_IF1629.26 GBDownload
UD-IQ2_M_IF16/MiniMax-M3-UD-IQ2_M-IF16_MSA-00003-of-00006.ggufGGUFIQ2_M_IF1629.28 GBDownload
UD-IQ2_M_IF16/MiniMax-M3-UD-IQ2_M-IF16_MSA-00004-of-00006.ggufGGUFIQ2_M_IF1629.00 GBDownload
UD-IQ2_M_IF16/MiniMax-M3-UD-IQ2_M-IF16_MSA-00005-of-00006.ggufGGUFIQ2_M_IF1629.54 GBDownload
UD-IQ2_M_IF16/MiniMax-M3-UD-IQ2_M-IF16_MSA-00006-of-00006.ggufGGUFIQ2_M_IF167.97 GBDownload
UD-IQ2_M_IF32/MiniMax-M3-UD-IQ2_M_IF32-00001-of-00006.ggufGGUFIQ2_M_IF327.9 MBDownload
UD-IQ2_M_IF32/MiniMax-M3-UD-IQ2_M_IF32-00002-of-00006.ggufGGUFIQ2_M_IF3229.35 GBDownload
UD-IQ2_M_IF32/MiniMax-M3-UD-IQ2_M_IF32-00003-of-00006.ggufGGUFIQ2_M_IF3229.38 GBDownload
UD-IQ2_M_IF32/MiniMax-M3-UD-IQ2_M_IF32-00004-of-00006.ggufGGUFIQ2_M_IF3229.10 GBDownload
UD-IQ2_M_IF32/MiniMax-M3-UD-IQ2_M_IF32-00005-of-00006.ggufGGUFIQ2_M_IF3229.64 GBDownload
UD-IQ2_M_IF32/MiniMax-M3-UD-IQ2_M_IF32-00006-of-00006.ggufGGUFIQ2_M_IF328.00 GBDownload
imatrix.ggufGGUFGGUF441.4 MBDownload

Model Details

Model IDSciguy429/MiniMax-M3-MSA-GGUF
AuthorSciguy429
Pipeline
License
Base modelMiniMaxAI/MiniMax-M3
Last modified2026-07-28T20:45:18.000Z

Model README

---

base_model:

  • MiniMaxAI/MiniMax-M3

---

NOTICE:

This repo is being kept up for archival purposes for now. These quants were used during the development of the llama.cpp PR which added MSA support for MiniMax-M3. They are now outdated and it is recommended to use AesSedai or Unsloth's updated quant mixes. I will likely not be updating these further (even though they should be working still) as MiniMax-M3 dose not survive the 2-bit quant process very well. I would not recommend the use of these low bit quants for real world tasks.

---

OLD DESCRIPTION:

This is an experimental 2-bit GGUF quant of MiniMax-M3 which supports MSA via this llama.cpp PR:

https://github.com/ggml-org/llama.cpp/pull/24908

I will keep these quants (and any more I create) up until the PR merges and the bigger players update there repos to match, then I will likely replace this repo with IK quants once ik_llama catches up.

These are up to date with the PR as of 7-10-2026.

Models:

IF16/IF32?

I am still running tests on this at the moment, but for now. The indexer tensors within the official MiniMax-M3 release are stored as both FP32 and BF16 tensors. Specifically the projection tensors (index_{q,k}_proj) are what use BF16. There is some conflicting information in the PR about what these values should be quantized too. The main description recommends F32, but an example llama-quantize command further down in the chain just uses F16.

~~(Once I get some proper perp testing done I will update this section with any differences I noted. I will leave both model's up regardless!)~~

I have done some testing via runpod on the IF16/32 issue. The repo now has a flat Q8 and Q6_K quant in it if anyone else wants to avoid generating them. I also generated a 32K context Q8 logit dump via the following command (which is also in the repo):

./llama-perplexity --flash-attn on -lv 4 --file /workspace/Models/wiki.test.raw --save-all-logits /workspace/Models/MiniMax-M3-Q8_0-IF32-32768ctx-wiki.test.raw.bin --batch-size 2048 --ubatch-size 2048 -c 32768 --fit on --no-mmap --model /workspace/Models/MiniMax-M3-Q8_0.gguf

(Sadly, runpod ate the command output by randomly crashing the SSH session, so I don't have the final PPL for the Q8 file as I lost the entire command output...)

Regardless, I went on to run a perplexity test on my UD-IQ2_M-IF32/16 quants and did find a slight difference between the two:

(Full outputs in the repo.)

IF16:
Mean    KLD:   0.457160 ±   0.002882

IF32:
Mean    KLD:   0.455649 ±   0.002860

The difference is within the error margin, but it is there. So while I can't say with certainty the exact importance of the F32 conversion, I think it is safe to say that it dose have an effect.

UD-IQ2_M_IF16/IF32:

These are clones of Unsloth's UD-IQ2_M mix, with FP16 and FP32 indexer tensors respectively. They behave fairly well compared to the original quant from Unsloth and are likely the highest 'quality' 2-bit quant you can get right now until they update their mix once PR #24098 lands or ik_llama implements it's own support and allows for better 2-bit SOTA quant types.

Mean KLD: 0.488235 ± 0.002577

IQ2_M_IF16:

This is a flat automatic IQ2_M quant directly from llama-quantize without any tensor overrides. I just made this one to compare against the UD quant.

NOT RECOMENDED FOR ACTUALL USE!

Mean KLD: 0.490047 ± 0.002621

Run Sciguy429/MiniMax-M3-MSA-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models