GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

scragnog/MiniMax-Music3-GGUF overview

MiniMax Music3 — GGUF GGUF conversion of MiniMaxAI/MiniMax Music3 https://huggingface.co/MiniMaxAI/MiniMax Music3 for HOT Step CPP https://github.com/scragnog/…

ggufmusic-generationtext-to-musichot-step-cppbase_model:MiniMaxAI/MiniMax-Music3base_model:quantized:MiniMaxAI/MiniMax-Music3license:otherregion:us

Runs locally from ~5.1 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
68,608
Likes
5
Pipeline
Author

Repository Files & Downloads

63 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
mm3-cond-f16.ggufGGUFF1648.0 MBDownload
mm3-depth-MXFP4.ggufGGUFGGUF365.9 MBDownload
mm3-depth-NVFP4.ggufGGUFGGUF382.9 MBDownload
mm3-depth-Q2_K.ggufGGUFQ2_K275.5 MBDownload
mm3-depth-Q3_K_L.ggufGGUFQ3_K_L351.0 MBDownload
mm3-depth-Q3_K_M.ggufGGUFQ3_K_M313.8 MBDownload
mm3-depth-Q3_K_S.ggufGGUFQ3_K_S293.2 MBDownload
mm3-depth-Q4_K_M.ggufGGUFQ4_K_M386.1 MBDownload
mm3-depth-Q4_K_S.ggufGGUFQ4_K_S385.5 MBDownload
mm3-depth-Q5_K_M.ggufGGUFQ5_K_M444.1 MBDownload
mm3-depth-Q5_K_S.ggufGGUFQ5_K_S433.5 MBDownload
mm3-depth-Q6_K.ggufGGUFQ6_K505.7 MBDownload
mm3-depth-f16.ggufGGUFF161.20 GBDownload
mm3-depth-q8_0.ggufGGUFQ8_0654.9 MBDownload
mm3-dit-MXFP4.ggufGGUFGGUF1.22 GBDownload
mm3-dit-NVFP4.ggufGGUFGGUF1.29 GBDownload
mm3-dit-Q2_K.ggufGGUFQ2_K792.4 MBDownload
mm3-dit-Q3_K_L.ggufGGUFQ3_K_L1.17 GBDownload
mm3-dit-Q3_K_M.ggufGGUFQ3_K_M1.06 GBDownload
mm3-dit-Q3_K_S.ggufGGUFQ3_K_S1011.4 MBDownload
mm3-dit-Q4_K_M.ggufGGUFQ4_K_M1.36 GBDownload
mm3-dit-Q4_K_S.ggufGGUFQ4_K_S1.29 GBDownload
mm3-dit-Q5_K_M.ggufGGUFQ5_K_M1.61 GBDownload
mm3-dit-Q5_K_S.ggufGGUFQ5_K_S1.57 GBDownload
mm3-dit-Q6_K.ggufGGUFQ6_K1.87 GBDownload
mm3-dit-f16.ggufGGUFF164.53 GBDownload
mm3-dit-q8_0.ggufGGUFQ8_02.41 GBDownload
mm3-enc-f16.ggufGGUFF1685.3 MBDownload
mm3-lm-IQ2_XS-imat.ggufGGUFIQ2_XS2.76 GBDownload
mm3-lm-IQ2_XXS-imat.ggufGGUFIQ2_XXS2.56 GBDownload
mm3-lm-IQ3_XXS-imat.ggufGGUFIQ3_XXS3.56 GBDownload
mm3-lm-IQ4_XS-imat.ggufGGUFIQ4_XS4.70 GBDownload
mm3-lm-MXFP4.ggufGGUFGGUF5.07 GBDownload
mm3-lm-NVFP4.ggufGGUFGGUF5.27 GBDownload
mm3-lm-Q2_K-imat.ggufGGUFQ2_K3.43 GBDownload
mm3-lm-Q3_K_L-imat.ggufGGUFQ3_K_L4.66 GBDownload
mm3-lm-Q3_K_M-imat.ggufGGUFQ3_K_M4.27 GBDownload
mm3-lm-Q3_K_S-imat.ggufGGUFQ3_K_S4.04 GBDownload
mm3-lm-Q4_K_M-imat.ggufGGUFQ4_K_M5.13 GBDownload
mm3-lm-Q4_K_S-imat.ggufGGUFQ4_K_S4.92 GBDownload
mm3-lm-Q5_K_M-imat.ggufGGUFQ5_K_M5.83 GBDownload
mm3-lm-Q5_K_S-imat.ggufGGUFQ5_K_S5.71 GBDownload
mm3-lm-Q6_K-imat.ggufGGUFQ6_K6.57 GBDownload
mm3-lm-bf16.ggufGGUFBF1616.00 GBDownload
mm3-lm-f16.ggufGGUFF1616.00 GBDownload
mm3-lm-q8_0.ggufGGUFQ8_08.50 GBDownload
mm3-lm.imatrix.ggufGGUFGGUF5.1 MBDownload
mm3-rec7-f16.ggufGGUFF16599.2 MBDownload
mm3-rvq-53kpooled-f32.ggufGGUFF32644.7 MBDownload
mm3-synth-MXFP4.ggufGGUFGGUF1.82 GBDownload
mm3-synth-NVFP4.ggufGGUFGGUF1.91 GBDownload
mm3-synth-Q2_K.ggufGGUFQ2_K1.29 GBDownload
mm3-synth-Q3_K_L.ggufGGUFQ3_K_L1.76 GBDownload
mm3-synth-Q3_K_M.ggufGGUFQ3_K_M1.62 GBDownload
mm3-synth-Q3_K_S.ggufGGUFQ3_K_S1.52 GBDownload
mm3-synth-Q4_K_M.ggufGGUFQ4_K_M1.98 GBDownload
mm3-synth-Q4_K_S.ggufGGUFQ4_K_S1.92 GBDownload
mm3-synth-Q5_K_M.ggufGGUFQ5_K_M2.29 GBDownload
mm3-synth-Q5_K_S.ggufGGUFQ5_K_S2.24 GBDownload
mm3-synth-Q6_K.ggufGGUFQ6_K2.61 GBDownload
mm3-synth-f16.ggufGGUFF165.98 GBDownload
mm3-synth-q8_0.ggufGGUFQ8_03.30 GBDownload
mm3-voc-f16.ggufGGUFF16206.6 MBDownload

Model Details

Model IDscragnog/MiniMax-Music3-GGUF
Authorscragnog
Pipeline
Licenseother
Base modelMiniMaxAI/MiniMax-Music3
Last modified2026-09-09T20:00:37.000Z

Model README

---

license: other

license_name: minimax-music3-community-license

license_link: https://huggingface.co/MiniMaxAI/MiniMax-Music3/blob/main/LICENSE

base_model: MiniMaxAI/MiniMax-Music3

tags:

  • gguf
  • music-generation
  • text-to-music
  • hot-step-cpp

---

MiniMax-Music3 — GGUF

GGUF conversion of MiniMaxAI/MiniMax-Music3

for HOT-Step CPP, a fully local desktop app for

AI music generation with a native C++/GGML engine. As far as we know this is the first GGUF

conversion of this model.

These files are downloaded automatically by HOT-Step's Model Manager when the MiniMax-Music3

backend is selected. They are not usable with llama.cpp alone: mm3-lm-*.gguf is

structurally a Qwen3 GGUF, but music generation requires the full five-module pipeline

(LM → RVQ depth decoder → condition encoder → flow-matching DiT → vocoder) implemented in

HOT-Step's engine.

Split format (one GGUF per component)

Since 2026-08-14 the repo carries one file per pipeline component, so each can be picked

at its own quantisation and swapped without re-downloading the others:

| File family | Component | Quants |

|---|---|---|

| mm3-lm-<quant>.gguf | Global LM (8.59B, Qwen3 arch, 200k vocab incl. 16,384 semantic audio codes) + full tokenizer | f16 · q8_0 · Q6_K · Q5_K_M/S · Q4_K_M/S · NVFP4 · MXFP4 · Q3_K_L/M/S · Q2_K |

| mm3-dit-<quant>.gguf | Flow-matching DiT (2.4B) | same ladder |

| mm3-depth-<quant>.gguf | RVQ depth decoder (0.6B) | same ladder (Q8_0 is the validated quality floor) |

| mm3-cond-f16.gguf | Condition encoder (25M) | f16 only — never quantised |

| mm3-voc-f16.gguf | Vocoder (54M) | f16 only — never quantised |

| LICENSE | MiniMax-Music3 Community License (governs the weights) | — |

Suggested combos: quality = everything f16 (~24 GB VRAM); recommended = LM/DiT/depth

q8_0 + f16 frontends (~13 GB); balanced = LM q8_0 + DiT Q4_K_M + depth q8_0 (~12 GB

download, the split's headline mix); fast on RTX 50-series = LM/DiT NVFP4 + depth q8_0.

Audio-code LMs degrade audibly below Q5 — the sub-Q5 LM files exist for experiments, not

listening.

The legacy two-file layout (mm3-synth-<quant>.gguf bundling depth+cond+dit+voc) remains

available for older HOT-Step versions; current versions load either, preferring split files.

Output: 44.1 kHz stereo, up to 5 minutes.

Training files (added 2026-08-31)

Two more files, needed only by HOT-Step's Training Studio when training an MM3 LoRA. Neither

is loaded during generation, and neither is in the generation packs. Download them together

as the Model Manager's MiniMax-Music3 Training pack.

| File | Component | Why it is needed |

|---|---|---|

| mm3-rvq-53kpooled-f32.gguf | Audio → RVQ codes encoder (169M) | The codes stage turns your dataset's audio into the code streams the LM trains on |

| mm3-enc-f16.gguf | DAV audio encoder (44.1 kHz stereo → 128-channel flow latents) | Input stage of the same codes job |

| mm3-lm-bf16.gguf | The LM in its source BF16 precision (17.2 GB) | OPTIONAL, training only (added 2026-09-03). Picked as the training base it lets the trainer run the projection GEMMs on BF16 tensor cores (--weights bf16) instead of the F32 fallback every other base uses: 1.4x faster per step than q8_0 on a 5090 at matched crop, for ~9 GB more VRAM, so it wants a 40 GB+ card at the default crop. Not better than f16 for generation; render on q8_0 as always |

| mm3-rec7-f16.gguf | rec7 state encoder (audio → LM frame hiddens, 170M) | OPTIONAL — only the codes stage's "Cover-launder" option (dense-mix training fix, 2026-08-31). Carries the LM's two semantic table slices so laundering never runs the 8B. PurpleOrc's m3-rec7-encoder (MIT), converted with convert-rvq-encoder.py --head --m3 |

Training also wants mm3-depth-f16.gguf specifically. The quantised depth files that ship

with the generation packs are for generation.

MiniMax never released the official audio tokeniser, so the RVQ encoder here is a community

reimplementation: PurpleOrc's open-rvq encoder,

SimpleTuner's v4 architecture trained from scratch on a 53k-track multilingual corpus,

converted to GGUF and mirrored here so the Model Manager has one place to fetch from. Its

reported holdout scores and training code are on PurpleOrc's repo, which is the place to

read before drawing conclusions about it.

Codes are encoder-specific. An adapter trained on this encoder's codes has to keep using this

encoder at inference; swapping encoders means re-exporting every code cache.

Conversion provenance

Converted with HOT-Step's engine/tools/convert-mm3.py

from the bf16/fp16 safetensors published by MiniMax (via the Comfy-Org repackage), then split

per component with engine/tools/split-mm3.py

(byte-exact tensor passthrough — a split file's tensors are bit-identical to the bundle's).

Vocoder weight-norm folded at conversion; vocoder and DiT Fourier/RoPE bases pinned F32; all

911 tensors shape-validated. The HOT-Step implementation is parity-validated against the

official diffusers reference (per-module correlation ≥ 0.9999 vs fp32; full-pipeline replay

0.9988). The split-model approach follows

ServeurpersoCom/minimaxmusic.cpp,

whose author kindly sanctioned reuse of his design.

Credits

All credit for the model belongs to MiniMax.

MiniMax-Music3 is their work — the architecture, the training, and the release. This

repository contains nothing but a format conversion of their weights, and the Structured

Caption format that drives the model is their design.

Thanks also to:

acestep.cpp, the C++/GGML engine

HOT-Step extends, and for minimaxmusic.cpp,

whose per-component split design this repo's layout follows — kindly sanctioned for reuse.

HOT-Step is built around.

open-rvq encoder

mirrored here, and for offering it to the community in the first place; and to

bghira / SimpleTuner for the v4 encoder

architecture it was trained from.

MOSS-Music-8B-Instruct,

which writes the Structured Captions that prompt this model — from listening to a

reference track rather than from text. GGUF conversion:

scragnog/MOSS-Music-8B-Instruct-GGUF.

License

The model weights are subject to the MiniMax-Music3 Community License (included here as

LICENSE, per its notice-preservation requirement). Notable terms: prominent display of

"MiniMax-Music3" in commercial products, separate authorization above US$20M annual revenue,

acceptable-use policy, and clear disclosure of AI generation for publicly distributed outputs.

The conversion adds no additional restrictions.

Run scragnog/MiniMax-Music3-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models