scragnog/MiniMax-Music3-GGUF overview
MiniMax Music3 — GGUF GGUF conversion of MiniMaxAI/MiniMax Music3 https://huggingface.co/MiniMaxAI/MiniMax Music3 for HOT Step CPP https://github.com/scragnog/…
Runs locally from ~5.1 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| mm3-cond-f16.gguf | GGUF | F16 | 48.0 MB | Download |
| mm3-depth-MXFP4.gguf | GGUF | GGUF | 365.9 MB | Download |
| mm3-depth-NVFP4.gguf | GGUF | GGUF | 382.9 MB | Download |
| mm3-depth-Q2_K.gguf | GGUF | Q2_K | 275.5 MB | Download |
| mm3-depth-Q3_K_L.gguf | GGUF | Q3_K_L | 351.0 MB | Download |
| mm3-depth-Q3_K_M.gguf | GGUF | Q3_K_M | 313.8 MB | Download |
| mm3-depth-Q3_K_S.gguf | GGUF | Q3_K_S | 293.2 MB | Download |
| mm3-depth-Q4_K_M.gguf | GGUF | Q4_K_M | 386.1 MB | Download |
| mm3-depth-Q4_K_S.gguf | GGUF | Q4_K_S | 385.5 MB | Download |
| mm3-depth-Q5_K_M.gguf | GGUF | Q5_K_M | 444.1 MB | Download |
| mm3-depth-Q5_K_S.gguf | GGUF | Q5_K_S | 433.5 MB | Download |
| mm3-depth-Q6_K.gguf | GGUF | Q6_K | 505.7 MB | Download |
| mm3-depth-f16.gguf | GGUF | F16 | 1.20 GB | Download |
| mm3-depth-q8_0.gguf | GGUF | Q8_0 | 654.9 MB | Download |
| mm3-dit-MXFP4.gguf | GGUF | GGUF | 1.22 GB | Download |
| mm3-dit-NVFP4.gguf | GGUF | GGUF | 1.29 GB | Download |
| mm3-dit-Q2_K.gguf | GGUF | Q2_K | 792.4 MB | Download |
| mm3-dit-Q3_K_L.gguf | GGUF | Q3_K_L | 1.17 GB | Download |
| mm3-dit-Q3_K_M.gguf | GGUF | Q3_K_M | 1.06 GB | Download |
| mm3-dit-Q3_K_S.gguf | GGUF | Q3_K_S | 1011.4 MB | Download |
| mm3-dit-Q4_K_M.gguf | GGUF | Q4_K_M | 1.36 GB | Download |
| mm3-dit-Q4_K_S.gguf | GGUF | Q4_K_S | 1.29 GB | Download |
| mm3-dit-Q5_K_M.gguf | GGUF | Q5_K_M | 1.61 GB | Download |
| mm3-dit-Q5_K_S.gguf | GGUF | Q5_K_S | 1.57 GB | Download |
| mm3-dit-Q6_K.gguf | GGUF | Q6_K | 1.87 GB | Download |
| mm3-dit-f16.gguf | GGUF | F16 | 4.53 GB | Download |
| mm3-dit-q8_0.gguf | GGUF | Q8_0 | 2.41 GB | Download |
| mm3-enc-f16.gguf | GGUF | F16 | 85.3 MB | Download |
| mm3-lm-IQ2_XS-imat.gguf | GGUF | IQ2_XS | 2.76 GB | Download |
| mm3-lm-IQ2_XXS-imat.gguf | GGUF | IQ2_XXS | 2.56 GB | Download |
| mm3-lm-IQ3_XXS-imat.gguf | GGUF | IQ3_XXS | 3.56 GB | Download |
| mm3-lm-IQ4_XS-imat.gguf | GGUF | IQ4_XS | 4.70 GB | Download |
| mm3-lm-MXFP4.gguf | GGUF | GGUF | 5.07 GB | Download |
| mm3-lm-NVFP4.gguf | GGUF | GGUF | 5.27 GB | Download |
| mm3-lm-Q2_K-imat.gguf | GGUF | Q2_K | 3.43 GB | Download |
| mm3-lm-Q3_K_L-imat.gguf | GGUF | Q3_K_L | 4.66 GB | Download |
| mm3-lm-Q3_K_M-imat.gguf | GGUF | Q3_K_M | 4.27 GB | Download |
| mm3-lm-Q3_K_S-imat.gguf | GGUF | Q3_K_S | 4.04 GB | Download |
| mm3-lm-Q4_K_M-imat.gguf | GGUF | Q4_K_M | 5.13 GB | Download |
| mm3-lm-Q4_K_S-imat.gguf | GGUF | Q4_K_S | 4.92 GB | Download |
| mm3-lm-Q5_K_M-imat.gguf | GGUF | Q5_K_M | 5.83 GB | Download |
| mm3-lm-Q5_K_S-imat.gguf | GGUF | Q5_K_S | 5.71 GB | Download |
| mm3-lm-Q6_K-imat.gguf | GGUF | Q6_K | 6.57 GB | Download |
| mm3-lm-bf16.gguf | GGUF | BF16 | 16.00 GB | Download |
| mm3-lm-f16.gguf | GGUF | F16 | 16.00 GB | Download |
| mm3-lm-q8_0.gguf | GGUF | Q8_0 | 8.50 GB | Download |
| mm3-lm.imatrix.gguf | GGUF | GGUF | 5.1 MB | Download |
| mm3-rec7-f16.gguf | GGUF | F16 | 599.2 MB | Download |
| mm3-rvq-53kpooled-f32.gguf | GGUF | F32 | 644.7 MB | Download |
| mm3-synth-MXFP4.gguf | GGUF | GGUF | 1.82 GB | Download |
| mm3-synth-NVFP4.gguf | GGUF | GGUF | 1.91 GB | Download |
| mm3-synth-Q2_K.gguf | GGUF | Q2_K | 1.29 GB | Download |
| mm3-synth-Q3_K_L.gguf | GGUF | Q3_K_L | 1.76 GB | Download |
| mm3-synth-Q3_K_M.gguf | GGUF | Q3_K_M | 1.62 GB | Download |
| mm3-synth-Q3_K_S.gguf | GGUF | Q3_K_S | 1.52 GB | Download |
| mm3-synth-Q4_K_M.gguf | GGUF | Q4_K_M | 1.98 GB | Download |
| mm3-synth-Q4_K_S.gguf | GGUF | Q4_K_S | 1.92 GB | Download |
| mm3-synth-Q5_K_M.gguf | GGUF | Q5_K_M | 2.29 GB | Download |
| mm3-synth-Q5_K_S.gguf | GGUF | Q5_K_S | 2.24 GB | Download |
| mm3-synth-Q6_K.gguf | GGUF | Q6_K | 2.61 GB | Download |
| mm3-synth-f16.gguf | GGUF | F16 | 5.98 GB | Download |
| mm3-synth-q8_0.gguf | GGUF | Q8_0 | 3.30 GB | Download |
| mm3-voc-f16.gguf | GGUF | F16 | 206.6 MB | Download |
Model Details
Model README
---
license: other
license_name: minimax-music3-community-license
license_link: https://huggingface.co/MiniMaxAI/MiniMax-Music3/blob/main/LICENSE
base_model: MiniMaxAI/MiniMax-Music3
tags:
- gguf
- music-generation
- text-to-music
- hot-step-cpp
---
MiniMax-Music3 — GGUF
GGUF conversion of MiniMaxAI/MiniMax-Music3
for HOT-Step CPP, a fully local desktop app for
AI music generation with a native C++/GGML engine. As far as we know this is the first GGUF
conversion of this model.
These files are downloaded automatically by HOT-Step's Model Manager when the MiniMax-Music3
backend is selected. They are not usable with llama.cpp alone: mm3-lm-*.gguf is
structurally a Qwen3 GGUF, but music generation requires the full five-module pipeline
(LM → RVQ depth decoder → condition encoder → flow-matching DiT → vocoder) implemented in
HOT-Step's engine.
Split format (one GGUF per component)
Since 2026-08-14 the repo carries one file per pipeline component, so each can be picked
at its own quantisation and swapped without re-downloading the others:
| File family | Component | Quants |
|---|---|---|
| mm3-lm-<quant>.gguf | Global LM (8.59B, Qwen3 arch, 200k vocab incl. 16,384 semantic audio codes) + full tokenizer | f16 · q8_0 · Q6_K · Q5_K_M/S · Q4_K_M/S · NVFP4 · MXFP4 · Q3_K_L/M/S · Q2_K |
| mm3-dit-<quant>.gguf | Flow-matching DiT (2.4B) | same ladder |
| mm3-depth-<quant>.gguf | RVQ depth decoder (0.6B) | same ladder (Q8_0 is the validated quality floor) |
| mm3-cond-f16.gguf | Condition encoder (25M) | f16 only — never quantised |
| mm3-voc-f16.gguf | Vocoder (54M) | f16 only — never quantised |
| LICENSE | MiniMax-Music3 Community License (governs the weights) | — |
Suggested combos: quality = everything f16 (~24 GB VRAM); recommended = LM/DiT/depth
q8_0 + f16 frontends (~13 GB); balanced = LM q8_0 + DiT Q4_K_M + depth q8_0 (~12 GB
download, the split's headline mix); fast on RTX 50-series = LM/DiT NVFP4 + depth q8_0.
Audio-code LMs degrade audibly below Q5 — the sub-Q5 LM files exist for experiments, not
listening.
The legacy two-file layout (mm3-synth-<quant>.gguf bundling depth+cond+dit+voc) remains
available for older HOT-Step versions; current versions load either, preferring split files.
Output: 44.1 kHz stereo, up to 5 minutes.
Training files (added 2026-08-31)
Two more files, needed only by HOT-Step's Training Studio when training an MM3 LoRA. Neither
is loaded during generation, and neither is in the generation packs. Download them together
as the Model Manager's MiniMax-Music3 Training pack.
| File | Component | Why it is needed |
|---|---|---|
| mm3-rvq-53kpooled-f32.gguf | Audio → RVQ codes encoder (169M) | The codes stage turns your dataset's audio into the code streams the LM trains on |
| mm3-enc-f16.gguf | DAV audio encoder (44.1 kHz stereo → 128-channel flow latents) | Input stage of the same codes job |
| mm3-lm-bf16.gguf | The LM in its source BF16 precision (17.2 GB) | OPTIONAL, training only (added 2026-09-03). Picked as the training base it lets the trainer run the projection GEMMs on BF16 tensor cores (--weights bf16) instead of the F32 fallback every other base uses: 1.4x faster per step than q8_0 on a 5090 at matched crop, for ~9 GB more VRAM, so it wants a 40 GB+ card at the default crop. Not better than f16 for generation; render on q8_0 as always |
| mm3-rec7-f16.gguf | rec7 state encoder (audio → LM frame hiddens, 170M) | OPTIONAL — only the codes stage's "Cover-launder" option (dense-mix training fix, 2026-08-31). Carries the LM's two semantic table slices so laundering never runs the 8B. PurpleOrc's m3-rec7-encoder (MIT), converted with convert-rvq-encoder.py --head --m3 |
Training also wants mm3-depth-f16.gguf specifically. The quantised depth files that ship
with the generation packs are for generation.
MiniMax never released the official audio tokeniser, so the RVQ encoder here is a community
reimplementation: PurpleOrc's open-rvq encoder,
SimpleTuner's v4 architecture trained from scratch on a 53k-track multilingual corpus,
converted to GGUF and mirrored here so the Model Manager has one place to fetch from. Its
reported holdout scores and training code are on PurpleOrc's repo, which is the place to
read before drawing conclusions about it.
Codes are encoder-specific. An adapter trained on this encoder's codes has to keep using this
encoder at inference; swapping encoders means re-exporting every code cache.
Conversion provenance
Converted with HOT-Step's engine/tools/convert-mm3.py
from the bf16/fp16 safetensors published by MiniMax (via the Comfy-Org repackage), then split
per component with engine/tools/split-mm3.py
(byte-exact tensor passthrough — a split file's tensors are bit-identical to the bundle's).
Vocoder weight-norm folded at conversion; vocoder and DiT Fourier/RoPE bases pinned F32; all
911 tensors shape-validated. The HOT-Step implementation is parity-validated against the
official diffusers reference (per-module correlation ≥ 0.9999 vs fp32; full-pipeline replay
0.9988). The split-model approach follows
ServeurpersoCom/minimaxmusic.cpp,
whose author kindly sanctioned reuse of his design.
Credits
All credit for the model belongs to MiniMax.
MiniMax-Music3 is their work — the architecture, the training, and the release. This
repository contains nothing but a format conversion of their weights, and the Structured
Caption format that drives the model is their design.
Thanks also to:
acestep.cpp, the C++/GGML engine
HOT-Step extends, and for minimaxmusic.cpp,
whose per-component split design this repo's layout follows — kindly sanctioned for reuse.
- ACE Studio / StepFun for ACE-Step, the model
HOT-Step is built around.
- PurpleOrc for the
mirrored here, and for offering it to the community in the first place; and to
bghira / SimpleTuner for the v4 encoder
architecture it was trained from.
which writes the Structured Captions that prompt this model — from listening to a
reference track rather than from text. GGUF conversion:
scragnog/MOSS-Music-8B-Instruct-GGUF.
License
The model weights are subject to the MiniMax-Music3 Community License (included here as
LICENSE, per its notice-preservation requirement). Notable terms: prominent display of
"MiniMax-Music3" in commercial products, separate authorization above US$20M annual revenue,
acceptable-use policy, and clear disclosure of AI generation for publicly distributed outputs.
The conversion adds no additional restrictions.
Run scragnog/MiniMax-Music3-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models