AesSedai/MiMo-V2.6-Flash-GGUF overview
Updates 09/22/26: Added 3.5, 2.5, and 2.0 BPW quants using ed's bpw size PR https://github.com/ggml org/llama.cpp/pull/15550 . The FFNs are the primarily quant…
Runs locally from ~5.7 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| BPW2.0/MiMo-V2.6-Flash-RL-BPW2.0-00001-of-00003.gguf | GGUF | GGUF | 5.7 MB | Download |
| BPW2.0/MiMo-V2.6-Flash-RL-BPW2.0-00002-of-00003.gguf | GGUF | GGUF | 46.38 GB | Download |
| BPW2.0/MiMo-V2.6-Flash-RL-BPW2.0-00003-of-00003.gguf | GGUF | GGUF | 20.61 GB | Download |
| BPW2.5/MiMo-V2.6-Flash-RL-BPW2.5-00001-of-00003.gguf | GGUF | GGUF | 5.7 MB | Download |
| BPW2.5/MiMo-V2.6-Flash-RL-BPW2.5-00002-of-00003.gguf | GGUF | GGUF | 46.32 GB | Download |
| BPW2.5/MiMo-V2.6-Flash-RL-BPW2.5-00003-of-00003.gguf | GGUF | GGUF | 43.82 GB | Download |
| BPW3.5/MiMo-V2.6-Flash-RL-BPW3.5-00001-of-00004.gguf | GGUF | GGUF | 5.7 MB | Download |
| BPW3.5/MiMo-V2.6-Flash-RL-BPW3.5-00002-of-00004.gguf | GGUF | GGUF | 46.56 GB | Download |
| BPW3.5/MiMo-V2.6-Flash-RL-BPW3.5-00003-of-00004.gguf | GGUF | GGUF | 45.73 GB | Download |
| BPW3.5/MiMo-V2.6-Flash-RL-BPW3.5-00004-of-00004.gguf | GGUF | GGUF | 33.90 GB | Download |
| IQ2_S/MiMo-V2.6-Flash-RL-IQ2_S-00001-of-00004.gguf | GGUF | IQ2_S | 5.7 MB | Download |
| IQ2_S/MiMo-V2.6-Flash-RL-IQ2_S-00002-of-00004.gguf | GGUF | IQ2_S | 46.43 GB | Download |
| IQ2_S/MiMo-V2.6-Flash-RL-IQ2_S-00003-of-00004.gguf | GGUF | IQ2_S | 46.53 GB | Download |
| IQ2_S/MiMo-V2.6-Flash-RL-IQ2_S-00004-of-00004.gguf | GGUF | IQ2_S | 13.34 GB | Download |
| MXFP4/MiMo-V2.6-Flash-RL-MXFP4-00001-of-00005.gguf | GGUF | GGUF | 5.7 MB | Download |
| MXFP4/MiMo-V2.6-Flash-RL-MXFP4-00002-of-00005.gguf | GGUF | GGUF | 45.69 GB | Download |
| MXFP4/MiMo-V2.6-Flash-RL-MXFP4-00003-of-00005.gguf | GGUF | GGUF | 45.69 GB | Download |
| MXFP4/MiMo-V2.6-Flash-RL-MXFP4-00004-of-00005.gguf | GGUF | GGUF | 45.69 GB | Download |
| MXFP4/MiMo-V2.6-Flash-RL-MXFP4-00005-of-00005.gguf | GGUF | GGUF | 25.83 GB | Download |
| Q3_K/MiMo-V2.6-Flash-RL-Q3_K-00001-of-00004.gguf | GGUF | Q3_K | 5.7 MB | Download |
| Q3_K/MiMo-V2.6-Flash-RL-Q3_K-00002-of-00004.gguf | GGUF | Q3_K | 45.85 GB | Download |
| Q3_K/MiMo-V2.6-Flash-RL-Q3_K-00003-of-00004.gguf | GGUF | Q3_K | 46.04 GB | Download |
| Q3_K/MiMo-V2.6-Flash-RL-Q3_K-00004-of-00004.gguf | GGUF | Q3_K | 45.86 GB | Download |
| imatrix.gguf | GGUF | GGUF | 473.3 MB | Download |
| mmproj-MiMo-V2.6-Flash-RL-BF16.gguf | GGUF | BF16 | 2.56 GB | Download |
| mmproj-MiMo-V2.6-Flash-RL-F16.gguf | GGUF | F16 | 2.56 GB | Download |
| mmproj-MiMo-V2.6-Flash-RL-F32.gguf | GGUF | F32 | 4.91 GB | Download |
| mmproj-MiMo-V2.6-Flash-RL-Q8_0.gguf | GGUF | Q8_0 | 1.46 GB | Download |
Model Details
Model README
---
base_model:
- XiaomiMiMo/MiMo-V2.6-Flash-RL
---
Updates
- 09/22/26: Added 3.5, 2.5, and 2.0 BPW quants using ed's bpw-size PR. The FFNs are the primarily quantized feature, rest of the model remains in Q8_0 / Q6_K
This repo contains specialized MoE-quants for XiaomiMiMo/MiMo-V2.6-Flash-RL. The idea being that given the huge size of the FFN tensors compared to the rest of the tensors in the model, it should be possible to achieve a better quality while keeping the overall size of the entire model smaller compared to a similar naive quantization. To that end, the quantization type default is kept in high quality and the FFN UP + FFN GATE tensors are quanted down along with the FFN DOWN tensors.
The MXFP4 quant is the "full quality" version, as the model has MXFP4 experts.
| Quant | Size | Mixture | PPL | 1-(Mean PPL(Q)/PPL(base)) | KLD |
| :----- | :-------------------- | :------------------------------------ | :------------------ | :------------------------ | :------------------- |
| MXFP4 | 162.89 GiB (4.52 BPW) | BF16 / MXFP4 | 5.149210 ± 0.030596 | +0.0715% | -0.000000 ± 0.000000 |
| Q3_K | 137.75 GiB (3.82 BPW) | Q8_0 / Q3_K / Q3_K / MXFP4 | 5.177623 ± 0.030900 | +0.6237% | 0.123607 ± 0.000648 |
| BPW3.5 | 126.19 GiB (3.50 BPW) | Q8_0 / varies | 5.232510 ± 0.031063 | +1.6904% | 0.138808 ± 0.000718 |
| IQ2_S | 106.31 GiB (2.95 BPW) | Q6_K / IQ2_S / IQ2_S / Q3_K | 5.404774 ± 0.032158 | +5.0382% | 0.178092 ± 0.000878 |
| BPW2.5 | 90.14 GiB (2.50 BPW) | Q8_0 / varies | 5.739578 ± 0.034587 | +11.5449% | 0.243713 ± 0.001158 |
| BPW2.0 | 66.99 GiB (1.86 BPW) | Q6_K / varies | 7.290888 ± 0.046722 | +41.6936% | 0.477636 ± 0.002111 |
Run AesSedai/MiMo-V2.6-Flash-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models