GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

AesSedai/MiMo-V2.6-Flash-GGUF overview

Updates 09/22/26: Added 3.5, 2.5, and 2.0 BPW quants using ed's bpw size PR https://github.com/ggml org/llama.cpp/pull/15550 . The FFNs are the primarily quant…

ggufbase_model:XiaomiMiMo/MiMo-V2.6-Flash-RLbase_model:quantized:XiaomiMiMo/MiMo-V2.6-Flash-RLendpoints_compatibleregion:usimatrixconversational

Runs locally from ~5.7 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
2
Likes
8
Pipeline
—
Author

Repository Files & Downloads

28 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
BPW2.0/MiMo-V2.6-Flash-RL-BPW2.0-00001-of-00003.ggufGGUFGGUF5.7 MBDownload
BPW2.0/MiMo-V2.6-Flash-RL-BPW2.0-00002-of-00003.ggufGGUFGGUF46.38 GBDownload
BPW2.0/MiMo-V2.6-Flash-RL-BPW2.0-00003-of-00003.ggufGGUFGGUF20.61 GBDownload
BPW2.5/MiMo-V2.6-Flash-RL-BPW2.5-00001-of-00003.ggufGGUFGGUF5.7 MBDownload
BPW2.5/MiMo-V2.6-Flash-RL-BPW2.5-00002-of-00003.ggufGGUFGGUF46.32 GBDownload
BPW2.5/MiMo-V2.6-Flash-RL-BPW2.5-00003-of-00003.ggufGGUFGGUF43.82 GBDownload
BPW3.5/MiMo-V2.6-Flash-RL-BPW3.5-00001-of-00004.ggufGGUFGGUF5.7 MBDownload
BPW3.5/MiMo-V2.6-Flash-RL-BPW3.5-00002-of-00004.ggufGGUFGGUF46.56 GBDownload
BPW3.5/MiMo-V2.6-Flash-RL-BPW3.5-00003-of-00004.ggufGGUFGGUF45.73 GBDownload
BPW3.5/MiMo-V2.6-Flash-RL-BPW3.5-00004-of-00004.ggufGGUFGGUF33.90 GBDownload
IQ2_S/MiMo-V2.6-Flash-RL-IQ2_S-00001-of-00004.ggufGGUFIQ2_S5.7 MBDownload
IQ2_S/MiMo-V2.6-Flash-RL-IQ2_S-00002-of-00004.ggufGGUFIQ2_S46.43 GBDownload
IQ2_S/MiMo-V2.6-Flash-RL-IQ2_S-00003-of-00004.ggufGGUFIQ2_S46.53 GBDownload
IQ2_S/MiMo-V2.6-Flash-RL-IQ2_S-00004-of-00004.ggufGGUFIQ2_S13.34 GBDownload
MXFP4/MiMo-V2.6-Flash-RL-MXFP4-00001-of-00005.ggufGGUFGGUF5.7 MBDownload
MXFP4/MiMo-V2.6-Flash-RL-MXFP4-00002-of-00005.ggufGGUFGGUF45.69 GBDownload
MXFP4/MiMo-V2.6-Flash-RL-MXFP4-00003-of-00005.ggufGGUFGGUF45.69 GBDownload
MXFP4/MiMo-V2.6-Flash-RL-MXFP4-00004-of-00005.ggufGGUFGGUF45.69 GBDownload
MXFP4/MiMo-V2.6-Flash-RL-MXFP4-00005-of-00005.ggufGGUFGGUF25.83 GBDownload
Q3_K/MiMo-V2.6-Flash-RL-Q3_K-00001-of-00004.ggufGGUFQ3_K5.7 MBDownload
Q3_K/MiMo-V2.6-Flash-RL-Q3_K-00002-of-00004.ggufGGUFQ3_K45.85 GBDownload
Q3_K/MiMo-V2.6-Flash-RL-Q3_K-00003-of-00004.ggufGGUFQ3_K46.04 GBDownload
Q3_K/MiMo-V2.6-Flash-RL-Q3_K-00004-of-00004.ggufGGUFQ3_K45.86 GBDownload
imatrix.ggufGGUFGGUF473.3 MBDownload
mmproj-MiMo-V2.6-Flash-RL-BF16.ggufGGUFBF162.56 GBDownload
mmproj-MiMo-V2.6-Flash-RL-F16.ggufGGUFF162.56 GBDownload
mmproj-MiMo-V2.6-Flash-RL-F32.ggufGGUFF324.91 GBDownload
mmproj-MiMo-V2.6-Flash-RL-Q8_0.ggufGGUFQ8_01.46 GBDownload

Model Details

Model IDAesSedai/MiMo-V2.6-Flash-GGUF
AuthorAesSedai
Pipeline—
License—
Base modelXiaomiMiMo/MiMo-V2.6-Flash-RL
Last modified2026-09-22T22:07:30.000Z

Model README

---

base_model:

  • XiaomiMiMo/MiMo-V2.6-Flash-RL

---

Updates

  • 09/22/26: Added 3.5, 2.5, and 2.0 BPW quants using ed's bpw-size PR. The FFNs are the primarily quantized feature, rest of the model remains in Q8_0 / Q6_K

This repo contains specialized MoE-quants for XiaomiMiMo/MiMo-V2.6-Flash-RL. The idea being that given the huge size of the FFN tensors compared to the rest of the tensors in the model, it should be possible to achieve a better quality while keeping the overall size of the entire model smaller compared to a similar naive quantization. To that end, the quantization type default is kept in high quality and the FFN UP + FFN GATE tensors are quanted down along with the FFN DOWN tensors.

The MXFP4 quant is the "full quality" version, as the model has MXFP4 experts.

| Quant | Size | Mixture | PPL | 1-(Mean PPL(Q)/PPL(base)) | KLD |

| :----- | :-------------------- | :------------------------------------ | :------------------ | :------------------------ | :------------------- |

| MXFP4 | 162.89 GiB (4.52 BPW) | BF16 / MXFP4 | 5.149210 ± 0.030596 | +0.0715% | -0.000000 ± 0.000000 |

| Q3_K | 137.75 GiB (3.82 BPW) | Q8_0 / Q3_K / Q3_K / MXFP4 | 5.177623 ± 0.030900 | +0.6237% | 0.123607 ± 0.000648 |

| BPW3.5 | 126.19 GiB (3.50 BPW) | Q8_0 / varies | 5.232510 ± 0.031063 | +1.6904% | 0.138808 ± 0.000718 |

| IQ2_S | 106.31 GiB (2.95 BPW) | Q6_K / IQ2_S / IQ2_S / Q3_K | 5.404774 ± 0.032158 | +5.0382% | 0.178092 ± 0.000878 |

| BPW2.5 | 90.14 GiB (2.50 BPW) | Q8_0 / varies | 5.739578 ± 0.034587 | +11.5449% | 0.243713 ± 0.001158 |

| BPW2.0 | 66.99 GiB (1.86 BPW) | Q6_K / varies | 7.290888 ± 0.046722 | +41.6936% | 0.477636 ± 0.002111 |

!kld_graph

!ppl_graph

Run AesSedai/MiMo-V2.6-Flash-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models