GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

ggml-org/MiMo-V2.6-Flash-RL-GGUF overview

MiMo V2.6 Flash RL Run with https://llama.app bash llama serve hf ggml org/MiMo V2.6 Flash RL GGUF Source models https://huggingface.co/XiaomiMiMo/MiMo V2.6 Fl…

ggufquantizedimage-text-to-textbase_model:XiaomiMiMo/MiMo-V2.6-Flash-RLbase_model:quantized:XiaomiMiMo/MiMo-V2.6-Flash-RLlicense:mitendpoints_compatibleregion:usconversational

Runs locally from ~5.7 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
5,966
Likes
12
Pipeline
image-text-to-text
Author

Repository Files & Downloads

11 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
MiMo-V2.6-Flash-RL-MXFP4-00001-of-00002.ggufGGUFGGUF5.7 MBDownload
MiMo-V2.6-Flash-RL-MXFP4-00002-of-00002.ggufGGUFGGUF155.87 GBDownload
MiMo-V2.6-Flash-RL-Q2_K-00001-of-00002.ggufGGUFQ2_K5.7 MBDownload
MiMo-V2.6-Flash-RL-Q2_K-00002-of-00002.ggufGGUFQ2_K117.54 GBDownload
dflash-MiMo-V2.6-Flash-RL-BF16.ggufGGUFBF162.74 GBDownload
dflash-MiMo-V2.6-Flash-RL-Q8_0.ggufGGUFQ8_01.46 GBDownload
mmproj-MiMo-V2.6-Flash-RL-BF16.ggufGGUFBF162.56 GBDownload
mmproj-MiMo-V2.6-Flash-RL-Q8_0.ggufGGUFQ8_01.46 GBDownload
mtp-MiMo-V2.6-Flash-RL-BF16.ggufGGUFBF164.17 GBDownload
mtp-MiMo-V2.6-Flash-RL-Q4_0.ggufGGUFQ4_01.18 GBDownload
mtp-MiMo-V2.6-Flash-RL-Q8_0.ggufGGUFQ8_02.22 GBDownload

Model Details

Model IDggml-org/MiMo-V2.6-Flash-RL-GGUF
Authorggml-org
Pipelineimage-text-to-text
Licensemit
Base modelXiaomiMiMo/MiMo-V2.6-Flash-RL
Last modified2026-09-25T13:55:37.000Z

Model README

---

license: mit

pipeline_tag: image-text-to-text

tags:

  • gguf
  • quantized

base_model:

  • XiaomiMiMo/MiMo-V2.6-Flash-RL

---

MiMo-V2.6-Flash-RL

Run with https://llama.app

llama serve -hf ggml-org/MiMo-V2.6-Flash-RL-GGUF

Source models

  • https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL

Notes

  • The MXFP4 output keeps the routed experts at their native MXFP4 precision.
  • The Q2_K output keeps the expert down projections at MXFP4, and quantizes the gate/up projections to Q2_K.
  • Includes MTP sidecars (Q4_0 and Q8_0) for speculative decoding (--mtp).
  • Includes a DFlash drafter sidecar (BF16 and Q8_0) for speculative decoding, converted from the dflash/ subdirectory of the source repo.
  • Includes a Q8_0 mmproj for the vision and audio encoders.
  • Currently, the Q2 models do not use an imatrix calibration due to lack of one.

> [!IMPORTANT]

> This model is automatically converted using https://github.com/ggml-org/convert

Run ggml-org/MiMo-V2.6-Flash-RL-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models