GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Baekpica/MiMo-V2.6-Flash-RL-Mixed-Quant-GGUF overview

MiMo V2.6 Flash RL Mixed Quant GGUF Mixed weights with original representation calibration. The four shard language model passed its 508 tensor structural and …

ggufmixed-quantmimo_v2multimodalaudiovideo-understandingimatrixdgx-sparkimage-text-to-textconversationalenzhbase_model:XiaomiMiMo/MiMo-V2.6-Flash-RLbase_model:quantized:XiaomiMiMo/MiMo-V2.6-Flash-RLlicense:mitendpoints_compatibleregion:us

Runs locally from ~473.3 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
777
Likes
5
Pipeline
image-text-to-text
Author

Repository Files & Downloads

7 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
MQ-IQ2-XXS-XS-Q8-MM-BF16/MiMo-V2.6-Flash-RL-DFlash-Q8_0.ggufGGUFIQ21.46 GBDownload
MQ-IQ2-XXS-XS-Q8-MM-BF16/MiMo-V2.6-Flash-RL-MQ-IQ2-XXS-XS-Q8-00001-of-00004.ggufGGUFIQ220.93 GBDownload
MQ-IQ2-XXS-XS-Q8-MM-BF16/MiMo-V2.6-Flash-RL-MQ-IQ2-XXS-XS-Q8-00002-of-00004.ggufGGUFIQ220.92 GBDownload
MQ-IQ2-XXS-XS-Q8-MM-BF16/MiMo-V2.6-Flash-RL-MQ-IQ2-XXS-XS-Q8-00003-of-00004.ggufGGUFIQ220.92 GBDownload
MQ-IQ2-XXS-XS-Q8-MM-BF16/MiMo-V2.6-Flash-RL-MQ-IQ2-XXS-XS-Q8-00004-of-00004.ggufGGUFIQ219.91 GBDownload
MQ-IQ2-XXS-XS-Q8-MM-BF16/calibration/imatrix.ggufGGUFIQ2473.3 MBDownload
MQ-IQ2-XXS-XS-Q8-MM-BF16/mmproj-MiMo-V2.6-Flash-RL-BF16.ggufGGUFIQ22.56 GBDownload

Model Details

Model IDBaekpica/MiMo-V2.6-Flash-RL-Mixed-Quant-GGUF
AuthorBaekpica
Pipelineimage-text-to-text
Licensemit
Base modelXiaomiMiMo/MiMo-V2.6-Flash-RL
Last modified2026-09-23T23:48:09.000Z

Model README

---

license: mit

base_model: XiaomiMiMo/MiMo-V2.6-Flash-RL

base_model_relation: quantized

pipeline_tag: image-text-to-text

library_name: gguf

language:

- en

- zh

tags:

- gguf

- mixed-quant

- mimo_v2

- multimodal

- audio

- video-understanding

- imatrix

- dgx-spark

---

MiMo-V2.6-Flash-RL Mixed-Quant GGUF

> Mixed weights with original-representation calibration. The four-shard language model passed its 508-tensor structural and quantization audit.

>

> Role-aware precision: expert gate/up IQ2_XXS, expert down IQ2_XS, shared dense paths and embedded MTP Q8_0, media BF16, and numerical controls F32.

>

> Measured GGUF storage: 93.092 GB / 86.699 GiB, including main model, multimodal projector and separate DFlash weights. Auxiliary files and runtime memory are additional.

>

> Text, image, audio, video and audiovisual assets are retained. ds4-dfm-rs on one DGX Spark GB10: serial text is qualified at context 524288 with prefill chunk 4096 and embedded MTP when no DFlash file is loaded. One process at context 262144 also qualifies the external five-layer DFlash draft, this projector, prefill chunk 4096, and finite image, audio, and video completions. max_seqs 2 is omitted because the graph is serial. Context 1048576 is not qualified. Context 524288 together with the projector and DFlash was not measured. Those 262K/512K serving checks do not establish throughput or general quality retention.

Prefill and decode on DGX Spark

Current default f09c1862 reaches 663.67 tok/s at 64K incremental prefill

and 16.22 tok/s decode, versus 61.45 / 1.70 for the historical PR #56

baseline and 634.66 / 1.70 for the earlier P1–P5 candidate.

Across the curve, the median of run means is 854.95 prefill and 18.96

decode tok/s (+342.5% and +326.9% versus #56).

The 64K/2K prefill ratio is 66.2% (#56 7.5%, P1–P5 67.6%).

!MiMo 2K–64K prefill and decode

One GB10, MQ-IQ2-XXS-XS-Q8-MM-BF16, promessi_sposi.txt.

2,048-token incremental steps and 128 greedy tokens at each frontier.

Curves are per-frontier medians; bands are min–max; the legend is the median of run means.

Three fresh processes, one warm session per process. MTP and DFlash off.

The current default adds window HMMA and the adopted decode rounds on top of

global HMMA and async KV. P1–P5 left window HMMA and the decode split off.

Busy SM clocks on this sweep were 2190–2197 MHz, inside 300–2200 MHz.

Not a same-hour paired A/B.

Measurements.

Decode rounds on DGX Spark

Later 8K pairs, 128 greedy tokens, one fresh process per arm. MTP and

DFlash stay off. The base stack adds P6 SWA HMMA and the full-attention

decode split. Busy clocks were 2190–2197 MHz. KV size and the 8K argmax

matched on each pair.

| Round | Switch | 8K decode tok/s | Result |

| --- | --- | ---: | --- |

| D1 | DS4_MIMO2_SWA_DECODE | 21.32 → 21.83 | adopted |

| D2 | DS4_MIMO2_SPLIT_VEC | 21.83 → 22.04 | unadopted |

| D3 | DS4_MIMO2_SWA_VEC | 21.79 → 23.16 | adopted |

One cold 64K run of the adopted stack measured 884.21 tok/s prefill

and 17.10 tok/s decode, clocks 2184–2197 MHz. That single process

does not replace the incremental median above.

Later prefill rounds on DGX Spark

Three same-binary comparisons at one 8,192-token frontier, 128 greedy

tokens, promessi_sposi.txt, MTP and DFlash off. Three fresh processes

per side after a warmup. Busy SM clocks stayed inside 2177–2197 MHz.

That 8,192-token frontier is not the incremental curve above.

| Round | Kill switch | 8K prefill tok/s | Prefill time upper | Decode time envelope |

| --- | --- | ---: | ---: | --- |

| 1 | DS4_MIMO2_NO_PREFILL_HMMA | 630.51 → 871.19 | −27.5% | −1.03% to +0.66% |

| 2 | DS4_MIMO2_NO_PREFILL_ASYNC | 870.38 → 1071.07 | −18.6% | −1.59% to +0.14% |

| 3 | DS4_MIMO2_NO_SWA_HMMA | 1073.29 → 1158.35 | −7.0% | −0.84% to +0.42% |

Unset, windowless and window-128 prefill of 32 or more rows use tensor

cores, and the windowless path stages KV asynchronously. One-row decode

stays on the previous kernels. Rounds 1 and 3 change summation order:

relative RMS was 0.073 and 0.077 under --logit-rel-rms 0.10, and every

frontier argmax matched. Round 2 matched logits and greedy tokens.

Support my work

I work on making large language models practical on constrained hardware through mixed quantization, inference optimization, and serving experiments. Contributions help cover calibration, GPU compute, storage, and testing so these results can be published openly.

<a href="https://www.buymeacoffee.com/baekpica" target="_blank"><img src="https://cdn.buymeacoffee.com/buttons/v2/default-yellow.png" alt="Buy Me a Coffee" style="height: 60px !important;width: 217px !important;"></a> <a href="https://github.com/sponsors/Baekpica" target="_blank"><img src="https://img.shields.io/badge/Sponsor-EA4AAA?style=for-the-badge&logo=githubsponsors&logoColor=white" alt="Sponsor Baekpica on GitHub" style="height: 60px !important;width: 217px !important;"></a>

This is a mixed-precision conversion of XiaomiMiMo/MiMo-V2.6-Flash-RL, pinned to revision 3b38d063180c3e4aed9691fdc735f3d10b266ee4.

Why this model needs an asymmetric layout

The source stores routed expert weights as packed MXFP4. Two logical weights occupy each packed byte, so counting stored tensor elements can produce a roughly 159B figure. The expanded language trunk contains approximately 308.78B logical parameters; the root checkpoint including embedded MTP and media components contains approximately 310.76B, excluding the separate DFlash package. The packed representation does not make this a 159B logical model.

Routed experts account for 302.80B parameters. They dominate the storage budget, so the recipe concentrates compression there and gives their output projections a higher tier. Shared attention and dense paths, the output head, and multimodal components receive substantially more precision.

Variant and availability

| Component | Precision | Measured GB | Measured GiB |

|---|---|---:|---:|

| Main model including three embedded MTP blocks | IQ2_XXS / IQ2_XS / Q8_0 / F32 | 88.778 | 82.681 |

| Multimodal projector | BF16 / F32 | 2.749 | 2.560 |

| Separate DFlash | Q8_0 / F32 | 1.566 | 1.458 |

| All GGUF weights | Mixed | 93.092 | 86.699 |

Main plus media without the optional separate drafter: 91.526 GB / 85.240 GiB.

| File | Bytes |

|---|---:|

| MiMo-V2.6-Flash-RL-MQ-IQ2-XXS-XS-Q8-00001-of-00004.gguf | 22,470,644,192 |

| MiMo-V2.6-Flash-RL-MQ-IQ2-XXS-XS-Q8-00002-of-00004.gguf | 22,464,695,232 |

| MiMo-V2.6-Flash-RL-MQ-IQ2-XXS-XS-Q8-00003-of-00004.gguf | 22,464,695,232 |

| MiMo-V2.6-Flash-RL-MQ-IQ2-XXS-XS-Q8-00004-of-00004.gguf | 21,377,729,024 |

| mmproj-MiMo-V2.6-Flash-RL-BF16.gguf | 2,748,509,792 |

| MiMo-V2.6-Flash-RL-DFlash-Q8_0.gguf | 1,565,911,104 |

Download the complete variant, including template, DFlash configuration and provenance:

hf download Baekpica/MiMo-V2.6-Flash-RL-Mixed-Quant-GGUF \
  --include 'MQ-IQ2-XXS-XS-Q8-MM-BF16/*' --local-dir ./mimo-mixed
cd ./mimo-mixed/MQ-IQ2-XXS-XS-Q8-MM-BF16
sha256sum -c SHA256SUMS

These are download/verification commands. Serving commands await runtime qualification.

Quantization targets

| Model region | Target | Reason |

|---|---|---|

| Routed expert gate and up | IQ2_XXS | Largest share of the parameter budget |

| Routed expert down | IQ2_XS | Higher precision at the return to the residual stream |

| Attention projections and dense FFN matrices | Q8_0 | Shared computation on every token |

| Token embedding and output head | Q8_0 | Preserve input and logit precision |

| Three embedded MTP blocks, eligible matrices | Q8_0 | Preserve checkpoint predictors for later runtime integration |

| Routers, norms, attention sinks, and control tensors | F32 | Preserve routing and numerical control |

| Multimodal encoder/projector matrices | BF16 | Preserve media representation precision |

| Separate DFlash draft matrices | Q8_0 | Optional draft package, separate from embedded MTP |

The per-tensor inventory and recipe accompany the weights. BF16/F32 in this table describes the output storage format; source tensors already stored in FP8 are expanded from that source precision.

Importance matrix and calibration

Calibration uses the original checkpoint representation: expert MXFP4 values are repacked into GGUF without an additional quantization step, while the source FP8 dense tensors are expanded to BF16. This reference is not an original full-BF16 checkpoint and is not calibrated from the final IQ2 artifact.

The corpus starts from Baekpica/Inkling-Small-Multimodal-Calibration, with media recovered from the source datasets and prompts tokenized using MiMo's own tokenizer. Inkling token IDs and embeddings are not reused. Video clips are added from FineVideo.

| Prepared subset | Records | Current state |

|---|---:|---|

| Text reasoning and code/tool text | 667 | Completed 640 × 1,024-token chunks (655,360 tokens) |

| Chart/document images | 486 | Native trunk calibration completed |

| Audio | 309 | Native trunk calibration completed |

| FineVideo clips | 79 | Native trunk calibration completed |

| FineVideo joint audiovisual clips | 79 | Native audiovisual trunk calibration completed |

| FineVideo holdout | 9 | Separated by source video; not part of calibration |

The visual-only clips are paired with a separate set of 79 joint audiovisual inputs using restored original audio and production per-frame-pair interleaving. Every prepared audio token is consumed exactly once in each joint input. All 953 native media records completed trunk calibration with finite final logits. Coverage is 12,030 of 12,032 layer/expert pairs; block 7 experts 13 and 184 remain unobserved. The original recipe is unchanged: raw zero counts are retained and the pinned quantizer uses uniform importance weights for those experts. The final imatrix SHA256 is 265cc19bc1470b95a9157d3b2fab893f325a9e2b6ef414cdbdfac4326c9ef8e3.

Text calibration used raw GGUF Qwen2 BPE without the HF tokenizer NFC normalizer. Six decomposed Y-macron occurrences in four corpus lines differ from NFC input; the collected statistics are preserved as observed. Native media prompts used the original HF tokenizer. The ds4 MiMo tokenizer applies NFC.

Multimodal input contract

Input processing is being aligned with the SGLang MiMo implementation, alongside the pinned checkpoint configuration and tokenizer.

| Input | Required handling | Validation state |

|---|---|---|

| Text | MiMo tokenizer, chat template, and special tokens | Reference text decode passed |

| Image | Production normalization, spatial patch/merge order, vision boundary tokens | Native smoke and one-image GGUF comparison completed; broader checks pending |

| Audio | Original audio tokenizer, local encoder/projector, audio boundary tokens | Native preparation and trunk calibration completed |

| Video | Two-frame temporal patches, MiMo video boundaries, MM:SS timestamps | Native temporal preparation and trunk calibration completed |

| Video with audio | Source timing and native audiovisual token layout | Token-layout, audio-coverage and trunk calibration completed |

Generic video ingestion is insufficient here: using independent image frames or generic timestamps changes the model's input contract. Media imatrix collection must consume the correctly prepared original-model embeddings.

Memory metrics

Exact GGUF file sizes are listed above and in artifact-manifest.json. Main tensor payload is 88,771,782,144 bytes; file sizes additionally include headers, tokenizer metadata and alignment.

Device usage includes KV cache, media activations, allocator overhead, workspaces and the operating system. The qualified serving set is the same one named above. Selective residency beyond that measured process is not claimed.

Conversion and verification

| Check | Current evidence |

|---|---|

| Source download | All 90 source repository files passed Hub checksum verification |

| MXFP4 repacking and TP=4 QKV ordering | Local numerical checks passed |

| Routed-expert source audit | All 36,096 expert matrices sampled across 108,257 deterministic rows; repacked bytes matched |

| Original-representation reference | Four shards and media/DFlash assets published and verified |

| Reference text smoke | GPU load and arithmetic decode passed |

| Text imatrix | Completed 640 chunks / 655,360 tokens |

| Multimodal imatrix and coverage | Completed 953 records / 464,968 tokens; merged raw sums/counts verified |

| Final mixed tensor inventory and checksums | 508-tensor audit passed; six GGUF files listed in SHA256SUMS |

| Embedded MTP | Qualified on CUDA when no DFlash file is loaded |

| DFlash runtime | External five-layer Q8_0 draft. A rejected draft stops at the accepted prefix |

| DGX Spark / ds4-dfm-rs | 524288 serial text, and the 262144 DFlash plus image, audio, and video plan. max_seqs 2 and context 1048576 are not qualified |

The sampled expert audit is not a whole-file byte comparison. A successful reference smoke is not a quality benchmark for the final mixed model. These conversion checks do not establish accuracy or general quality retention. Scoped throughput measurements are reported above.

The original-representation reference and its audit are available in the separate intermediate GGUF repository.

Reproduction and release contents

The variant includes four mixed language shards, a BF16/F32 multimodal projector (including input audio-codec tensors), the separate Q8_0 DFlash package and mask embedding, original chat template, tensor recipe, final imatrix, calibration provenance and coverage, conversion scripts, checksums and validation reports.

See reproduction instructions. The synthesis decoder and training-only audio-codebook statistics are not part of this input-modality projector. No audio-output serving capability is claimed. Original gated source media is not redistributed.

The separate intermediate repository contains the calibration reference. No full Q8_0 or BF16 language baseline was needed.

Chat template

The upstream chat_template.jinja is published alongside this card, copied byte for byte from the pinned source revision. Use it with MiMo’s own tokenizer and special-token mapping. Media preprocessing and embedding insertion remain part of the runtime input contract; the template alone does not implement them.

Runtime support

The current work covers calibration, mixed-quant construction, validation, and the measured ds4-dfm-rs serving set named above. This card does not add a second serving recipe.

License

MIT, inherited from the pinned upstream model. Calibration datasets retain their respective licenses and access conditions.

Run Baekpica/MiMo-V2.6-Flash-RL-Mixed-Quant-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models