MiMo-V2.6-Flash-MOPD — EXL3 2.27 bpw
XiaomiMiMo/MiMo-V2.6-Flash-MOPD (309B total / 15B active MoE) quantized to EXL3 at 2.27 bpw, so it fits on a single 128 GB machine such as an NVIDIA DGX Spark.
MOPD is Xiaomi's update to MiMo-V2.6-Flash-RL. It mainly fixes tool-call repetition in agentic use: the model issuing the same tool calls over and over without making progress. The quant of the earlier checkpoint is benthecarman/MiMo-V2.6-Flash-RL-exl3.
The repo includes the vision tower, the MTP heads, and Xiaomi's DFlash speculative-decoding drafter, fixed so it loads and quantized to 4 bpw.
| Weights | 85.28 GiB, 12 shards, including the MTP heads and the vision tower |
| Bitrate | 2.27 bpw (excluding head), head 6 bpw |
| Perplexity | 5.43, wikitext-2 test, 64 x 2048 tokens |
| Context | 262,144 tokens on a DGX Spark, with the drafter and vision loaded (same footprint as the Flash-RL quant) |
| Modalities | Text and images (no audio) |
Requires exllamav3 v1.5.2 or later; see How to run.
Files
model-*.safetensors, model.safetensors.index.json EXL3 weights
quantization_config.json per-tensor storage record
config.json, tokenizer files, chat_template.jinja, preprocessor_config.json
dflash/ drafter, 4 bpw EXL3 (use this one)
dflash-bf16/ the same drafter, unquantized
Bitrate
Converted with exllamav3 v1.5.2, convert.py -b 2.25 -hq -cr 250 -cc 2048. The per-module
result:
| module | bpw |
|---|---|
| routed experts, layers 12–35 | 2.0 |
| routed experts, layers 1–11 and 36–47 | 2.5 |
| attention | 4.0 |
| dense MLP (layer 0) | 3.0 |
lm_head |
6.0 |
| MTP heads | 4.0 (eh_proj 5.0) |
| embeddings, norms, router, vision tower | BF16 |
This is the same allocation as the Flash-RL quant, so the two are directly comparable.
Layer 47: some of its experts produce intermediate values past the fp16 limit.
exllamav3 scales up_proj down by 128 in that layer (interm_div) and restores the scale in
fp32. The scale is folded into the weights.
Quality
Perplexity (wikitext-2 test, 64 x 2048): 5.4335. The Flash-RL quant measures 5.4003 the same way.
Vision was checked with exllamav3's API on two images. A synthetic image (a red circle, a blue square and the text "INVOICE #4821 TOTAL $317.50") was described exactly. On a photo of a man ironing a shirt on the back of a taxi, it said he was adjusting an umbrella on the roof. The Flash-RL quant, with a byte-identical vision tower and the same script, described the ironing correctly, so the difference comes from the MOPD language model. That is one photo, so treat it as a caution rather than a measurement.
Speed
Measured on the conversion machine, one RTX PRO 6000 (96 GB), exllamav3 v1.5.2, batch 1, greedy, 512 generated tokens, median of 3 runs, in tok/s:
| prompt | no drafter | DFlash drafter | MTP heads |
|---|---|---|---|
| coding | 100.4 | 198.5 | 164.1 |
| prose | 103.4 | 144.7 | 120.1 |
| reasoning | 103.8 | 234.6 | 169.4 |
The DFlash drafter accepted 62–70% of drafted tokens and is the faster drafter on every prompt. The target model verifies every drafted token, so neither drafter affects output quality.
On a DGX Spark, expect roughly a third of these speeds: the Flash-RL quant, which has the same architecture and size, decodes at ~31 tok/s without a drafter and 35–80 tok/s with DFlash there.
How to run
Supported in exllamav3 v1.5.2 and later. Serve it with TabbyAPI.
hf download benthecarman/MiMo-V2.6-Flash-MOPD-exl3 --local-dir mimo-mopd-exl3
There are no prebuilt exllamav3 wheels for aarch64, so on a DGX Spark install it from source.
TabbyAPI also needs pip install uvloop there.
TabbyAPI looks for the drafter by name inside draft_model_dir, so put dflash/ in its own
directory next to the model:
models/
mimo-mopd-exl3/ everything except dflash/ and dflash-bf16/
mimo-dflash-draft/ the contents of dflash/
config.yml:
model:
model_dir: models
model_name: mimo-mopd-exl3
max_seq_len: 262144
cache_size: 262144
cache_mode: FP16
chunk_size: 2048
max_batch_size: 1
vision: true
reasoning: true
reasoning_start_token: "<think>"
reasoning_end_token: "</think>"
draft_model:
draft_mode: model
draft_model_dir: models
draft_model_name: mimo-dflash-draft
draft_cache_mode: FP16
dynamic_draft: true
To use the MTP heads instead of the DFlash drafter:
draft_model:
draft_mode: mtp
dynamic_draft: true
Without vision: true the vision tower is not loaded, which saves about 1.3 GiB.
Thinking can be turned off per request with "chat_template_kwargs": {"enable_thinking": false}.
Tool calls use the qwen3_coder format, which TabbyAPI detects automatically.
Credits
- Xiaomi MiMo team: the model and the DFlash drafter (MIT).
- turboderp: ExLlamaV3 and the EXL3 format.
- vcruz305: the aarch64 build fixes.
- theroyallab: TabbyAPI.
- Downloads last month
- 118
Model tree for benthecarman/MiMo-V2.6-Flash-MOPD-exl3
Base model
XiaomiMiMo/MiMo-V2.6-Flash-MOPD