MiMo-V2.6-Flash-MOPD — EXL3 2.27 bpw

XiaomiMiMo/MiMo-V2.6-Flash-MOPD (309B total / 15B active MoE) quantized to EXL3 at 2.27 bpw, so it fits on a single 128 GB machine such as an NVIDIA DGX Spark.

MOPD is Xiaomi's update to MiMo-V2.6-Flash-RL. It mainly fixes tool-call repetition in agentic use: the model issuing the same tool calls over and over without making progress. The quant of the earlier checkpoint is benthecarman/MiMo-V2.6-Flash-RL-exl3.

The repo includes the vision tower, the MTP heads, and Xiaomi's DFlash speculative-decoding drafter, fixed so it loads and quantized to 4 bpw.

Weights 85.28 GiB, 12 shards, including the MTP heads and the vision tower
Bitrate 2.27 bpw (excluding head), head 6 bpw
Perplexity 5.43, wikitext-2 test, 64 x 2048 tokens
Context 262,144 tokens on a DGX Spark, with the drafter and vision loaded (same footprint as the Flash-RL quant)
Modalities Text and images (no audio)

Requires exllamav3 v1.5.2 or later; see How to run.

Files

model-*.safetensors, model.safetensors.index.json   EXL3 weights
quantization_config.json                            per-tensor storage record
config.json, tokenizer files, chat_template.jinja, preprocessor_config.json
dflash/        drafter, 4 bpw EXL3 (use this one)
dflash-bf16/   the same drafter, unquantized

Bitrate

Converted with exllamav3 v1.5.2, convert.py -b 2.25 -hq -cr 250 -cc 2048. The per-module result:

module bpw
routed experts, layers 12–35 2.0
routed experts, layers 1–11 and 36–47 2.5
attention 4.0
dense MLP (layer 0) 3.0
lm_head 6.0
MTP heads 4.0 (eh_proj 5.0)
embeddings, norms, router, vision tower BF16

This is the same allocation as the Flash-RL quant, so the two are directly comparable.

Layer 47: some of its experts produce intermediate values past the fp16 limit. exllamav3 scales up_proj down by 128 in that layer (interm_div) and restores the scale in fp32. The scale is folded into the weights.

Quality

Perplexity (wikitext-2 test, 64 x 2048): 5.4335. The Flash-RL quant measures 5.4003 the same way.

Vision was checked with exllamav3's API on two images. A synthetic image (a red circle, a blue square and the text "INVOICE #4821 TOTAL $317.50") was described exactly. On a photo of a man ironing a shirt on the back of a taxi, it said he was adjusting an umbrella on the roof. The Flash-RL quant, with a byte-identical vision tower and the same script, described the ironing correctly, so the difference comes from the MOPD language model. That is one photo, so treat it as a caution rather than a measurement.

Speed

Measured on the conversion machine, one RTX PRO 6000 (96 GB), exllamav3 v1.5.2, batch 1, greedy, 512 generated tokens, median of 3 runs, in tok/s:

prompt no drafter DFlash drafter MTP heads
coding 100.4 198.5 164.1
prose 103.4 144.7 120.1
reasoning 103.8 234.6 169.4

The DFlash drafter accepted 62–70% of drafted tokens and is the faster drafter on every prompt. The target model verifies every drafted token, so neither drafter affects output quality.

On a DGX Spark, expect roughly a third of these speeds: the Flash-RL quant, which has the same architecture and size, decodes at ~31 tok/s without a drafter and 35–80 tok/s with DFlash there.

How to run

Supported in exllamav3 v1.5.2 and later. Serve it with TabbyAPI.

hf download benthecarman/MiMo-V2.6-Flash-MOPD-exl3 --local-dir mimo-mopd-exl3

There are no prebuilt exllamav3 wheels for aarch64, so on a DGX Spark install it from source. TabbyAPI also needs pip install uvloop there.

TabbyAPI looks for the drafter by name inside draft_model_dir, so put dflash/ in its own directory next to the model:

models/
  mimo-mopd-exl3/      everything except dflash/ and dflash-bf16/
  mimo-dflash-draft/   the contents of dflash/

config.yml:

model:
  model_dir: models
  model_name: mimo-mopd-exl3
  max_seq_len: 262144
  cache_size: 262144
  cache_mode: FP16
  chunk_size: 2048
  max_batch_size: 1
  vision: true
  reasoning: true
  reasoning_start_token: "<think>"
  reasoning_end_token: "</think>"

draft_model:
  draft_mode: model
  draft_model_dir: models
  draft_model_name: mimo-dflash-draft
  draft_cache_mode: FP16
  dynamic_draft: true

To use the MTP heads instead of the DFlash drafter:

draft_model:
  draft_mode: mtp
  dynamic_draft: true

Without vision: true the vision tower is not loaded, which saves about 1.3 GiB.

Thinking can be turned off per request with "chat_template_kwargs": {"enable_thinking": false}. Tool calls use the qwen3_coder format, which TabbyAPI detects automatically.

Credits

  • Xiaomi MiMo team: the model and the DFlash drafter (MIT).
  • turboderp: ExLlamaV3 and the EXL3 format.
  • vcruz305: the aarch64 build fixes.
  • theroyallab: TabbyAPI.
Downloads last month
118
Safetensors
Model size
46B params
Tensor type
BF16
·
F16
·
I16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for benthecarman/MiMo-V2.6-Flash-MOPD-exl3

Quantized
(22)
this model