MiMo-V2.6-Flash-MOPD-NVFP4

Lossless MXFP4 → NVFP4 transcode of XiaomiMiMo/MiMo-V2.6-Flash-MOPD (revision 2479e2d0029eca9a34cc7e7f55a121925f81908e). No calibration and no re-quantization: the dequantized expert weights are bit-identical to the official checkpoint.

What was changed

  • Routed experts (model.layers.*.mlp.experts.*.{gate,up,down}_proj): the FP4 (E2M1) values are copied bit-for-bit. Each 32-element E8M0 block scale becomes two 16-element E4M3 block scales of value 2^(e − 127 − k), with a per-tensor weight_scale_2 = 2^k and k = max block exponent − 8. gate_proj and up_proj of the same expert share k, because they are fused into one w13 matrix at load time. input_scale = 1.0.
  • Everything else is byte-identical to the official checkpoint: FP8 block-scaled linear layers (weight_scale_inv), BF16 tensors, MTP layers, vision and audio encoders, the DFlash draft model (dflash/), tokenizer and chat template.
  • Quantization config: ModelOpt MIXED_PRECISION (experts NVFP4, group size 16; other linear layers FP8_BLOCK_SCALES).

Verification

  • Tensor layout (145,273 tensors: names, dtypes, shapes) is identical to tiyuvta/MiMo-V2.6-Flash-RL-NVFP4, which uses the same transcode of the RL checkpoint (the same procedure reproduces its tensors byte-for-byte from the RL weights on all sampled tensors).
  • 200 randomly sampled tensors are byte-identical to the independent transcode AxionML/MiMo-V2.6-Flash-MOPD-NVFP4.
  • SHA256SUMS lists every file.
  • Served with vLLM, tensor parallel 4: 17×23 = 391; 5/5 tool-calling checks (single tool, 8 tools, streaming, tool result, two calls per turn); needle retrieval at 16K / 40K / 114K tokens; GSM8K (first 100, temperature 0) 98/100.

Serving notes (vLLM)

  • Use --trust-remote-code --reasoning-parser mimo --tool-call-parser mimo --enable-auto-tool-choice.
  • Use --generation-config vllm: the shipped generation_config.json sets max_new_tokens: 2048, which truncates long reasoning answers. Recommended sampling: temperature=1.0, top_p=0.95.
  • Some vLLM builds route ModelOpt FP8_BLOCK_SCALES layers in mixed-precision checkpoints to the unquantized method, or fail with KeyError: ...qkv_proj.weight_scale_inv; these layers need vLLM's block-FP8 linear method.
  • Check the log line Using '...' NvFp4 MoE backend: FLASHINFER_CUTLASS runs native FP4; MARLIN is a W4A16 fallback (vLLM 0.30.x selects it whenever --enable-lora is set).
  • With the DFlash draft enabled, vLLM rejects min_p and logit_bias.

License

MIT, same as the base model.

Downloads last month
79
Safetensors
Model size
159B params
Tensor type
F32
·
BF16
·
F8_E4M3
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for revorm/MiMo-V2.6-Flash-MOPD-NVFP4

Quantized
(22)
this model