Text Generation
Transformers
Safetensors
mimo_v2
nvfp4
modelopt
quantized
multimodal
conversational
custom_code
8-bit precision
Instructions to use revorm/MiMo-V2.6-Flash-MOPD-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use revorm/MiMo-V2.6-Flash-MOPD-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="revorm/MiMo-V2.6-Flash-MOPD-NVFP4", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("revorm/MiMo-V2.6-Flash-MOPD-NVFP4", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use revorm/MiMo-V2.6-Flash-MOPD-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "revorm/MiMo-V2.6-Flash-MOPD-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "revorm/MiMo-V2.6-Flash-MOPD-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/revorm/MiMo-V2.6-Flash-MOPD-NVFP4
- SGLang
How to use revorm/MiMo-V2.6-Flash-MOPD-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "revorm/MiMo-V2.6-Flash-MOPD-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "revorm/MiMo-V2.6-Flash-MOPD-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "revorm/MiMo-V2.6-Flash-MOPD-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "revorm/MiMo-V2.6-Flash-MOPD-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use revorm/MiMo-V2.6-Flash-MOPD-NVFP4 with Docker Model Runner:
docker model run hf.co/revorm/MiMo-V2.6-Flash-MOPD-NVFP4
MiMo-V2.6-Flash-MOPD-NVFP4
Lossless MXFP4 → NVFP4 transcode of XiaomiMiMo/MiMo-V2.6-Flash-MOPD
(revision 2479e2d0029eca9a34cc7e7f55a121925f81908e). No calibration and no re-quantization: the dequantized expert
weights are bit-identical to the official checkpoint.
What was changed
- Routed experts (
model.layers.*.mlp.experts.*.{gate,up,down}_proj): the FP4 (E2M1) values are copied bit-for-bit. Each 32-element E8M0 block scale becomes two 16-element E4M3 block scales of value 2^(e − 127 − k), with a per-tensorweight_scale_2= 2^k and k = max block exponent − 8.gate_projandup_projof the same expert share k, because they are fused into one w13 matrix at load time.input_scale= 1.0. - Everything else is byte-identical to the official checkpoint: FP8 block-scaled linear layers (
weight_scale_inv), BF16 tensors, MTP layers, vision and audio encoders, the DFlash draft model (dflash/), tokenizer and chat template. - Quantization config: ModelOpt
MIXED_PRECISION(expertsNVFP4, group size 16; other linear layersFP8_BLOCK_SCALES).
Verification
- Tensor layout (145,273 tensors: names, dtypes, shapes) is identical to tiyuvta/MiMo-V2.6-Flash-RL-NVFP4, which uses the same transcode of the RL checkpoint (the same procedure reproduces its tensors byte-for-byte from the RL weights on all sampled tensors).
- 200 randomly sampled tensors are byte-identical to the independent transcode AxionML/MiMo-V2.6-Flash-MOPD-NVFP4.
SHA256SUMSlists every file.- Served with vLLM, tensor parallel 4: 17×23 = 391; 5/5 tool-calling checks (single tool, 8 tools, streaming, tool result, two calls per turn); needle retrieval at 16K / 40K / 114K tokens; GSM8K (first 100, temperature 0) 98/100.
Serving notes (vLLM)
- Use
--trust-remote-code --reasoning-parser mimo --tool-call-parser mimo --enable-auto-tool-choice. - Use
--generation-config vllm: the shippedgeneration_config.jsonsetsmax_new_tokens: 2048, which truncates long reasoning answers. Recommended sampling:temperature=1.0,top_p=0.95. - Some vLLM builds route ModelOpt
FP8_BLOCK_SCALESlayers in mixed-precision checkpoints to the unquantized method, or fail withKeyError: ...qkv_proj.weight_scale_inv; these layers need vLLM's block-FP8 linear method. - Check the log line
Using '...' NvFp4 MoE backend:FLASHINFER_CUTLASSruns native FP4;MARLINis a W4A16 fallback (vLLM 0.30.x selects it whenever--enable-lorais set). - With the DFlash draft enabled, vLLM rejects
min_pandlogit_bias.
License
MIT, same as the base model.
- Downloads last month
- 79
Model tree for revorm/MiMo-V2.6-Flash-MOPD-NVFP4
Base model
XiaomiMiMo/MiMo-V2.6-Flash-MOPD