# Nemotron 3 Elastic MLX Conversion Tools Tools to convert NVIDIA Nemotron 3 Elastic NVFP4 models to MLX format for Apple Silicon (Metal) inference. ## 🚀 Quick Start **Get the MLX model from HuggingFace:** ```bash pip install -U "huggingface_hub[cli]" huggingface-cli download ljupco/Nemotron-3-Elastic-30B-MLX --local-dir ./nemotron-30b-mlx ``` **Run chat (8-bit KV cache by default):** ```bash pip install mlx-lm python chat_mlx.py --model ./nemotron-30b-mlx --max-kv-size 1048576 ``` ## 📁 What This Repo Contains - **convert_to_mlx.py** - Converts NVFP4 HuggingFace checkpoints to MLX format - **chat_mlx.py** - Interactive chat with MLX models on Apple Silicon - **chat.py** - CUDA/Linux inference (requires NVIDIA GPU) - **zero_shot_slicing.py** - Extract 12B/23B variants from 30B elastic checkpoint ## 📖 Documentation - **CHAT_COMMANDS.md** - Chat usage, KV cache options, commands - **CONVERSION_RUN.md** - Conversion details and bug fixes ## 🎯 Supported Models | Model | Size | Platform | Status | |-------|------|----------|--------| | Nemotron 3 Elastic 30B NVFP4 | 30B | Apple Silicon | ✅ Ready | | Nemotron 3 Elastic 12B NVFP4 | 12B | Apple Silicon | ✅ Slice + convert | | Nemotron 3 Elastic 23B NVFP4 | 23B | Apple Silicon | ✅ Slice + convert | ## 🔧 Requirements ### Apple Silicon (MLX) - macOS with M1/M2/M3 chip - 96GB RAM recommended for 30B model - Python 3.13+, mlx-lm ### Linux/NVIDIA (CUDA) - NVIDIA Hopper/Blackwell GPU - CUDA, mamba-ssm, compressed-tensors ## 🏗️ Conversion Process The converter handles ModelOpt NVFP4 format: 1. Loads all shards (fixes cross-shard scale references) 2. Dequantizes NVFP4 → bfloat16 3. Loads into MLX Nemotron H model 4. Re-quantizes to MLX NVFP4 (4.5 bits/weight) **Key fix:** Original sharding splits weights and scales across different files. The converter loads all shards before dequantizing to handle this correctly. ## 📊 Memory Usage | Config | Peak RAM | Model Size | |--------|----------|------------| | 30B conversion | ~59 GB | ~16.5 GB MLX NVFP4 | | 12B conversion | ~24 GB | ~6-7 GB MLX NVFP4 | | 30B inference (1M context, 8-bit KV) | ~70-80 GB | - | ## 🔗 Links - **MLX Model**: https://huggingface.co/ljupco/Nemotron-3-Elastic-30B-MLX - **Original Model**: https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Elastic-30B-A3B-NVFP4 ## 📝 License Same as original NVIDIA model. See LICENSE.md. ## 🤝 Contributing This repository contains conversion tools adapted for Apple Silicon MLX inference. For model-specific issues, refer to NVIDIA's original repository.