--- license_name: nvidia-open-model-license license_link: https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/ pipeline_tag: text-generation language: - en - es - fr - de - ja - it tags: - nvidia - mlx - apple-silicon - metal base_model: - nvidia/NVIDIA-Nemotron-Labs-3-Elastic-30B-A3B-NVFP4 --- # Nemotron 3 Elastic 30B - MLX Format (Apple Silicon) NVIDIA Nemotron 3 Elastic 30B model converted to MLX format for efficient inference on Apple Silicon (Metal). ## 🚀 Quick Start **Get the MLX model from HuggingFace:** ```bash pip install -U "huggingface_hub[cli]" huggingface-cli download ljupco/Nemotron-3-Elastic-30B-MLX --local-dir ./nemotron-30b-mlx ``` **Run chat (8-bit KV cache by default):** ```bash pip install mlx-lm python chat_mlx.py --model . --max-kv-size 1048576 ``` **For larger context with 4-bit KV cache (more memory savings):** ```bash python chat_mlx.py --model . --max-kv-size 1048576 --kv-bits 4 ``` ## 📊 Model Details - **Format**: MLX NVFP4 (4.5 bits/weight) - **Size**: ~16.5 GB - **Context**: Up to 1M tokens (design limit, hardware-dependent) - **Platform**: Apple Silicon (M1/M2/M3) with macOS ## 💾 Memory Requirements | Config | Peak RAM / Model Size | |--------|----------------------| | 30B conversion | ~59 GB → ~16.5 GB MLX NVFP4 | | 30B inference (1M context, 8-bit KV) | ~70-80 GB | | 30B inference (1M context, 4-bit KV) | ~50-60 GB | | 30B inference (1M context, 16-bit KV) | ~90-100 GB | | Variant | Size | Platform | Status | |---------|------|----------|--------| | Nemotron 3 Elastic 30B NVFP4 | 30B | Apple Silicon | ✅ Ready | | Nemotron 3 Elastic 12B NVFP4 | 12B | Apple Silicon | ✅ Slice + convert | | Nemotron 3 Elastic 23B NVFP4 | 23B | Apple Silicon | ✅ Slice + convert | ## 🎯 Features - **Hybrid Architecture**: Mamba-2 + MoE (Mixture of Experts) + Attention layers - **Elastic Variants**: Supports 12B/23B/30B configurations - **Long Context**: Designed for up to 1M token context window - **Reasoning**: Thinking traces enabled by default ## 📖 Usage ### Basic Chat ```bash python chat_mlx.py --model . ``` ### Large Text Input For texts larger than ~4KB: ``` You> /paste Now paste your text. After the text is pasted, to process the text, in empty line enter /endpaste [paste your large text] /endpaste Assistant> [processes your text] ``` ### In-Chat Commands - `/paste` - Multi-line input mode for large text - `/quit` - Exit chat - `/reset` - Clear conversation history - `/thinking on|off` - Toggle reasoning traces ## 🔧 Conversion Tools This repo includes tools to convert NVIDIA Nemotron 3 Elastic NVFP4 models to MLX format: - **convert_to_mlx.py** - Converts NVFP4 HuggingFace checkpoints to MLX format - **chat_mlx.py** - Interactive chat with MLX models on Apple Silicon - **zero_shot_slicing.py** - Extract 12B/23B variants from 30B elastic checkpoint The converter handles ModelOpt NVFP4 format: 1. Loads all shards (fixes cross-shard scale references) 2. Dequantizes NVFP4 → bfloat16 3. Loads into MLX Nemotron H model 4. Re-quantizes to MLX NVFP4 (4.5 bits/weight) **Key fix:** Original sharding splits weights and scales across different files. The converter loads all shards before dequantizing to handle this correctly. ## 🔗 Links - **Original Model**: https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Elastic-30B-A3B-NVFP4 - **Conversion Tools**: https://github.com/ljubomirj/nemotron-3-elastic-mlx - **NVIDIA License**: https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/ ## 📝 License NVIDIA Open Model License. See [LICENSE.md](LICENSE.md) for details. ## 🙏 Acknowledgments - **Model Developer**: NVIDIA - **MLX Framework**: Apple MLX team - **Conversion**: Adapted from mlx-lm Nemotron H implementation