Instructions to use ljupco/Nemotron-3-Elastic-30B-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use ljupco/Nemotron-3-Elastic-30B-MLX with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("ljupco/Nemotron-3-Elastic-30B-MLX") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use ljupco/Nemotron-3-Elastic-30B-MLX with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ljupco/Nemotron-3-Elastic-30B-MLX"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ljupco/Nemotron-3-Elastic-30B-MLX" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use ljupco/Nemotron-3-Elastic-30B-MLX with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "ljupco/Nemotron-3-Elastic-30B-MLX"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "ljupco/Nemotron-3-Elastic-30B-MLX" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ljupco/Nemotron-3-Elastic-30B-MLX", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use ljupco/Nemotron-3-Elastic-30B-MLX with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ljupco/Nemotron-3-Elastic-30B-MLX"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ljupco/Nemotron-3-Elastic-30B-MLX
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ljupco/Nemotron-3-Elastic-30B-MLX with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ljupco/Nemotron-3-Elastic-30B-MLX"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ljupco/Nemotron-3-Elastic-30B-MLX" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Nemotron 3 Elastic 30B - MLX Format (Apple Silicon)
NVIDIA Nemotron 3 Elastic 30B model converted to MLX format for efficient inference on Apple Silicon (Metal).
🚀 Quick Start
Get the MLX model from HuggingFace:
pip install -U "huggingface_hub[cli]"
huggingface-cli download ljupco/Nemotron-3-Elastic-30B-MLX --local-dir ./nemotron-30b-mlx
Run chat (8-bit KV cache by default):
pip install mlx-lm
python chat_mlx.py --model . --max-kv-size 1048576
For larger context with 4-bit KV cache (more memory savings):
python chat_mlx.py --model . --max-kv-size 1048576 --kv-bits 4
📊 Model Details
- Format: MLX NVFP4 (4.5 bits/weight)
- Size: ~16.5 GB
- Context: Up to 1M tokens (design limit, hardware-dependent)
- Platform: Apple Silicon (M1/M2/M3) with macOS
💾 Memory Requirements
| Config | Peak RAM / Model Size |
|---|---|
| 30B conversion | ~59 GB → ~16.5 GB MLX NVFP4 |
| 30B inference (1M context, 8-bit KV) | ~70-80 GB |
| 30B inference (1M context, 4-bit KV) | ~50-60 GB |
| 30B inference (1M context, 16-bit KV) | ~90-100 GB |
| Variant | Size | Platform | Status |
|---|---|---|---|
| Nemotron 3 Elastic 30B NVFP4 | 30B | Apple Silicon | ✅ Ready |
| Nemotron 3 Elastic 12B NVFP4 | 12B | Apple Silicon | ✅ Slice + convert |
| Nemotron 3 Elastic 23B NVFP4 | 23B | Apple Silicon | ✅ Slice + convert |
🎯 Features
- Hybrid Architecture: Mamba-2 + MoE (Mixture of Experts) + Attention layers
- Elastic Variants: Supports 12B/23B/30B configurations
- Long Context: Designed for up to 1M token context window
- Reasoning: Thinking traces enabled by default
📖 Usage
Basic Chat
python chat_mlx.py --model .
Large Text Input
For texts larger than ~4KB:
You> /paste
Now paste your text. After the text is pasted, to process the text, in empty line enter /endpaste
[paste your large text]
/endpaste
Assistant> [processes your text]
In-Chat Commands
/paste- Multi-line input mode for large text/quit- Exit chat/reset- Clear conversation history/thinking on|off- Toggle reasoning traces
🔧 Conversion Tools
This repo includes tools to convert NVIDIA Nemotron 3 Elastic NVFP4 models to MLX format:
- convert_to_mlx.py - Converts NVFP4 HuggingFace checkpoints to MLX format
- chat_mlx.py - Interactive chat with MLX models on Apple Silicon
- zero_shot_slicing.py - Extract 12B/23B variants from 30B elastic checkpoint
The converter handles ModelOpt NVFP4 format:
- Loads all shards (fixes cross-shard scale references)
- Dequantizes NVFP4 → bfloat16
- Loads into MLX Nemotron H model
- Re-quantizes to MLX NVFP4 (4.5 bits/weight)
Key fix: Original sharding splits weights and scales across different files. The converter loads all shards before dequantizing to handle this correctly.
🔗 Links
- Original Model: https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Elastic-30B-A3B-NVFP4
- Conversion Tools: https://github.com/ljubomirj/nemotron-3-elastic-mlx
- NVIDIA License: https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/
📝 License
NVIDIA Open Model License. See LICENSE.md for details.
🙏 Acknowledgments
- Model Developer: NVIDIA
- MLX Framework: Apple MLX team
- Conversion: Adapted from mlx-lm Nemotron H implementation
- Downloads last month
- 264
4-bit
Model tree for ljupco/Nemotron-3-Elastic-30B-MLX
Base model
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16