Text Generation
MLX
Safetensors
nemotron_h
nvidia
apple-silicon
metal
conversational
custom_code
4-bit precision
Instructions to use ljupco/Nemotron-3-Elastic-30B-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use ljupco/Nemotron-3-Elastic-30B-MLX with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("ljupco/Nemotron-3-Elastic-30B-MLX") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use ljupco/Nemotron-3-Elastic-30B-MLX with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ljupco/Nemotron-3-Elastic-30B-MLX"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ljupco/Nemotron-3-Elastic-30B-MLX" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use ljupco/Nemotron-3-Elastic-30B-MLX with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "ljupco/Nemotron-3-Elastic-30B-MLX"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "ljupco/Nemotron-3-Elastic-30B-MLX" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ljupco/Nemotron-3-Elastic-30B-MLX", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use ljupco/Nemotron-3-Elastic-30B-MLX with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ljupco/Nemotron-3-Elastic-30B-MLX"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ljupco/Nemotron-3-Elastic-30B-MLX
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ljupco/Nemotron-3-Elastic-30B-MLX with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ljupco/Nemotron-3-Elastic-30B-MLX"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ljupco/Nemotron-3-Elastic-30B-MLX" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
|
Download CHAT_COMMANDS.md from ljupco/Nemotron-3-Elastic-30B-MLX: direct link, hf CLI and curl.
- Browser
- Download file 2.76 kB
-
https://huggingface.co/ljupco/Nemotron-3-Elastic-30B-MLX/resolve/main/CHAT_COMMANDS.md
- Command line
-
hf download hf://ljupco/Nemotron-3-Elastic-30B-MLX/CHAT_COMMANDS.md
-
curl -L -o CHAT_COMMANDS.md https://huggingface.co/ljupco/Nemotron-3-Elastic-30B-MLX/resolve/main/CHAT_COMMANDS.md
2.76 kB
Nemotron 30B MLX Chat Commands
Basic Usage
# Activate venv
source ~/python3-venv/torch313-metal/bin/activate
# Run with 1M context (8-bit KV cache is now default)
python chat_mlx.py --model ./nemotron-30b-mlx --max-kv-size 1048576
In-Chat Commands
| Command | Description |
|---|---|
/paste |
Enter multi-line mode for large text pastes (bypasses terminal input limit) |
/quit or /q |
Exit chat |
/reset |
Clear chat history |
| `/thinking on | off` |
Large Text Input (/paste)
The standard terminal input is limited to ~4KB. For larger texts:
You> /paste
Now paste your text. After the text is pasted, to process the text, in empty line enter /endpaste
[paste your large text here - can be many lines]
[paste more if needed]
[end with /endpaste on its own line]
Assistant> [processes your text]
[stats: XXX tokens in X.Xs = XX.X tokens/sec]
Command Line Options
| Option | Default | Description |
|---|---|---|
--model |
./nemotron-12b-mlx |
Path to MLX model |
--max-context-len |
1048576 | Total context window (tokens) |
--max-kv-size |
1048576 | KV cache size limit |
--kv-bits |
8 | KV cache quantization: 4, 8, or 16 |
--kv-group-size |
64 | KV cache quantization group size |
--max-tokens |
128000 | Max tokens per generation |
--no-thinking |
False | Disable reasoning traces |
--temperature |
1.0 | Sampling temperature |
--top-p |
1.0 | Top-p nucleus sampling |
--seed |
0 | PRNG seed |
KV Cache Quantization (Memory Saving)
8-bit KV cache is now the default - provides ~50% memory reduction with minimal quality impact.
8-bit (Default)
python chat_mlx.py --model ./nemotron-30b-mlx --max-kv-size 1048576
4-bit (Aggressive memory saving)
python chat_mlx.py --model ./nemotron-30b-mlx --max-kv-size 1048576 --kv-bits 4
16-bit (No quantization)
python chat_mlx.py --model ./nemotron-30b-mlx --max-kv-size 1048576 --kv-bits 16
Response Stats
Each response ends with performance stats:
[stats: 150 tokens in 3.2s = 46.9 tokens/sec]
Memory Estimates
For 1M context with 30B model:
- 16-bit KV cache (no quantization): ~90-100 GB unified memory
- 8-bit KV cache (default, recommended): ~70-80 GB unified memory
- 4-bit KV cache: ~50-60 GB unified memory
Notes
- The model weights are NVFP4 (4-bit) regardless of KV cache setting
- KV cache quantization only affects the keys/values stored during generation
- 8-bit KV cache is the recommended default for most use cases
- Use 4-bit only if 8-bit still causes OOM on your system
- Use 16-bit if you need maximum quality and have sufficient memory