Nemotron-3-Elastic-30B-MLX / CHAT_COMMANDS.md
ljupco's picture
Upload folder using huggingface_hub
ff9de85 verified
|
Raw History Blame Contribute Delete
2.76 kB

Nemotron 30B MLX Chat Commands

Basic Usage

# Activate venv
source ~/python3-venv/torch313-metal/bin/activate

# Run with 1M context (8-bit KV cache is now default)
python chat_mlx.py --model ./nemotron-30b-mlx --max-kv-size 1048576

In-Chat Commands

Command Description
/paste Enter multi-line mode for large text pastes (bypasses terminal input limit)
/quit or /q Exit chat
/reset Clear chat history
`/thinking on off`

Large Text Input (/paste)

The standard terminal input is limited to ~4KB. For larger texts:

You> /paste
Now paste your text. After the text is pasted, to process the text, in empty line enter /endpaste
[paste your large text here - can be many lines]
[paste more if needed]
[end with /endpaste on its own line]

Assistant> [processes your text]
[stats: XXX tokens in X.Xs = XX.X tokens/sec]

Command Line Options

Option Default Description
--model ./nemotron-12b-mlx Path to MLX model
--max-context-len 1048576 Total context window (tokens)
--max-kv-size 1048576 KV cache size limit
--kv-bits 8 KV cache quantization: 4, 8, or 16
--kv-group-size 64 KV cache quantization group size
--max-tokens 128000 Max tokens per generation
--no-thinking False Disable reasoning traces
--temperature 1.0 Sampling temperature
--top-p 1.0 Top-p nucleus sampling
--seed 0 PRNG seed

KV Cache Quantization (Memory Saving)

8-bit KV cache is now the default - provides ~50% memory reduction with minimal quality impact.

8-bit (Default)

python chat_mlx.py --model ./nemotron-30b-mlx --max-kv-size 1048576

4-bit (Aggressive memory saving)

python chat_mlx.py --model ./nemotron-30b-mlx --max-kv-size 1048576 --kv-bits 4

16-bit (No quantization)

python chat_mlx.py --model ./nemotron-30b-mlx --max-kv-size 1048576 --kv-bits 16

Response Stats

Each response ends with performance stats:

[stats: 150 tokens in 3.2s = 46.9 tokens/sec]

Memory Estimates

For 1M context with 30B model:

  • 16-bit KV cache (no quantization): ~90-100 GB unified memory
  • 8-bit KV cache (default, recommended): ~70-80 GB unified memory
  • 4-bit KV cache: ~50-60 GB unified memory

Notes

  • The model weights are NVFP4 (4-bit) regardless of KV cache setting
  • KV cache quantization only affects the keys/values stored during generation
  • 8-bit KV cache is the recommended default for most use cases
  • Use 4-bit only if 8-bit still causes OOM on your system
  • Use 16-bit if you need maximum quality and have sufficient memory