Instructions to use 0xSojalSec/DeepSeek-V4.1-Flash-MLX-MAC with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use 0xSojalSec/DeepSeek-V4.1-Flash-MLX-MAC with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # if on a CUDA device, also pip install mlx[cuda] # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("0xSojalSec/DeepSeek-V4.1-Flash-MLX-MAC") prompt = "Once upon a time in" text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use 0xSojalSec/DeepSeek-V4.1-Flash-MLX-MAC with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Generate some text mlx_lm.generate --model "0xSojalSec/DeepSeek-V4.1-Flash-MLX-MAC" --prompt "Once upon a time"
- Atomic Chat
Download runtime/README.md from 0xSojalSec/DeepSeek-V4.1-Flash-MLX-MAC: direct link, hf CLI and curl.
- Browser
- Download file 2.48 kB
-
https://huggingface.co/0xSojalSec/DeepSeek-V4.1-Flash-MLX-MAC/resolve/main/runtime/README.md
- Command line
-
hf download hf://0xSojalSec/DeepSeek-V4.1-Flash-MLX-MAC/runtime/README.md
-
curl -L -o README.md https://huggingface.co/0xSojalSec/DeepSeek-V4.1-Flash-MLX-MAC/resolve/main/runtime/README.md
Experimental MLX text runtime
This standalone adapter implements the source-layout checkpoint in this repository. It is not an oMLX plugin, HTTP server or production inference engine. Python 3.11 on Apple silicon was used for testing; the dependency versions are pinned to the tested environment.
From the downloaded model directory:
python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install -r runtime/requirements.txt
python -m unittest discover -s runtime -p 'test_*.py'
python runtime/generate.py --model . --resident-backbone --prompt "What is 2+2? Answer briefly." --max-tokens 16 --repeat 2
The first run includes model and kernel warm-up. The second reuses the loaded weights and Engram row cache, matching the warm-cache measurement scope in the model card. Decode speed excludes model loading and prompt processing; compare the printed token counts as well as tokens/sec. The 9.5 headline is a rounded short-test result, not a throughput guarantee for other prompts or contexts.
--resident-backbone attempts to materialize about 161 GiB of backbone tensors
after a conservative memory check. A 256 GiB Mac with other heavy applications
closed is the tested setup. Omitting it uses slow disk streaming, which does
not reproduce the headline speed. The full weights still require about 239 GB
of disk space. Engram tables remain disk-backed; the default in-process cache
is bounded to 16,384 rows and can be disabled with --engram-cache-rows 0.
Add --mtp to exercise the model's own three-stage DSpark head. It loads about
4.15 GiB of additional weights and verifies proposals serially against the
target model. This mode was slower in the measured tests; it is provided for
experimentation, not as an acceleration claim. No external drafter is used.
This runner deliberately caps prompt plus output at 128 tokens and generated
output at 64 tokens. It uses greedy decoding and the upstream chat prompt
mode. Vision, tool execution, long context and concurrent requests are not
supported. Process memory is released when the command exits. The scripts do
not upload prompts, start a network listener, or alter system memory settings.
encoding.py is the upstream DeepSeek-V4.1 encoder, included unchanged under
the root MIT licence. The text and DSpark formulas were ported from DeepSeek's
MIT-licensed inference reference. The root model card documents the limited
validation scope and quantisation-related quality risks.