# Experimental MLX text runtime This standalone adapter implements the source-layout checkpoint in this repository. It is not an oMLX plugin, HTTP server or production inference engine. Python 3.11 on Apple silicon was used for testing; the dependency versions are pinned to the tested environment. From the downloaded model directory: ```bash python3.11 -m venv .venv source .venv/bin/activate python -m pip install -r runtime/requirements.txt python -m unittest discover -s runtime -p 'test_*.py' python runtime/generate.py --model . --resident-backbone --prompt "What is 2+2? Answer briefly." --max-tokens 16 --repeat 2 ``` The first run includes model and kernel warm-up. The second reuses the loaded weights and Engram row cache, matching the warm-cache measurement scope in the model card. Decode speed excludes model loading and prompt processing; compare the printed token counts as well as tokens/sec. The 9.5 headline is a rounded short-test result, not a throughput guarantee for other prompts or contexts. `--resident-backbone` attempts to materialize about 161 GiB of backbone tensors after a conservative memory check. A 256 GiB Mac with other heavy applications closed is the tested setup. Omitting it uses slow disk streaming, which does not reproduce the headline speed. The full weights still require about 239 GB of disk space. Engram tables remain disk-backed; the default in-process cache is bounded to 16,384 rows and can be disabled with `--engram-cache-rows 0`. Add `--mtp` to exercise the model's own three-stage DSpark head. It loads about 4.15 GiB of additional weights and verifies proposals serially against the target model. This mode was slower in the measured tests; it is provided for experimentation, not as an acceleration claim. No external drafter is used. This runner deliberately caps prompt plus output at 128 tokens and generated output at 64 tokens. It uses greedy decoding and the upstream `chat` prompt mode. Vision, tool execution, long context and concurrent requests are not supported. Process memory is released when the command exits. The scripts do not upload prompts, start a network listener, or alter system memory settings. `encoding.py` is the upstream DeepSeek-V4.1 encoder, included unchanged under the root MIT licence. The text and DSpark formulas were ported from DeepSeek's MIT-licensed inference reference. The root model card documents the limited validation scope and quantisation-related quality risks.