FP4 kv path for a single 5090

#5
by sHEL1562 - opened

Added an FP4 kv route via a prior user's implementation for FP4 on the sm120 path.

Adds a lot more wiggle room, allowing for 3 sessions at 90k. Probably extendable, but greatly worthwhile speed-wise now that I've tested the speculative MTP.

Try it yourself (--kv-cache-dtype nvfp4) after applying patches from this particular PR: https://github.com/hikarioyama/vllm-nvfp4-kv-sm120/pull/2

OrcaRouter org

Nice — thanks for the sm120 KV writeup. FP4 KV is exactly the headroom lever on a single 5090, so 3× 90k tracks, and good it plays well with the MTP drafter.

Before we link this from the model card, could you share a few numbers so others can judge the tradeoff? Mainly any long-context quality hit (a rough PPL or recall check, FP4 KV vs FP8/FP16 KV on the same box), the tok/s and MTP accept rate you're seeing, and the vLLM commit you patched on top of. If the quality holds up we'll link this thread + your PR with credit to you. 🙌

Certainly, also thanks for this release guys! Orca did amazing here.

Well first and foremost it's an impressive model. With the fp8 scales I was a little blown away with how accurate the routing was at 90k. And—that's about where it stood. With 30tok/s average, I could manage out to about 3 sessions. Activating the drafter meant vllm throwing the warning that there was even less headroom than implied. At a 50k recommendation.

After patching this FP4 path I didn't notice the typical tradeoffs in vllm cache quantization (estimation of scales to 1.0) — because the right FP8 path already existed.
As stated in the PR, there were 3 options, disable the drafter (30tok/s acceptance), lower the context (get drafter bonuses but drop to 40-50k context), or patch some random found sm120 FP4 target (pointed out to me by this very qwen model, actually, when using firecrawl in a chat session)

So FP4 is beneficial on this model because the checkpoint ships with calibrated FP8 scales for the FP4 weights — vLLM isn't estimating anything at runtime, and that's exactly what determines how far you can push context before drift.

The "FP8 falloff" people complain about — drift at 8–16k, degradation at 32k, collapse past 50k — is an artifact of the KV cache running on estimated or static scales (vLLM's 1.0 fallback, missing scale tensors, misaligned groups). KV error grows roughly linearly with sequence length. The tell is log lines like estimated q scales at 1.0 error / fallback to static scale 1.0 — i.e. "we don't have real scales, so we're guessing."

This run doesn't do that. vLLM selected CutlassNvFp4LinearKernel for the NVFP4 GEMM, and it only does that when the FP4 weights, FP8 scales, and NVFP4 metadata are all present and valid. If the scales were missing or broken you'd see Selected CutlassFP8ScaledMMLinearKernel plus the static-scale fallback lines instead.

So what you're getting is the normal FP8-KV drift — roughly 5–10% degradation at 80–100k, mild late-layer precision loss, occasional repetition under aggressive sampling — not the estimation-based collapse. Stable to ~80–120k, which is why the 90k context runs are holding up.

click to expand

Video of it working (Odysseus Harness):

It also wrote this post when on a run of seeing "whether it's worth it" to switch to single-session SGLang variants.
https://a.shel.sh/#ss:~y5g,#'&?784gOO#FPW'@Y@R!eFJ$(OKIbUsX!Az-ggdaT'#

Seems like overall it's very solid for the target range. At 90k one wouldn't expect much falloff. 1-2% tops if no weird routing problems create some super early spontaneous chain of failures.
Very impressive that this little 27B is still beating the 180B MoE on most coding specialized tasks. Getting over the hump with the tradeoff in speed is what a loooot of users are focused on. Glad to have this model now.

Mainly any long-context quality hit (a rough PPL or recall check, FP4 KV vs FP8/FP16 KV on the same box)

None that I could see, I'll have to keep updated and update this comment as I try more coding sessions specifically. Writing/research has had zero errors as of yet

the tok/s and MTP accept rate you're seeing

60 in-harness, 80 reported by vllm at times. Likewise MTP (not at all scientifically) is taken out of console logs:

(APIServer pid=2485494) INFO 09-07 00:10:27 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 3.13, Accepted throughput: 54.00 tokens/s, Drafted throughput: 75.90 tokens/s, Accepted: 540 token
s, Drafted: 759 tokens, Per-position acceptance rate: 0.854, 0.688, 0.593, Avg Draft acceptance rate: 71.1%
(APIServer pid=2485494) INFO 09-07 00:10:37 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 63.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage:
 32.6%, Prefix cache hit rate: 0.0%, MM cache hit rate: 0.0%
(APIServer pid=2485494) INFO 09-07 00:10:37 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.53, Accepted throughput: 38.50 tokens/s, Drafted throughput: 75.30 tokens/s, Accepted: 385 token
s, Drafted: 753 tokens, Per-position acceptance rate: 0.721, 0.466, 0.347, Avg Draft acceptance rate: 51.1%
(APIServer pid=2485494) INFO 09-07 00:10:47 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 74.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage:
 32.6%, Prefix cache hit rate: 0.0%, MM cache hit rate: 0.0%
(APIServer pid=2485494) INFO 09-07 00:10:47 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.94, Accepted throughput: 49.40 tokens/s, Drafted throughput: 76.20 tokens/s, Accepted: 494 token
s, Drafted: 762 tokens, Per-position acceptance rate: 0.819, 0.622, 0.504, Avg Draft acceptance rate: 64.8%
(APIServer pid=2485494) INFO 09-07 00:10:57 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 76.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage:
 32.6%, Prefix cache hit rate: 0.0%, MM cache hit rate: 0.0%
(APIServer pid=2485494) INFO 09-07 00:10:57 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 3.13, Accepted throughput: 52.29 tokens/s, Drafted throughput: 73.79 tokens/s, Accepted: 523 token
s, Drafted: 738 tokens, Per-position acceptance rate: 0.829, 0.703, 0.593, Avg Draft acceptance rate: 70.9%
(APIServer pid=2485494) INFO 09-07 00:11:07 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 95.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage:
 34.8%, Prefix cache hit rate: 0.0%, MM cache hit rate: 0.0%
(APIServer pid=2485494) INFO 09-07 00:11:07 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 3.80, Accepted throughput: 70.29 tokens/s, Drafted throughput: 75.29 tokens/s, Accepted: 703 token
s, Drafted: 753 tokens, Per-position acceptance rate: 0.968, 0.928, 0.904, Avg Draft acceptance rate: 93.4%
(APIServer pid=2485494) INFO 09-07 00:11:17 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 8.2 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage:
0.0%, Prefix cache hit rate: 0.0%, MM cache hit rate: 0.0%
(APIServer pid=2485494) INFO 09-07 00:11:17 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 3.42, Accepted throughput: 5.80 tokens/s, Drafted throughput: 7.20 tokens/s, Accepted: 58 tokens,
Drafted: 72 tokens, Per-position acceptance rate: 0.917, 0.792, 0.708, Avg Draft acceptance rate: 80.6%
(APIServer pid=2485494) INFO 09-07 00:11:27 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage:
0.0%, Prefix cache hit rate: 0.0%, MM cache hit rate: 0.0%
(APIServer pid=2485494) INFO:     127.0.0.1:54578 - "OPTIONS /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=2485494) INFO:     127.0.0.1:54578 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=2485494) INFO 09-07 00:11:37 [loggers.py:310] Engine 000: Avg prompt throughput: 363.1 tokens/s, Avg generation throughput: 1.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage
: 30.4%, Prefix cache hit rate: 0.0%, MM cache hit rate: 0.0%
(APIServer pid=2485494) INFO 09-07 00:11:37 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.40, Accepted throughput: 0.35 tokens/s, Drafted throughput: 0.75 tokens/s, Accepted: 7 tokens, D
rafted: 15 tokens, Per-position acceptance rate: 0.600, 0.400, 0.400, Avg Draft acceptance rate: 46.7%
(APIServer pid=2485494) INFO 09-07 00:11:47 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 67.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage:
 30.4%, Prefix cache hit rate: 0.0%, MM cache hit rate: 0.0%
(APIServer pid=2485494) INFO 09-07 00:11:47 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.66, Accepted throughput: 41.90 tokens/s, Drafted throughput: 75.90 tokens/s, Accepted: 419 token
s, Drafted: 759 tokens, Per-position acceptance rate: 0.787, 0.522, 0.348, Avg Draft acceptance rate: 55.2%
(APIServer pid=2485494) INFO 09-07 00:11:57 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 76.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage:
 30.4%, Prefix cache hit rate: 0.0%, MM cache hit rate: 0.0%
(APIServer pid=2485494) INFO 09-07 00:11:57 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 3.01, Accepted throughput: 51.29 tokens/s, Drafted throughput: 76.49 tokens/s, Accepted: 513 token
s, Drafted: 765 tokens, Per-position acceptance rate: 0.851, 0.655, 0.506, Avg Draft acceptance rate: 67.1%
(APIServer pid=2485494) INFO 09-07 00:12:07 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 92.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage:
 32.6%, Prefix cache hit rate: 0.0%, MM cache hit rate: 0.0%
(APIServer pid=2485494) INFO 09-07 00:12:07 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 3.59, Accepted throughput: 66.90 tokens/s, Drafted throughput: 77.40 tokens/s, Accepted: 669 token
s, Drafted: 774 tokens, Per-position acceptance rate: 0.953, 0.849, 0.791, Avg Draft acceptance rate: 86.4%
(APIServer pid=2485494) INFO 09-07 00:12:17 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 1.2 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage:
0.0%, Prefix cache hit rate: 0.0%, MM cache hit rate: 0.0%
(APIServer pid=2485494) INFO 09-07 00:12:17 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.40, Accepted throughput: 0.70 tokens/s, Drafted throughput: 1.50 tokens/s, Accepted: 7 tokens, D
rafted: 15 tokens, Per-position acceptance rate: 1.000, 0.200, 0.200, Avg Draft acceptance rate: 46.7%
(APIServer pid=2485494) INFO 09-07 00:12:27 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage:
0.0%, Prefix cache hit rate: 0.0%, MM cache hit rate: 0.0%
(APIServer pid=2485494) INFO:     127.0.0.1:47734 - "GET /v1 HTTP/1.1" 404 Not Found
(APIServer pid=2485494) INFO:     127.0.0.1:47738 - "GET /v1/models HTTP/1.1" 200 OK

the vLLM commit you patched on top of.

VLLM 0.27.1 / Flashinfer 0.6.16.post3
A --torch-backend=cu130 from a couple weeks ago (0.28.0 is current)

Further updated for 0.28.0 (latest vllm) via the separate tree.

Verified working on long routes and actually very minimal noticeable (if any—at all) errors

Got context up assuming you really want to push the barebones limits:

Three Sessions @ 90k ctx (~60tps)
export CUDA_VISIBLE_DEVICES=0
vllm serve orcarouter/Qwen3.8-27B-Uncensored-NVFP4 \
--host 0.0.0.0 \
--port 8000 \
--tensor-parallel-size 1 \
--max-model-len 92256 \
--kv-cache-dtype nvfp4 \
--gpu-memory-utilization 0.949 \
--max-num-seqs 3 \
--trust-remote-code \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--speculative-config '{"method": "mtp", "num_speculative_tokens": 3}'
Two Sessions @ 132k ctx (~60tps)
export CUDA_VISIBLE_DEVICES=0
vllm serve orcarouter/Qwen3.8-27B-Uncensored-NVFP4 \
--host 0.0.0.0 \
--port 8000 \
--tensor-parallel-size 1 \
--max-model-len 132240 \
--kv-cache-dtype nvfp4 \
--gpu-memory-utilization 0.949 \
--max-num-seqs 2 \
--trust-remote-code \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--speculative-config '{"method": "mtp", "num_speculative_tokens": 3}'
Image (memory info)

image

OrcaRouter org

will check

Upon my first 132k ctx session in our new static js harness, llm.shel.sh (reliant on Firecrawl/collapsable toolcalls), did first proper Blue Team research paper, on windows sandbox.

These are HyperV vulns (obviously high bug bounties) - Still, interesting read with what it gathered!

I just further updated my PR for the nvfp4 patch. Worth noting, on WSL in particular, VLLM's 0.29.0 release (now latest) doesn't pin the right memory in WSL. As well, the latest flashinfer can run OOM due to unbounded max jobs. I've found the sweet spot at how this quant works at 94% gpu mem. Example of the vllm serve command:

export CUDA_VISIBLE_DEVICES=0
export VLLM_USE_FLASHINFER_SAMPLER=1
export VLLM_WSL2_ENABLE_PIN_MEMORY=1
export MAX_JOBS=3
vllm serve orcarouter/Qwen3.8-27B-Uncensored-NVFP4 \
--served-model-name Qwen3.8-27B \
--host 0.0.0.0 \
--port 8000 \
--tensor-parallel-size 1 \
--max-model-len 112736 \
--kv-cache-dtype nvfp4 \
--gpu-memory-utilization 0.94 \
--max-num-seqs 2 \
--trust-remote-code \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--speculative-config '{"method": "mtp", "num_speculative_tokens": 3}'

kv modification: vllm-nvfp4-kv-sm120 PR#2

Sign up or log in to comment