Gemma · Benchmark run
Gemma 4 26B A4B FP8 MTP on RTX PRO 6000 Max-Q
Gemma 4 26B A4B FP8 served by vLLM reached 185.534 tok/s decode at a 4K shared repository prefix with the Gemma 4 assistant MTP5 path. The 256K-class point was fastest without MTP: 58.211 tok/s at a 261,504-token shared repository prefix, so this entry reports the best zero-failure points from the documented no-MTP baseline and MTP sweep.
Conditions
- Model
- gemma-4-26B-A4B-it-FP8-Dynamic
- Family
- Gemma
- Parameters
- 26B
- Quantization
- FP8
- Rig
- windows-blackwell-max-q
- Topology
- Single GPU
- Nodes
- 1
- GPUs
- 1
- Tensor parallelism
- N/A
- Speculative decoding
- MTP=5
- Context
- 4,096 tokens
- Latency
- 320 ms
- Peak GPU temp
- 85 °C
- Engine
- vLLM
- Runtime version
- 0.23.0
- Operating system
- Windows (WSL2)
- Driver
- NVIDIA 596.72; host CUDA 13.2; vLLM Docker image vllm/vllm-openai:latest with CUDA 13.0 and Torch 2.11.0+cu130
Topology notes
Single RTX PRO 6000 Blackwell Max-Q GPU exposed to vLLM through Docker Desktop on Windows with WSL2 GPU passthrough.
Recipe
This is a best-of-sweep result, not one continuous server configuration. MTP5 produced the fastest validated 4K through 64K points, but no-MTP was faster at 128K and the 261,504-token 256K-class bucket. A literal 262,144 shared-prefix dataset was rejected because Gemma tokenization made the input alone exceed the service window after the task suffix; the publishable 256K-class bucket is 261,504 shared-prefix tokens with a 512-token output budget. vLLM logged WSL pin_memory warnings and a warning that num_speculative_tokens greater than 1 reuses the same MTP layer multiple times.
Operator notes
The FP8 checkpoint is an LLM Compressor compressed-tensors model. vLLM inferred quantization=compressed-tensors when -Quantization was left blank, selected CutlassFP8ScaledMMLinearKernel for CompressedTensorsW8A8Fp8, and used the TRITON FP8 MoE backend. The no-MTP smoke at max_model_len 262144 and KV32G launched cleanly, and the server reported a GPU KV cache size of 1,240,740 tokens with 4.73x maximum concurrency for 262,144 tokens per request. The first full no-MTP run through a literal 262,144 shared-prefix bucket failed because the actual prompt inputs were about 262,233 to 262,245 tokens before generation, producing Bad Request responses. The corrected 261,504 shared-prefix bucket kept max input plus 512 output at 262,117 tokens. The MTP sweep used the compatible Gemma 4 assistant checkpoint. At 4K/8K, MTP1 measured 80.606/77.767 tok/s, MTP2 measured 96.418/89.497, MTP3 measured 98.742/94.416, MTP4 measured 103.810/102.994, and MTP5 measured 160.204/145.350 in the sweep. MTP6 at KV32G loaded the target and assistant but failed during KV allocation with CUDA OOM while trying to allocate 6.40 GiB. Reducing to KV24G let MTP6 and MTP7 run, but they were slower: MTP6 measured 92.475/97.061 and MTP7 measured 101.451/99.483. The full MTP5 ladder completed every corrected bucket with zero failed requests: 185.534 tok/s at 4K, 182.276 at 8K, 144.629 at 16K, 109.451 at 32K, 72.391 at 64K, 41.150 at 128K, and 15.642 at 261,504. Draft acceptance stayed around 49 to 56 percent with mean accepted length around 3.46 to 3.78 tokens, but the high-context overhead still made MTP slower above 64K. The full no-MTP ladder also completed every corrected bucket with zero failed requests: 98.885 tok/s at 4K, 98.072 at 8K, 88.474 at 16K, 77.827 at 32K, 66.931 at 64K, 52.900 at 128K, and 58.211 at 261,504. The final series therefore reports best validated zero-failure points by context: MTP5 for 4K, 8K, 16K, 32K, and 64K; no-MTP for 128K and 261,504. The practical conclusion is that Gemma 4 FP8 benefits heavily from MTP in the low and mid context bands on this rig, but MTP should be gated off for the deepest buckets.
Text generation sweep
Decode throughput at each measured context and concurrency, not just the headline number. Hover a point for its value, or toggle a series in the legend.