Qwen · Benchmark run
Qwen3.6 35B A3B FP8 on RTX PRO 6000 Max-Q
Qwen3.6 35B A3B FP8 served by vLLM reached 192.414 tok/s decode at a 4K shared repository prefix and completed the 261,504-token 256K-class bucket at 62.677 tok/s. The best public series is a context-gated result: MTP5 wins at 4K, MTP2 wins at 8K and 16K, and no-MTP wins from 32K through the deepest context.
Conditions
- Model
- Qwen3.6-35B-A3B-FP8
- Family
- Qwen
- Parameters
- 35B
- Quantization
- FP8
- Rig
- windows-blackwell-workstation
- Topology
- Single GPU
- Nodes
- 1
- GPUs
- 1
- Tensor parallelism
- N/A
- Speculative decoding
- MTP=5
- Context
- 4,096 tokens
- Latency
- 128 ms
- Peak GPU temp
- 80 °C
- Engine
- vLLM
- Runtime version
- 0.23.0
- Operating system
- Windows (WSL2)
- Driver
- NVIDIA 596.72; vLLM Docker image vllm/vllm-openai:latest with CUDA 13.0 and Torch 2.11.0+cu130
- Model source
- https://huggingface.co/Qwen/Qwen3.6-35B-A3B-FP8
Topology notes
Single RTX PRO 6000 Blackwell Max Q GPU exposed to vLLM through Docker Desktop on Windows with WSL2 GPU passthrough. Tensor parallelism was 1, and every measured point used one concurrent request.
Recipe
This is a best-of-sweep result, not one continuous server configuration. The Qwen3.6 model card recommends qwen3_next_mtp with two speculative tokens, but this workload benefited from MTP5 at 4K, MTP2 at 8K and 16K, and no speculative decoding above that. Deeper MTP settings regressed at high context.
Operator notes
The Qwen3.6 35B A3B FP8 model card describes a 35B total, 3B active MoE with 262,144 native context, 40 language layers, 256 experts, 8 routed experts plus one shared expert, fine-grained FP8 block quantization, and MTP trained with multiple steps. The vLLM section recommends vLLM 0.19 or newer, language-model-only serving for text-only workloads, Qwen3 parser flags, and qwen3_next_mtp for MTP. The 128-token smoke established that speculation helps only in a narrow range. No-MTP measured 199.468 tok/s at 4K and 198.385 at 8K. MTP2 measured 201.449 and 174.895. MTP4 measured 245.312 and 222.010. MTP5 measured 256.204 and 234.324. MTP6 regressed to 143.737 and 141.790. The full 512-token runs changed the conclusion. MTP5 stayed best at 4K with 192.414 tok/s but dropped to 110.538 at 8K and 104.014 at 16K, then fell behind no-MTP at deeper buckets. MTP2 was best at 8K and 16K with 159.885 and 127.182 tok/s, but it degraded badly at high context, reaching only 11.477 tok/s at 261,504. MTP3 and MTP4 did not win any public row. The no-MTP full ladder became the high-context recipe. It completed every bucket with zero failures and measured 115.571 tok/s at 4K, 113.345 at 8K, 112.531 at 16K, 102.382 at 32K, 91.734 at 64K, 79.785 at 128K, and 62.677 at 261,504. The public series therefore reports MTP5 for 4K, MTP2 for 8K and 16K, and no-MTP for 32K through 261,504. Telemetry stayed inside the single-GPU envelope. Across the relevant runs, peak VRAM was about 63.78 GiB, peak power about 305 W, and peak temperature about 80 C. The MTP2 full result reported acceptance around 80.877 percent at 4K, 80.930 at 8K, and 82.584 at 16K, with mean accepted lengths around 2.62 to 2.65 tokens. The MTP5 4K winner reported about 59.948 percent acceptance and a mean accepted length of 3.997 tokens. This model is a personal favorite for quick token generation and exceptional reasoning capabilities and tool calling for it's size, and it's a great daily model for DGX Spark owners.
Text generation sweep
Decode throughput at each measured context and concurrency, not just the headline number. Hover a point for its value, or toggle a series in the legend.