Qwen · Benchmark run
Qwen3.5 122B A10B Q4_K_M GGUF on RTX PRO 6000 Max-Q
Qwen3.5 122B A10B Q4_K_M GGUF served by llama.cpp reached 144.539 tok/s decode at a 4K shared repository prefix and kept a zero-failure ladder through the 261,504-token 256K-class bucket, where it measured 93.608 tok/s. The winning recipe used llama.cpp native MTP with draft depth 2, F16 KV cache, flash attention, a 262,144-token service window, and a smaller 512-token ubatch.
Conditions
- Model
- Qwen3.5-122B-A10B-MTP-GGUF
- Family
- Qwen
- Parameters
- 122B
- Quantization
- Q4_K_M
- Rig
- windows-blackwell-max-q
- Topology
- Single GPU
- Nodes
- 1
- GPUs
- 1
- Tensor parallelism
- N/A
- Speculative decoding
- MTP=2
- Context
- 4,096 tokens
- Latency
- 35 ms
- Peak GPU temp
- 86 °C
- Engine
- llama.cpp
- Runtime version
- b9775 (be4a6a63e)
- Operating system
- Windows
- Driver
- NVIDIA 596.72; llama.cpp CUDA 13.3 build
Topology notes
Single RTX PRO 6000 Blackwell Max-Q GPU running the native Windows CUDA build of llama.cpp.
Recipe
The 8K shared-prefix bucket needs a service context larger than 8192 because the task suffix raises the full prompt above the shared-prefix length. This entry reports the best validated zero-failure llama.cpp configuration from the sweep; no-MTP, MTP4 at the original ubatch, MTP6, and Q8 KV cache were all slower. The 86 C peak temperature indicates sustained thermals were near the top of the workstation envelope.
Operator notes
The Qwen3.5 122B GGUF model card points operators toward llama.cpp and MTP. The local tokenizer metadata reports a qwen3_5_moe model with 262,144 maximum positions, 48 layers, 256 experts, 8 experts per token, and one MTP hidden layer. The local llama.cpp b9775 Windows CUDA build exposes draft MTP, flash attention, KV cache type controls, host polling controls, and batch/ubatch controls, so the run used those knobs instead of a first-load baseline. The first smoke exposed a context-window detail: a server launched at c8192 can complete the 4K shared-prefix bucket but rejects the 8K bucket because the suffix pushes full prompt length to about 8,288 tokens. The smoke was rerun at c16384 before moving to the publishable c262144 ladder. MTP helped substantially, but more draft tokens were not always faster. With 128-token smoke outputs at c16384, no-MTP measured about 106 tok/s, MTP2 measured about 140 tok/s, MTP3 and MTP4 were close at low context, and MTP6 regressed. On the full 512-token ladder at b2048 ub256, MTP4 beat no-MTP through most buckets but MTP2 was competitive at the deepest context. The final zero-failure MTP2 b2048 ub512 ladder measured 144.539 tok/s at 4K, 142.273 at 8K, 138.275 at 16K, 133.575 at 32K, 121.718 at 64K, 109.930 at 128K, and 93.608 at 261,504 shared-prefix tokens. The same recipe peaked around 83.00 GiB VRAM, 308 W, and 86 C. Draft acceptance stayed useful across the run. Per-context result files show about 73 to 77 percent accepted draft tokens, and the cumulative llama.cpp server log ended with 10,019 generated draft tokens, 8,411 accepted draft tokens, 20,017 generated tokens, 15,003 accepted tokens, mean accepted length 2.50, and per-position acceptance of roughly 0.840 and 0.658. The losing configs are part of the conclusion. The no-MTP c262144 b2048 ub256 full ladder topped out at 102.829 tok/s at 4K and 66.536 tok/s at 261,504. MTP4 b2048 ub256 reached 128.419 at 4K and 79.533 at 261,504. MTP2 b2048 ub256 reached 124.709 at 4K and 81.853 at 261,504. Q8 KV cache reduced memory slightly but was slower at the low-context tune points, measuring 130.639 at 4K and 129.936 at 8K, so F16 KV remained the publishable recipe.
Text generation sweep
Decode throughput at each measured context and concurrency, not just the headline number. Hover a point for its value, or toggle a series in the legend.