User throughput Speculative decoding
Model (tok/sec) speedup Acceptance length Acceptance prob. (draft 1 / 2)
Nemotron 3 Nano (30B) 145.8
Nemotron 3.5 Lightning (30B) 146.71.00×
+ MTP decoding 247.51.70×2.94298% / 96.2%
Qwen 3.6 (35B) 192.41.32×2.66288.5% / 77.8%
Gemma 4 (26B + 4B drafter) 216.31.48×2.89896.4% / 93.4%