| User throughput | Speculative decoding | |||
|---|---|---|---|---|
| Model | (tok/sec) | speedup | Acceptance length | Acceptance prob. (draft 1 / 2) |
| Nemotron 3 Nano (30B) | 145.8 | – | – | – |
| Nemotron 3.5 Lightning (30B) | 146.7 | 1.00× | – | – |
| + MTP decoding | 247.5 | 1.70× | 2.942 | 98% / 96.2% |
| Qwen 3.6 (35B) | 192.4 | 1.32× | 2.662 | 88.5% / 77.8% |
| Gemma 4 (26B + 4B drafter) | 216.3 | 1.48× | 2.898 | 96.4% / 93.4% |