Why “mtp” is being discussed
The most recent Hacker News stories and comments contributing to this topic's mentions.
It really does get it, because MTP is usually run at "3 token" depth. It's pretty shocking to watch
by girvo · Sep 21, 2026
I think you’re missing that MTP can predict more than 1 token in advance.
by medvezhenok · Sep 21, 2026
On my m5 max 27b model does 75tps on 256k ctx and starts at 80 on the 8k ctx when you add https://huggingface.co/collections/z-lab/dflash-2 to it. So ye…
by alex7o · Sep 21, 2026
A 5090 has a 1.79TB/s memory bandwidth. Qwen 3.8 27B NVFP4 is 22GB. You cannot generate tokens faster than the weights can traverse the GPU memory, so that makes max genera…
by Eisenstein · Sep 21, 2026
Well I tested it on 7900XTX with the same prompts and their draft model gave me about 30t/s. Their own snippet of code with regular MTP model gave me 60t/s. Also model…
by npodbielski · Sep 18, 2026
I’m on a strix halo @ GPU-5 with MTP and I get 600 prefill and 30 TG which pushes it into a very usable range. The odd thing is that Dflash2 is really slow for me, like sub 10 TG.
by syntaxing · Sep 18, 2026
Interest
Proportion of Hacker News items mentioning "mtp" over time.
Mentions
Total number of Hacker News items mentioning "mtp" over time.