↑
logo

>1B tokens/minute/GPU by combining query planner and inference engine

Posted by charles_irl |an hour ago |1 comments

sharktheone 15 minutes ago

First of all, this is VERY impressive. I wonder how this might effect the output quality. I wrote my own AI execution engines in the past and from that I can tell that it is really hard to get right and event small mistakes can make the model way dumber than expected.

While qwen3-4b is smarter than I remembered it, newer models seem to be way better, K2 Horizon 3.7B for example. Is there a reason why that qwen model was choosen?

My last question would be if the same optimizations could be made to a Jev-like model. I mean Jev is already fast and probably has a high throughput per H100. So could we maybe get to >2B tokens/minute/GPU with such a model?