For instance, I would see Baesten claim 5 commited tokens per round with MTP heads, and I would consistently see barely 2 on WildChat
when I would re-run the test even copying over everything they are doing.
Here they claim they hit 280 tps https://www.baseten.co/blog/how-we-built-the-worlds-fastest-api-for-glm-52/ but their open router is at 61
With GLM 5.2