himata4113 5 hours ago
This is napkin math since I'm mostly just extrapolating from glm 5.2 by assuming it's twice as heavy to serve in every single measurement, but I believe you can easily achieve 2500tok/s aggregate compared to 4500tok/s and up to 8000tok/s for glm5.2.
with nvidia r100 you are likely going to be able to push that number even higher while the cost of hardware appears to be relatively the same, so far I am seeing 21% premium from supermicro which is twice as fast and has nearly twice the vram.
sroerick 5 hours ago
coder543 5 hours ago
The model weights are supposed to release tomorrow.
Over the next several weeks, I would expect competition among open weight providers to drive down the cost, as I've seen happen with other open weight model releases.
jszymborski 5 hours ago
ronsor 5 hours ago
cloudie78 5 hours ago
ofjcihen 5 hours ago
Additionally, I fully expect the frontier labs to continue increasing prices to meet the profit margins they need to to continue existing.
CamperBob2 5 hours ago
4 hours ago
Comment deletedTGS-AI-Chain 5 hours ago
seydor 5 hours ago
SwellJoe 5 hours ago