logo

A new ML compiler to run 70B+ LLMs on consumer GPUs with <1% accuracy loss

Posted by marcelbuilds |17 hours ago |1 comments

marcelbuilds 17 hours ago

We've built both a compiler and model runtime that has consistently compressed large open-source models so they run on consumer-grade GPUs with no accuracy or speed drawbacks. We believe it is currently an undiscovered approach in Machine Learning, not quantization or anything researched out there.

Our progress so far has shown our compressed 30B-500B+ models would be able to run on 40+ tok/sec.

Our largest complete success has been compressing a 30B Qwen-coder model to run on 2GB of VRAM while maintaining >99% accuracy.

We feel like this is something huge, and are actively progressing to continue on with 70B, 250B, 400B, all the way to the world's current frontier open-source models so that frontier intelligence can sit on our own machines instead of trusting third party APIs.

We've released a waitlist and website to see if this is something interesting for you guys, hope you would check it out: https://www.orafrontier.com/

Thanks for reading! Let me know if anybody is interested in knowing more