↑
logo

Local AI is real now. And it's blowing my mind

Posted by tylermhall |3 hours ago |1 comments

colingauvin 3 hours ago

The author's configuration is..... Interesting.

I don't disagree with the title. This spring folks were saying "I use Qwen 3.6 27B (or 35 A3B) as a daily driver". They either had low standards or were lying. The models were good enough technically, they capable enough, that you could see the potential, but the coherence was really not great compared to frontier. They'd lose track of what they were doing, mix things up, etc. Technical ability definitely arrived before coherence.

Then DeepSeek 4 Flash came out and suddenly it was like... Ok. You can run a local model that keeps coherence, and is decent technically. Then we got a rash of models that are both good technically and coherent: DS4 Flash 0731, Qwen 3.8 Flash Next, Qwen 3.8 27B, GLM 5.3 Flash, Hy4, MiMo 2.6 Flash... It just has no stopped. We have definitely crossed the threshold from interesting to very usable.

I like the solar bit. Not sure why the author is using a 4 model setup. My tests with Qwen 3.8 have convinced me it's completely fine general purpose, and dedicated small/vision/etc are useless. I can get 110 TPS single stream on 2x Sparks, with 5M total context, concurrent I think some folks are hitting 200-250 TPS. There's no reason to use anything else right now, it's really remarkable. For an Opus 4.6/4.7 quality model.

I am eagerly awaiting the release of Qwen 4 Flash.