krisoft an hour ago
Even absence of thinking this through you would think that some frogs will be from the front, some from the side. Just by chance. And yet all appears to go for the harder pose.
hn_throwaway_99 4 hours ago
Small side note, the first gemini-2.5-pro one totally reminded me of some sad faced meme or Pepe the frog from somewhere. Anyone know what I'm referring to, tried to find it.
thebigship 4 hours ago
also my favorite SVG was def the google/gemini-3.6-flash
edit: ok better now I think
wren6991 4 hours ago
gemini-3.6-flash runs 2 and 3 responded best to the royal portrait context.
evan_ 3 hours ago
getnormality 4 hours ago
ianberdin 3 hours ago
riazrizvi 2 hours ago
ricardobeat 3 hours ago
rush86999 3 hours ago
That's a pretty good benchmark
dehrmann 4 hours ago
linksnapzz 4 hours ago
gerdesj 3 hours ago
leumon 4 hours ago
MiroslavPokorny 3 hours ago
k1e 2 hours ago
throwuxiytayq 2 hours ago
epolanski 3 hours ago
Would've wanted to see also DS4 flash.
csomar 2 hours ago
I also did a timeline from 4.7 to 5.2: https://codeinput.com/s/7oK2IIA7qRO The improvements in models looks much less impressive with this test.
epolanski 2 hours ago
troupo 4 hours ago
Also try a fantasy archer with a proper bow who is not brooding, sitting in a fantasy wood :)
AlienRobot 2 hours ago
thebigship 5 hours ago
Seven of fourteen models silently imported royalty into a prompt that named only an anatomical feature. Two of them knew they were extrapolating ("because Habsburg") and did it anyway.
Mistral returned byte-identical output across separate calls.
Gemini narrates its work in 65 comments; Llama says nothing.
If you're deciding which model to trust with instructions, "how much does it embellish beyond what I asked" and "does it behave deterministically" are directly practical questions.
sixtyj 4 hours ago
kindawinda 3 hours ago
NemoNobody 2 hours ago