jampa 37 minutes ago
They understand all the rules and best practices, they can (sometimes) spot a bad idea in a floor plan, they can describe a good floor plan.
But ask them to make one, even if you give it every detail (even a "node graph" of rooms), they will still output nonsense. Same for text and image models.
Floor plans should be the new Pelican Benchmark.
ghostpepper 2 hours ago
nhl toronto scores nhl hockey toronto scores "nhl hockey" toronto score today nhl "hockey score toronto" "hockey" who won toronto
etc.
Somehow being good at semantic search makes them bad at keyword search, for whatever reason.
mojuba 28 minutes ago
We tend to think that the AI has some sort of self-knowledge and should be good at designing prompts for itself but it's really not.
Been struggling with a task that heavily depended on prompts, ended up rewriting all my prompts from scratch in my own words, and it finally worked. Then every time I ask Claude to fix something in the prompts, it invariably makes it worse.
A very strange phenomenon that can probably be explained by the quality of prompt design advice that made it to the training dataset. Bottomline, all the prompt design advice that you can find on the internet is really not great.
tartoran 2 hours ago
sandcat_ 2 hours ago
Anno 1800 was a recent one I had trouble with, using Claude Opus. Completely made up game mechanics. Rainbow Six Siege, too.
elliotto an hour ago
I asked a bot why it thought it wasn't funny once, and it told me it has been trained to avoid being misinterpreted or offensive, so anything that might be considered edgy would have been RLHF'd out of it. I thought this was very introspective.
TiccyRobby 2 hours ago
kanzure 2 hours ago
jstrieb 38 minutes ago
On math or programming problems, they are overfit to solving the entire thing end to end (presumably for benchmarks). I have had very poor results asking for pointers and hints that don't give away key insights. This has been the case across models I have tested.
An architecture with a "judge" that gates responses and ensures a lack of spoilers would probably work better. But this is a simple thing that they keep messing up.
da-x 34 minutes ago
humanrebar 2 hours ago
znnajdla an hour ago
spike021 2 hours ago
I've had some luck on the web app side if I use playwright or similar for the model to interact with but still far from efficient.
newsomix9xl 2 hours ago
sghiassy 2 hours ago
More of an image model than a LLM model tho
alexandra_au an hour ago
dowonseo 38 minutes ago
dSebastien 21 minutes ago
dhruv3006 2 hours ago
TZubiri 2 hours ago
SubiculumCode 2 hours ago
dorianpruski 2 hours ago
blinkbat 2 hours ago
Oh, you said simple. Speaking like a human
ipaddr 25 minutes ago
maxsavin 2 hours ago
eli 2 hours ago
Tasks it writes are typically too easy but also it utterly fails to see how a different model might misunderstand a vague part of the prompt.
flippy_flops 2 hours ago
respectattentio 2 hours ago
newsomix9xl 2 hours ago
rufi 2 hours ago
shoopadoop 2 hours ago
You wouldn't tolerate this kind of duplicity from a human coworker, but AI is so fast and efficient at lying, so it's OK.
bpodgursky 2 hours ago
Conol_ai 2 hours ago
senectus1 2 hours ago