↑
logo

Astra, Opus 5.5 Demonstrate Jagged Performance on Web to Robotics Tasks

Posted by loclol101 |an hour ago |2 comments

loclol101 an hour ago

Author here. We ran eight frontier models on web tasks across offline physical domain tasks from Bench2Drive, VLABench, IndEgo and Assembly101.

We expected these models to be jagged, but the shape of it surprised us. In all 15 model pairs, the lower-scoring model solves at least 3 tasks the higher-scoring one fails. We find that an update inside one model family moved the mean by -1.1pp, while flipping 36 of 177 episodes in both directions. The same was found on the four physical domains as well. Whether models would succeed or fail on a task is hard to predict before hand, since human labeled benchmark difficulty levels don’t necessarily mean the same to frontier models.

We also release the per-item results and model traces for exploration: https://huggingface.co/datasets/figai/RIDGE.

bubba123 an hour ago

thank you for sharing. Do you have plans to test this on more robust computer control or robotics tasks? It would be interesting to see how the performance scales