logo

I checked 30 frontier model cards. Here are the benchmarks labs report

Posted by ktwu01 |5 hours ago |7 comments

mkagenius 5 hours ago[1 more]

> What the two layers say

> Stated findings

> Findings derived from two curated layers: which model cards mention each benchmark, and which scores could be read verbatim from those documents. Each finding names the evidence behind it.

This is 100% AI generated but the problem is it's difficult to understand - what layers is it talking about, is it the llm model layers or what.

jrflo 5 hours ago

As an overall pro-AI person... Please don't let it do the entire UI design for your site. 99% of this site is completely useless.

claiir 5 hours ago

This text on this page is so aggressively LLM-written (Claude) I am struggling to understand what I am even looking at.

ttul 5 hours ago

I am waiting with bated breath to read, “load-bearing” somewhere… The latest models are very capable, but sometimes they seem to get so deep in the details that they lose the overall plot.

What the hell is the point of this page? Can you put in a single bit of human prose explaining why it exists and what we are supposed to learn?

bilbo-b-baggins 5 hours ago

Broken on mobile safari

augment_me 5 hours ago

Good idea, would be interesting to cross-examine the benchmarks, but the page information is completely obscured by the AI slop. The benchmarks comparison and should start immediately instead of having random completely arbitrary complex headers and labels

ktwu01 5 hours ago

[flagged]