Well it's not exactly saturated when OAI refused to use the harness explicitly provided by ARC-AGI. I'm not really familiar enough with the benchmark to declare whether it's a perfect measure for AGI, but I kind of doubt it is.
I think you're misunderstanding. Astra is at the top of the official ARC-AGI leaderboard, with an ARC-AGI approved harness. It's not a harness specialized for ARC-AGI. It just does the same thing the regular ChatGPT interface does: keeps conversation history across turns and compacts when it gets too long. Without the harness, it loses its entire context window every move. That's not how humans work and it's not how any real AI service works.
From what I've seen it only made people mad, not hyped, so the person that thought it was a good idea miscalculated a bit. Now waiting for Anthropic's post about their usage promo or something similar to redirect people to them.
Greentext is eh. Very formulaic, in fact very similar to the bottomless pit one, which I'd argue is better because of it's absurdity. I have to ask, did you mention the older GPT version to fable in the prompt?
Yes I've seen this before, and while the critiques are fair and high quality (and unfortunately not unique to METR) we're missing the forest for the trees here.
First of all, if you take the articles critiques and work out the implications on the METR graph, all you're doing is shifting the curve up or down, it doesn't change the fact that progress is scaling exponentially. While it is technically possible the universe could be throwing a massive pathological curveball to change the conclusion from METR data (which is we've been seeing exponential growth over the last 6 years), I think that seems very far from likely. The fact that we see the same behavior from a variety of sources over a wide variety of tasks and domains is a pretty clear indication that METR while certainly far from perfect is actually painting a consistent picture at least in terms of the rate of progress.
You can look at ECI for a summary benchmark statistic, which does NOT use METR's benchmark, and you see a similar trend. Same with SWE-bench where the task distribution is far more in domain for real world problems. It is a bummer that this METR data can't be better funded. It would probably take $1M or so to really beef it up properly which any of these labs probably have in their couch cushions.
reply