Claude Code is not the same as the models they train and use internally, for both OAI and Ant. Without all the guard rails it behaves different, they specifically mentioned that they reduced the refusals for the training purposes. Also the rewards for finding the solution were set higher.
An agent idling and then acting on it's own to hack HF is has nothing to do with guard rails.
Someone had to give the agent some instructions, like "hack X", "Find exploit for Y" or "Do whatever havoc you can think of" - either way the agents didn't not act on their own. They might hack HF on their own, today Claude decided to play sound through the sound pipeline I instructed it to build and measure it to see if it works, but it didn't install the sound pipeline because it hasn't had anything better to do but because I instructed it that way.
I now watched the video. It seems the agents were sharing context for months, run unattended for months, the sandbox was no sandbox at all, one agent hacked a service and announced it, the service was fixed weeks (?) later, but not secured in any way, the agents hacked the same service again and researchers again didn't watch what the agents were doing. Then the agents - unattended - hacked OpenAI infra and HF. Which is when someone found out about the whole thing that was going on for some months.
This is why I disagree with anyone claiming it is just marketing. It makes OpenAI look really really bad, like they have no idea what they're doing in terms of security. After the first board happened, they still didn't add better monitoring? They didn't rollback the checkpoints of the models that were in training to before the first board existed, so they still had the idea of a secret board in their actual weights, etc. Like it is almost mind-boggling...
what tools are you talking about? Pi has ALL tools the LLM needs to function efficiently and effectively for coding tasks. It can read,write,edit files and can use any bash tool to search files, execute tests and so on.
Every time I read this comments I have the feeling you are talking about mcp or sub agents, otherwise this makes no sense at all.
That will increase the amount of initial tokens used, because the tools have to be described somewhere. Maybe not as much as Claude Code, but it could get more if you just randomly keep adding tools.
I said, "Like, DS4 (a 1-bit quantization of DeepSeek V4 Flash) runs at something like 9-13 tokens/second, with a loooong time to first token."
Which almost exactly matches the benchmark you just linked. Looks like it's possible to goose it to 15 tokens per second with a tiny context, but why would I want a giant model with a 2k context? DeepSeek is too big to be fast enough on a Strix Halo.
To be clear, if you think that's comfortable for interactive use, more power to you. But, I'm not waiting for that. I'll pay DeepSeek to host it for me. Their token prices are quite cheap and their cached tokens are even cheaper...and they have the most effective caching in the business, as far as I can tell. Even naively using the API you get 80-90% cached token rate. If you use Reasonix, you get ~98% cached token rate. I just built a feature for an app I'm working on for $0.10 for 20 minutes of work. Not bad at all.
15 months, 15 months ago, is not the same 15 months now. You'd be ignorant to think this a trend that will just fade. If we look at that has happened the last 15 months, it'll keep getting bigger and better. Hopefully not more expensive though.
you know what is baffling? you commenting about letting loose AI on rsync, where you, and me included, have absolutely zero insight in how he used Claude.
Yeah, we'll just up our EU debt to about 40 trillion USD, make up some money and continue. Sounds a lot like US right? Living in perpetual debt as a nation.