My guess based purely off of vibes from previous models is that boosting the frontier math ability of a model is not that difficult.
Most models trained for general use are ingrained with certain tendencies that are usually very useful like "if you're stuck and bashing your head against the wall stop and tell the user". You generally don't want Claude Code to go off and work for weeks on something when if it had just asked for help you could've clarified or provided more information or just picked a different approach.
When you're solving extremely difficult math problems though you generally do want a model to be more persistent and keep trying even when the model can't clearly see a way forward. OpenAI appears to have done this with lots of previous models. The model they trained for the IMO competition seems to have been an RL maxxed version since they noted that while it did the math it couldn't write up its results on its own and just produced CoT [0]. The capabilities are in there lying dormant, you just to need to RL max the model to ruthlessly pursue the goal at all costs which destroys general use but improves frontier math.
We've also seen hints from OpenAI at least that they seem to train more persistent versions of all their models [1].
Also, Astra probably completed training at least one to two months before the public release so it's not like they only had a week to whip this version up.
Exactly, WFH and home schooling both require you to proactively seek out relationships. A lot of people haven’t developed this habit/skill, and without school or work to provide social time they don’t really develop relationships. Losing that “forcing function” means ripping off the bandaid of how few real relationships they actually have.
There have personally been times in my life where I’ve lost that bandaid (workplace, academic extracurricular activity, etc.) and thankfully I’ve usually been able to respond by realizing that I had a problem and proactively doing something about it.
That said, your reasoning is probably still correct. There's no placebo group, and a lot can change over a 2-year follow-up, with participants aged 7-17. Maybe they went to speech therapy or just matured and learned more coping behaviors. The 2019 followup also notes that 12 of the 18 participants made other diet and medication adjustments. They claim the adjustments were minor but that's still more noise, and it doesn't account for unreported social/environmental changes.
> There have been so many small scale trials showing amazing autism improvements that failed to replicate in larger, better controlled trials. I wouldn’t get excited yet.
Unfortunately yeah, it's unlikely anything exciting will replicate in a larger RCT other than maybe the gut biome improvement since that seems directly mechanistic, but that's just a gut feeling.
It seems like the purpose is to have the law and all the paperwork set up as a precaution for the future. Sure, right now it’s all voluntary and just rubber stamping, but if in the future they need to do something like Ukraine and lock down travel for military aged men, it’s much easier to flip a switch and start denying travel permits rather than having to set up and fund an entirely new system for requiring travel permits.
If it were to get closer to war (i.e., Spannungsfall, let alone the Verteidigungsfall) a set of laws would unlock that allow control of various areas of life and the economy anyway.
Yes, iPads (at least at my university) are incredibly common. I would guess they’re at least on-par with paper. So many people swear by Goodnotes because you get all the benefits of handwriting your notes without giving up the niceties of search-ability, auto correct, etc.
I don’t know anyone who uses any other tablet besides an iPad, they’ve basically conquered the market.
This is a great point. For example, near where I live there’s a massive Google cloud warehouse out in the middle of a field next to the highway. Inside of that warehouse there’s a separate section for servers belonging to the US government that can benefit from all the electricity contracts Google has negotiated, the physical security and fences that Google has set up, and the fiber optic cables they’ve laid.
It’s the best of both worlds, they get the decades of research Google has put into systems engineering and fault tolerance while retaining the security of having their own servers.
It’s incredible how much time and political maneuvering it took Sam Altman to get to this point. He took on the entire board and research scientists for every major department and somehow came out the winner. This reads more like an announcement of victory than anything else. It means Sam won. He’s going to do away with the non-profit charade and accept the billions in investment to abandon the vision of AGI for everyone and become a commercial AI company.
This is the mini existential crisis I have randomly. The attack area for a modern IT computer is mind bogglingly massive. Computers are pulling and executing code from a vast array of “trusted” sources without a sandbox. If any one of those “trusted” sources are compromised (package managers, cdns, OS updates, security software updates, just app updates in general, even specific utilities like xz) then you’re absolutely screwed.
It’s hard not to be a little nihilistic about security.
The more important thing would seem to be what actually led to the immigration in the first place. In Cuba’s case it seems like widespread corruption and wholesale mismanagement by the government. This particular case doesn’t seem like an inevitable result of globalization but the natural reaction to a terrible government.
Most models trained for general use are ingrained with certain tendencies that are usually very useful like "if you're stuck and bashing your head against the wall stop and tell the user". You generally don't want Claude Code to go off and work for weeks on something when if it had just asked for help you could've clarified or provided more information or just picked a different approach.
When you're solving extremely difficult math problems though you generally do want a model to be more persistent and keep trying even when the model can't clearly see a way forward. OpenAI appears to have done this with lots of previous models. The model they trained for the IMO competition seems to have been an RL maxxed version since they noted that while it did the math it couldn't write up its results on its own and just produced CoT [0]. The capabilities are in there lying dormant, you just to need to RL max the model to ruthlessly pursue the goal at all costs which destroys general use but improves frontier math.
We've also seen hints from OpenAI at least that they seem to train more persistent versions of all their models [1].
Also, Astra probably completed training at least one to two months before the public release so it's not like they only had a week to whip this version up.
[0]: https://x.com/OpenAI/status/1946594933470900631 [1]: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
reply