OpenAI's release explicitly says No. But then also caveats that with "we cannot rule out that de-identified data derived from their usage of our products" impacted things.
What's most striking to me, and what may or may not be true, is the "we cannot rule out" bit.
"We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)." https://openai.com/index/navier-stokes-solution/
NOTE: there are a couple duped threads around this. i replied on a different one first before seeing this one
I don't understand how they can say it's unlikely. It's objectively true that they train on de-identified user data (https://openai.com/policies/how-your-data-is-used-to-improve...), and objectively true that they encourage users to submit such data (the setting is on by default). Since it's de-identified I can believe that they can't give a straight yes or no answer here, but it seems more likely than not that at least some amount of his usage became training data. It takes an unusual level of awareness and effort for a user to ensure that all usage is opted out.
That was also striking to me too, the 'we cannot rule out'.
I also do not believe it. They can, trivially, rule out their model being trained on if
More to the point -- of _course_ they know what data their model is trained on and the lineage of that data, even if it ended up being anonymized and they cannot identify which precise user.
OpenAI's release explicitly says No. But then also caveats that with "we cannot rule out that de-identified data derived from their usage of our products" impacted things.
What's most striking to me, and what may or may not be true, is the "we cannot rule out" bit.
"We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)."
https://openai.com/index/navier-stokes-solution/
agree with this. i don't see the point of the terrarium if the fly never moves or really changes at all, other than seeing to twitch faster when sugar is on.
Such a fun story included here. It'd be great if this was also in a native web-app ebook too (i know it's harder to charge $ effectively, but most people end up wanting the physical or epub (for their e-reader) anyway if they like it.
CI/CD, rollback strategy, blue/green deployments, testing regions, automated testing suites... you need all of this and more.. because code reviews just aren't going to be happening soon. luckily, these LLMs can help in building that infrastructure needed to reliably ship changes quickly
There are a lot of systems at play here, and those things don't change overnight. I think certainly we'll look back in 10 years and clearly see this was the start of a revolution, but we seem to have a distorted view of what a revolution is. The industrial revolution didn't happen overnight, and many folks during the time didn't have any concept of what was actually happening.
Classic hindsight bias.. And it's funny that Sam is saying this.. just a few months ago i saw his interview where he claimed to be very good at seeing the future and predicting what will happen
Love the motivation behind this, and I think the execution is really quite nice. I really enjoy these little learning games/apps -- i think it's a great way to be using AI for development
"We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)." https://openai.com/index/navier-stokes-solution/
NOTE: there are a couple duped threads around this. i replied on a different one first before seeing this one
reply