Hacker Newsnew | past | comments | ask | show | jobs | submit | haldujai's commentslogin

While not the only ones, the NYT notoriously pushed a lot of false and poorly verified information from questionable sources as fact in an effort to be the first to publish. They subsequently issued a mea culpa.

https://www.nytimes.com/2004/05/26/world/from-the-editors-th...


They did the same thing with the Israeli holocaust

Not explicitly, but hacking HF is within the scope of “solve this problem at all costs” + no/poor guardrails + infinite budget + unsolvable problem.

Sounds like you are thinking they just need Asimov’s laws. But I think the point is, this can easily be weaponized by somebody with the willpower to do so.

No, just that “What if I just cheat to get the treat” is also something my lovable Boxer does when I hide rewards as well.

I suppose her energy is limited but even if infinite, and with caution, I estimate the chances of a mass extinction event due to Pickle taking over the world at < 1%.


> The US produces about 70% more electricity per capita.

And consumers use 4x as much per capita. Industrial generation per capita China comes out ~2x

> industrial electricity prices in China are roughly 34% higher than in the US

For which industrial customer and where? Chinese compute hubs are on par to slightly cheaper on pure electricity costs.

Conversely the US makes it more expensive with interconnect and upgrade fees as well as hefty take or pay contracts.

A 1GW datacenter in VA for example would add 5-10c kWh and a 12 year take or pay deal


1. I don’t think that’s a very strong argument. OpenAI and Anthropic don’t buy the vast majority of GPUs they use they rent capacity.

Nvidia could just the same rent those GPUs out for inference and actually have way better margins than they do right now. Antitrust and putting all your eggs in one basket are why they don’t, similar to TSMC.

2. Neither do AI labs. See Anthropic buying TPUs, deploying with AMD. OpenAI on Maia, Cerebras, their own wafers.


It’s not about whether or not Nvidia will be able to sell hardware to these providers - it’s about literally killing companies they are financially invested in.

Why would you invest money in a company, and then enter the market to compete with them?


About the same, 5-10, when you consider major (aka frontier) airlines.

Actually not a bad comparison. Both burn massive amounts of up front capital to protect an oligopoly in the hopes their commodity product eventually pays off.


1. The same isn’t necessarily true of the rest of the hardware stack which may be reused between accelerator generations.

2. You’re missing the “New Nvidia chips 10x B200, compute requirement grows less than 10*software improvements YoY -> buy less Nvidia.” Valuations are based on forward projections (>1T annual for NVDA) which can be revised down leading to a drop in valuation.

> If Amazon doesn't buy but Microsoft does

The big 3 all have their own proprietary accelerators. Meta is buying TPUs as well for now.

I would bet Nvidia’s major customers in 2 years are neoclouds and it seems that Jensen is making the same bet.


1. So this makes Burry’s argument even less convincing since those auxiliary hardware can last longer.

2. Jevons Paradox. More efficiency should lead to bigger models, faster inference, and more total tokens.

3. By all accounts, Trainium and Maia and Meta’s internal chip are struggling to keep up with Nvidia. That’s why they order as many Nvidia chips as possible. They’re not giving up but it isn’t as easy as buying stock Arm cores and taking them to TSMC.

Neoclouds may very well be Nvidia’s biggest customers and this probably what Nvidia wants.


1. Not really, current valuations are priced for persistent 80%+ margins based on spot. If auxiliary hardware lasts longer (I.e. next gen GPU reusing the same shell) then that reduces supply pressure and spot prices.

2. Jevon’s paradox is about total consumption, not margins. Valuations are about margins (and their projections). Many coal mine owners went bust despite increased total coal consumption.

3. Source? Gemini for example is 70% on TPU. I have yet to see data on Maia-300 beyond Microsoft PR. Remember it doesn’t have to be better it has to be more cost efficient. The overwhelming majority of inference spend does not care if token output is 20% slower if it is 50% cheaper.

> Neoclouds may very well be Nvidia’s biggest customers and this probably what Nvidia wants.

What Nvidia needs. Whether neoclouds can stay competitive vs hyperscalers paying Nvidia tax is far from clear, particularly when inference margins compress.


1. The whole Burry argument is that AI hardware becomes obsolete faster. If aux hardware can be reused, that works against the argument.

2. Total consumption drives more demand for the already supply constrained hardware. Can AI hardware market go bust? Sure it can. But being early is the same as being wrong in the investment market. When do you predict the bust to be?

3. Google, Amazon, Microsoft, Meta are all buying as many Nvidia GPUs as they possibly can. The biggest tell on how Nvidia is doing is that their share in inference has increased despite the increase in competition: https://archive.md/CKP0N. So while competition is getting bigger and bigger because the overall pie is getting exponentially bigger, Nvidia's growth is still higher than average.

Some other sources:

https://www.businessinsider.com/amazon-nvidia-aws-ai-chip-do...

https://www.businessinsider.com/startups-amazon-ai-chips-les...


> 1. The whole Burry argument is that AI hardware becomes obsolete faster. If aux hardware can be reused, that works against the argument.

Burry’s main argument is depreciation is being understated and the capex vintages will not be paid off before they are essentially useless. This can happen whether or not aux is reused.

> Total consumption drives more demand for the already supply constrained hardware.

Demand is the wrong metric.

Only number that matters is whether AI attributable revenue will be sufficient to pay back enough of each successive capex vintage (e.g. 750B this year, 1T next year, 1.2T in 2028) so that hyperscalers and neoclouds can either self-fund or continue to issue debt as bond markets are already straining and tax-payer backed sovereign debt is providing a high baseline. Otherwise they downgrade capex projections and the bubble pops.

Expensive compute needs expensive inference to justify 30-40B/year/GW of compute. There are many reasons why frontier API pricing which is what the industry is based on may not persist. It is also almost certainly the case that 2026 is the worst year of supply and demand mismatch to allow for 80%+ margins. HBF next year has the potential to single handedly pop the DRAM spot bubble.

> Can AI hardware market go bust? Sure it can.

This is the bear thesis. It is not that AI will crash or be useless.

> But being early is the same as being wrong in the investment market. When do you predict the bust to be?

Q4 27-Q2 28 is when the bill becomes due at the latest. There are sufficient financial levers left to buy time without returns until then.

> Google, Amazon, Microsoft, Meta are all buying as many Nvidia GPUs as they possibly can.

All of these companies have rock solid revenue streams and can easily swallow 500B of capex devaluation over time. Their buying of Nvidia today is not necessarily the indicator you are implying as there are strong competitive reasons to make the game more expensive for everyone else.


  Burry’s main argument is depreciation is being understated and the capex vintages will not be paid off before they are essentially useless. This can happen whether or not aux is reused.
And why does he think depreciation is understated? It is because he thinks newer Nvidia GPUs will make older ones obsolete faster. Hence, my entire post.

The rest of your argument centers around whether AI growth will meet the cap ex expenses. I don't see anything new in it.

  HBF next year has the potential to single handedly pop the DRAM spot bubble.
I'll believe it when I see it. Jevons paradox will apply here again in my opinion. HBF does not replace HBM.


A deal with a company Google has a share of, announced a week before their IPO, with very non-committal terms and ramp period protections delivered in one large block on short term notice priced likely at the high end of what Google charges for A4X instances anyway.

I don’t think this reflects desperation as much as strategy.


I dunno if it's a sound strategy that involves repeatedly telling investors [1] and employees [2] over multiple quarters that you are desperate for compute, including leaving a triple-digit billion backlog on the table [3], and then spending so much on CapEx that you have your first negative cash flow quarter ever and taking the inevitable hit to the stock [4], while turning away a large paying customer (who also happen to be a competitor) [5] ;-)

[1] https://www.mindstudio.ai/blog/sundar-pichai-google-compute-...

[2] https://www.cnbc.com/2025/11/21/google-must-double-ai-servin...

[3] https://www.bloomberg.com/news/articles/2026-07-22/google-sa...

[4] https://arstechnica.com/google/2026/07/google-just-had-its-f...

[5] https://thenextweb.com/news/google-caps-meta-gemini-compute-...


Backlog meaning RPO over 5 years. It’s not as if they would be able to collect 250B today from OpenAI and Anthropic if they were to have that compute.

In any case, my point is that the SpaceX deal specifically likely has ulterior motives.


The RPO can be anywhere from 3 - 6 years, sure, but even on an annual basis that’s like a hundred billion now. It was already in the double-digit billions since before AI took off and has only been spiking since then, which tells us 1) it’s been huge for 3+ years, and 2) it’s still growing faster than they can collect it. This matches what all the other hyperscalers are doing.

My point is that an ulterior motive is not necessary to assume when all their actions and statements point to them being severely crunched for compute.

I mean sure, if they had a choice between say, CoreWeave and SpaceX, they’d choose the latter for the nice bump to SpaceX’s financials and their stake… but not just for that, not when it contributes to their cash flow turning negative and their own stock taking a hit.


It is very suspicious when they’re paying for GB300 more than they turn around and charge those out as a4x instances, it was announced a week before IPO and they have a 90 day exit option.

> The RPO can be anywhere from 3 - 6 years, sure, but even on an annual basis that’s like a hundred billion now.

OpenAI and Anthropic deals are mostly 5 year and Anthropic’s starts in 2027, accordingly these RPOs are sized for projected compute needs and run rate in 2027 not today.

The only way either lab could pay 1 year of RPOs today (~80B for anthropic and ~150B for OpenAI) is with a lot more debt or circular financing, the former of which is difficult in this market.

Everyone spending crazy money on capex right now says they have a crushing backlog and need more compute to protect their share price. I highly doubt the demand exists today at current prices if the big 3 hyperscalers magically had an extra 2-3GW of compute. It’s not like Anthropic and OpenAI are turning away customers offering to pay API pricing..


Google literally limited Meta's Gemini usage because of capacity constrains, so yes paying customers are being turned away: https://www.ft.com/content/c5d52f72-71ef-40bc-bad3-61afdba8b...

> Everyone spending crazy money on capex right now says they have a crushing backlog and need more compute to protect their share price. I highly doubt the demand exists today at current prices if the big 3 hyperscalers magically had an extra 2-3GW of compute.

I don't get this though: The theory is all these hyperscalers are simultaneously spending buttloads of money on CapEx to the extent it affects their stock price, and then they would lie about the demand to protect their share price. Why would they do all that when they could just do nothing and keep their firehoses of existing business revenue untouched and maintain their stock prices on the upward trajectory they already were -- like Apple?

> It’s not like Anthropic and OpenAI are turning away customers offering to pay API pricing..

We don't know, but clearly Anthropic has been struggling to keep Claude's 9's better than GitHub's 9's even after paying through the nose for capacity from competitors like SpaceX and Google.

I suspect OpenAI is managing only because Altman scrounged for compute like a madman way in advance, and most of its traffic is free users who can be arbitrarily bumped down to weaker models whenever compute is low. Whenever Anthropic does that Claude Code degrades and people complain.


Paying customer at what price? My understanding of Meta’s request is API pricing per token (of course discounted for volume) for content moderation etc. which is a lower margin product than Gemini enterprise seats. They also prioritize internal training runs.

This doesn’t mean more compute at any price is worthwhile. It also doesn’t mean that the SpaceX compute deal would even be offered to Meta.

> I don't get this though: The theory is all these hyperscalers are simultaneously spending buttloads of money on CapEx to the extent it affects their stock price, and then they would lie about the demand to protect their share price.

It’s not lying - it’s optimistic revenue projections. Your Sam Altman point is an example, these RPOs are real but what’s questionable is whether the AI labs can generate enough premium token API revenue to actually pay those commitments. Today’s OpenAI annualized revenue estimate is only 40B. Will they actually be able to 10x that to pay those RPOs? I’m skeptical especially with offloading inference to cheaper models.

> Why would they do all that when they could just do nothing and keep their firehoses of existing business revenue untouched and maintain their stock prices on the upward trajectory they already were -- like Apple?

The hyperscalers with proven revenue streams and strong financials (Amazon, MSFT, Google, arguably Meta) benefit from making the game more expensive than everyone, will get at least 50% of their capex back from this peak supply/demand mismatch and maybe other than AWS could easily use any excess compute for internal needs.


It hasn’t bent already? 2022-2024 certainly seemed far more exponential than 2024 to present.


I find this astounding. 2024 to present thread is can write a coherent 15 line function to ... what exactly?

No future for research mathematicians othet than as tastemakers / agenda setters?


For frontier models, not local.

https://schema-harness.github.io/


Yes, saw that. They haven't yet released any code. Until they do, treat it with a huuuge grain of salt. In fact treat any 99% result in ML with a huge grain of salt.


No but the session traces are available. It passes the sniff test considering how AGI-3 is scored and how this wrapper works.

For example on bp35 it took fable 290M and >12k simulated turns for 566 real turns and finish more efficiently than a human.

Regardless of the true score I think the takeaway is the benchmark measures the wrapper rather than the model.

https://huggingface.co/schema-harness


Not my sniff test :)

> # FRAMEWORK ARTEFACT: the run's very first transition is replayed WITHOUT advancing state # (tools.py:954 and agent.py:468 both `continue` before `state = next_state`). So on the # level that contains that step (level 0) our counters start exactly one action behind. # That skipped step was action 1 with BOTH avatars moving, so seeding n=1, bumps=0 reproduces # the framework's lagged state exactly. # CAVEAT: this seed is only right while level 0 has never been RESET. If you ever RESET # level 0, change the seed to n=0 (after a reset the rollout re-inits and no longer skips).

from here - https://huggingface.co/datasets/schema-harness/arc-agi-3-sch...

That tells me that there is some leakage between runs. The idea of ARC3 is that agents start working blind, on new tasks, via API. A RESET is counted as one action. Without seeing the actual code that produced these traces we have no way of knowing how many iterations it took, if the "framework" played the same level multiple times (comment hint above makes it likely) and so on. That's why I said that before we actually see the code / can replicate / ARC team confirms it on new envs, this should be taken with a grain of salt.


The comment more likely means the harness source was read, not memory from a previous run and the first few turns of bp35 appear to be a cold start.

Sure none of this is certain without the source.

I do believe the authors that this schema significantly improves over the base, particularly given that it took 22x simulated turns over 14 hours, which is moving the trial and error to context rather than to game. I also don’t doubt there is some contamination.

Regardless, the approach is sound and I do believe it would significantly improve scores, even if that was +20-30 over baseline (49% in this case) it does imply the benchmark is measuring the harness more than the model.


Run policy search long enough with enough exploration and you can solve any of these games. But solving 110 games in 9 hours with a single RTX Pro 6000 doesn't seem likely. And if you did, you could keep the solution secret in exchange for the mountain of VC you would get to productize the approach. Not expecting it.


If you stop and think about the problem it really is quite simple. Just need to build a graph of the game state and then run A* to get to the end.


You really should play the 25 games before stating that it's "simple". The benchmark doesn't just track "completion", it also tracks the number of steps, and the score is based on the median steps took by human players. So in order to get 99% it would mean that the model solved every level of every game in less steps than the median humans. Which, having played the games and having setup harnesses for local models, I find hard to believe.

Also the models have to figure out what "end" means. And each game involves some kind of "gotchas" thrown in the harder levels. Some games are only solved by about 2/10 people trying them.

The 99% result most likely has some leakage somewhere, either in the preparation of the environments, or from session to session.

Seriously, play some of the games. They're fun.


You are not given the rules or the winning conditions. You are only given a potentially windowed and/or degenerate visualizer of the underlying game state along with the UI and told to just figure it out. And you as a human will, in a couple moves. An LLM? Not so much. But they do eventually solve them. And given enough moves, victory is inevitable, but you are penalized for taking more moves than a human, yet also slightly punished if you find a better solution by capping your reward to 115%.


> Say, Claude 7 creates a new, brilliant scientific idea every 1 out of 1000 times.

I would not hand wave that away - but Claude 5 feels closer to Claude 1 than the hypothetical Claude 7 you propose. I do not believe one can extrapolate LLMs that far ahead despite the very substantial progress so far.


Like two days ago Claude solved a century old math conjecture


If you are referring to the Jacobian conjecture Claude only provided a counterexample, not a disproof, neither of which is necessarily a “new brilliant scientific idea”.


A counterexample proves that the conjecture is false, not sure what you are talking about. And we can play word games all day about what counts as a "new brilliant scientific idea" but the fact is that no human had been able to solve it.


The whole "AI is a parrot" argument feels like moving the goalposts so quickly, you could actually hear the whooshing sound they make as they move.


No, that’s not my argument. Rather it is that a counterexample to a mathematical proof whether produced by human or AI may be a scientific discovery but it is not inherently new and brilliant by definition.

The case for AI is also weakened when the model is steered by an expert.


So what is your definition of "new and brilliant"?


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: