Hacker Newsnew | past | comments | ask | show | jobs | submit | Squarex's commentslogin

I don't know why, but the benchmarks still fails to cover the difference between large models and small ones. The small ones are great for many things, including general coding, but the larger ones, like fable and astra, have some kind of intelligence that is not present in the small ones.

More parameters = more facts stored. Knowledges are almost incompressible, where strong reasoning only requires a 3B core or so.

Weibo's VibeThinker manages with half of that: https://arxiv.org/abs/2511.06221 (They finetuned Qwen2.5-Math-1.5B for reasoning.)

It will always be until we have unlimited resources.

Costs can continue to be a factor in something, it is in most things, without it being the "driver".

It is possible to spend an infinite amount of money keeping a human being alive, and every year there are more ways to spend more money keeping someone alive a little longer. Once you're not spending an infinite amount of money for each person, you are in the territory of denied claims, "death panels", waiting lists, and all the other things people like to complain about depending on their political persuasion.

What does cost driven system and driver even mean? The cost of something is the ratio of demand to supply.

If something is not costing a lot, then it means there is sufficient supply that obtaining it is not going to require sacrifices elsewhere.

If something does cost a lot, then it means there isn't sufficient supply, and obtaining it is going to require sacrifices elsewhere.

Healthcare is obviously one of those things, especially with a top heavy population age histogram, where supply is not going to be sufficient so costs will remain high. And if costs are high, then obviously they will decide how resources get allocated, that is the very nature of something that is high cost.


It will always be the driver. The question is always going to be whether we spend X amount of resources to get Y probability of restoring Z health to someone (or someone to Z level of health.)

Costs are just another way to say resources. If we only have $100, and it takes $100 to have a 33% chance to give Bob a 50% better life, but $50 to have a 15% chance to give either Sara or Rose a 100% better life, we need a rubric to make that decision.

We can't just do first come first serve (maybe the first person in line takes all the health care for a 1% chance of getting 1% better), or pick the person we like better (who is we?). We have to figure out how to manage and grow the limited resources we have. We manage them by coming up with algorithms on how to spend money, and we grow them by cutting costs and investing. This is better when it is done in public through democratic means than when it is done behind closed doors (to avoid the above problems.)

There will always be death panels. There are currently death panels. The reason for single-payer is not to avoid trying (imperfectly) to maximize these equations, it's to get rid of massive amounts of administration and graft (cutting costs.) If all that savings goes to the Pentagon, we wouldn't be any happier than before single-payer.

But it's not good to come up with some fuzzy way to determine who gets treated. Efficient spending is a way of minimizing injustice. To say a health system is not being driven by costs is just another way of saying that it's going to be driven by undeclared and unexamined biases.


Has your job not be completely changed? Have you not noticed a worse job market?

I'm not sure I'd call that "meaningful" in the sense of "trillions of dollars and datacenters built all over the place" meaningful. Yes, I can now tell Claude to build a feature and it often does an okay job. I can ship faster. I can't tell you that it's worth spending the true cost of AI for that however. What'll be nice is 5 years of data showing that companies are raking in more money because they shipped more features. We haven't seen that yet.

Now, if Claude starts synthesizing more drugs that can help those with diseases (like OpenAI has just done) we might be trending in the right direction. For now, all I see is squirrels water-skiing. And while that's great... it's not meaningful.


there's a popular misconception that AI is so extremely costly that when the world finally comes to terms with it, we will stop using AI and things will die down. i'm sorry to burst your bubble: this won't happen.

the inference cost of AI has been dropping ~5x over every 6 months. AI today is around 100x cheaper than 2 years ago. it will continue to get cheap.


> AI today is around 100x cheaper than 2 years ago. it will continue to get cheap.

2 y.o model is probably 100x cheaper, but model size and requirements are also growing proportionally.

2 years ago, thinking was an optional, nowadays people use Opus with high effort almost by default.


That doesn't cause the amount that's already been spent to simply disappear. To recoup, they're going to need to sell a LOT of subscriptions. In about 3 years, we'll need to upgrade the GPUs that we use - if we're still chasing frontier models. That cost remains to be seen. There's a ton of unknowns here.

I don't have a bubble to burst.


My job hasn’t completely changed. It’s changed, but fundamentally I solve problems. Sometimes LLMs help me do that now.

The job market sucks, but there are so many confounding factors: pandemic over hiring, the loss of ZIRP, an economy that would be recession if it weren’t for data center investment, and may be in recession soon even with it.

This industry has been in down cycles before. It’s just been a while. So far it feels like a down cycle, not the sky falling.

It’s not clear what the future will hold, but I’m 100% sure it won’t evolve the way Anthropic predicts. Both because they have strong biases and because there's just no way to predict how this technology will play out. Humans are, if anything, remarkably adaptable.


I started my professional career in tech in the early 1990s. At that time, I sat in a chair with a keyboard and a screen to do my work. In 2026, I still work that way. So fundamentally how I work has not changed. Well, now I have two screens.

What? HN has always praised Library Genesis for example.

Libgen isn’t trying to sell you back the content for $100/month. No one would be complaining if OpenAI open sourced their models.

the second part is definitely not true

No one's complaining about stolen content from any of the labs releasing open source models

The ROI is probably billions of increased pre IPO valuation.

I don't know, in the Codex app, it burns the limit much faster.

They are all much larger and more expensive models. Google does not have a frontier model right now, but for cheap ones, they are better than event the chinese models now.

That's not being debated here. The initial reported numbers were false and this was simply pointed out. You're changing the subject.

Opus 5 medium has the same score as 3.8 flash on artificial analysis intelligence index.

Are you implying Google or Artificial Analysis are reporting false numbers? What's your source?


BTW you're comparing 3.8 flash high to opus 5 medium. 3.8 flash medium scores lower.

Flash models are on the order of 1/10th the size of Opus models, so some flex in the thinking level is fair.

When comparing closed models, the only thing that actually matters to anyone using them is some mix of cost and speed. Considering how much memory a server is using, when evaluating models that you'll never have access to in order to host yourself, doesn't really make sense.

Your comment is really strange, why are you defensive towards WarmWash when gemini flash 3.8 high is both 6 times faster and costs less, while having the same intelligence score as claude opus 5 medium?

>Considering how much memory a server is using, when evaluating models that you'll never have access to in order to host yourself, doesn't really make sense.

This entire sentence makes no sense given what is being discussed.

https://artificialanalysis.ai/models/gemini-3-8-flash

https://artificialanalysis.ai/models/claude-opus-5-medium


I was being pragmatic. These are closed models on closed systems that you cannot hope to host. They are only available as black boxes available over web APIs served by their owners. Within that black box perspective, that we're force to have, the size of the model is, quite literally, just how much memory that server is using.

intelligence/model size is not a useful metric for a black box user.

intelligence/cost and intelligence/speed is a useful metric for a black box user.

Yes, it's cool, but as a black box user, the amount of memory a model is using on a server that I do not own has exactly zero practical use to me.

Cheers!


Flash is just a name with no defined or consistent meaning even within labs, let alone between them. Considering both are closed weight, there is no way to truly assess how big the size delta between the two is. Then again, who cares about size, performance and end-to-end speed+cost are what matters along with task adherence, task assessment and so on.

Model size also can not be inferred by tokens/sec for a multitude of reasons, but to showcase two examples, Opus 5 and Sonnet 5, as well as Gemini 3.1 Pro Preview and 3.1 Flash have each very comparable output speeds when using the same deployment as a basis for comparison, despite it being very likely that within their generation, the former are larger than the latter. Feel the need to mention this, as I unfortunately stumble upon so many poorly reasoned, speculative hype post trying to infer model size via utterly unreliable metrics, not based in actual data.

It’s like comments below arguing about the reasoning levels not normalized to some metric (like cost, output token amount or duration) but just the labels or high, max, medium, etc. Those mean almost nothing even when comparing models based on the same pretrain (just compare GPT-5.4 to GPT-5.2), they mean less than nothing comparing different labs releases.


It's not totally a mystery

https://arxiv.org/html/2604.24827v1

The short of it is by using hard facts knowledge that is difficult to compress, and then quizzing models on these facts and calibrating against a bunch of open models, you can kind of feel out the size of closed models.


I really like that one, but it kinda highlights what I could have far better explained. Their 90% PI is three times in both directions. Between 3T and 24T for GPT-5.5.

That’s a massively wide, inaccurate and at best barely informative range, demonstrating that even the most well thought out method will yield little usable information.

Additionally, I got some private evaluation taking a similar approach towards gauging models in topics I’ve found either over or underfitted by labs. If we just used that to rank models (not get a potential size range but just a rough order) Thinking Machines Inkling would need to be lager than Fable 5.


> [...] shows an intelligence score of 59, the same as Opus 5 medium!

Nothing here is false, you are simply confused. You either didn't read what they wrote in its entirety or decided to reinterpret what they did write.


"Beating opus" is the false part, no?

Stop lying. mattlondon said "gemini-3-8-flash shows an intelligence score of 59" which is undeniably correct. You can't say that number is false. You're literally lying.

All you had to do is go hover your mouse over "Models" in the top bar, hover over Claude Opus 5 and and click on medium: https://imgur.com/mlRCrt1

When you do that you arrive on this page: https://artificialanalysis.ai/models/claude-opus-5-medium

The gemini flash page for reference: https://artificialanalysis.ai/models/gemini-3-8-flash

You have to be an incredibly dishonest person to see a 59 on both pages and say "the initial reported numbers were false and this was simply pointed out. You're changing the subject".


Better than even the Chinese models? That's a difficult-to-quantify, extremely rapidly moving target. Just today, Qwen 3.8 Max 0902 came out with a huge improvement over the previous Qwen 3.8 Max.

> "Google does not have a frontier model right now, but for cheap ones, they are better than event the chinese models now."

Just wow. Someone actually said this.


Google is targeting a different segment of the frontier.

They would, but it is not going to happen. We need solutions that don't count on that.


I don't like these comparisons. Sure it is impressive, but it does not have a world knowledge of larger models. It has most of theirs intelligence.


For world knowledge, you'd want it to find and reference the source material to be sure. At that point, it doesn't matter if the knowledge is embedded.


World knowledge also means knowing the various algorithms and ways particular programming problems are solved.

You can't search what you don't even know exists.


>You can't search what you don't even know exists.

that's not really entirely true -- one can google for "fast pathfinding' and stumble upon A-star , all that had to be queried was the intent/desire.

a lot of smaller agentic models and a lot of harnesses live on that premise.


Path finding is a very closed and well defined problem.


Keep in mind a web search might not include scanned books baked in the weights ;)


Are Chinese labs also acquiring and scanning books?


I think the big models have adequate recall, so tool use is probably unnecessary, but the user said the correctness of my response is important. Let me look up the data instead of relying on my memory.


If/when we can get larger context this will mostly be mitigated by these smaller models being able to search the internet.

Self-learning/improving would be even better but that's still a long way to go.


Search results suck because the web sucks these days. The big models from OpenAI/Anthropic have every book in existence baked into them


I don’t think that’s the right way to think about LLM ‘knowledge’. They don’t have absolute recall of everything in the training set. They have been trained so that they have weights that can predict what those books might say - that is, if they read them they would find the contents unsurprising. That doesn’t mean it wouldn’t be helpful to pull relevant passages of text directly into context for a particular task.


Does it really matter? What about including all relevant and up-to-date literature as skills for local models? I have no experience with this but I am pretty sure someone has already thought about it.


In a lot of spaces, this is actually preferable.

Ex - nodejs natively supports a huge set of typescript with built-in type stripping these days. But ask most hosted models to build a typescript project and they default to a heavy compile step, or a tool like tsx, ts-node, etc.

Models with lots of "world knowledge" have a good chunk of that knowledge go stale, and there's no real way to refresh it without training a new model.

Another classic example of this back in the day was to ask who the president of the US was, and watch different models happily give different answers based on the date they were trained.

---

Personally, I'm really interested to see if we're headed towards a spot where the model is entirely distinct from the knowledge store.

We're vaguely there with the ability for models to go search the web, but I think the reliability of that path is going to continue declining (more and more spam content, less and less genuine value).

I kinda want a paradigm where I can pick and engine and a knowledge bank, and combine them as I please.

Ex - if I'm doing gardening, I can pick "gardening for models (version 32)" as my knowledge store.

If I'm doing auto-repair... "cars for dummies (version 3)". etc...


> Personally, I'm really interested to see if we're headed towards a spot where the model is entirely distinct from the knowledge store.

This is what I've been trying to focus on with local AI for now. I've been trying to build all new documentation so it's more AI friendly. It's been pretty interesting. Qwen-35BA3B with a small prompt does a good job of surfacing what I'd consider institutional knowledge.

I've been trying to silo the docs I write from the model with a prompt that tells it not to use general knowledge unless asked to. From the anecdotal testing I did, Qwen-35BA3B is great for it. It does a really good job of following the prompt and calling tools, so I've been able to play around a lot to see what seems to work best.

Ultimately, I think one of the most effective uses of AI will be having a distinct knowledge store combined with an opinionated agent (and sub-agent) setup along with different models for each task.

Who owns the knowledge store is going to be the big caveat. Right now I think the big online models are trying for generic, persistent memory and I'd be very hesitant to let that happen. Think of having someone with a perfect memory following you around forever, but someone else has the ability to make them disappear. That's not a good situation.


One of the consequences of encountering a lot of LLM generated text which includes things the model vaguely remembers from its training is that honestly I have grown less tolerant even of human comments and documents that are based on mostly ‘I seem to recall that…’ level sourcing.

In a discussion on economic history, say, someone will opine that Alexander Hamilton had some particular opinion about tariff policy… based on their having a vague memory of a blog post where someone quoted a passage in support of some point. But wait - you can search the federalist papers, the text’s right there to be read, before you commit to saying online ‘Hamilton thought tariffs were a great idea’ you could take your internal ‘I seem to recall reading something about hamilton’s opinion on tariffs’ thought and turn it into a little RAG query where you pull up a source and check before you put another factoid out onto the internet.

And so I feel absolutely the same way about LLMs. I don’t care how much factual information was in the training data, when the LLM wants to rely on something it vaguely recalls having been trained on, it owes it to me to dig up a source and vet it.

There are limits to this, of course. I don’t want it to be thinking ‘but wait, maybe my memory of Python syntax is faulty. Is = used for assignment? <web search>…’.

But in general some caution about repeating vaguely recalled easily checked facts is warranted.


At 125B + 51B I'd expect it to have some degree of world knowledge, clearly in the middle between small models like qwen 27B, and huge trillion parameter models.


Probably per year, still a crazy amount.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: