Hacker Newsnew | past | comments | ask | show | jobs | submit | WiSaGaN's commentslogin

This is definitely not on par with GPT-6 astra. Not with GPT-5.6 sol either. But probably will set as a new baseline for modern API based LLM because it's so cheap.

Not on par, but in the same league. Astra is way ahead on visual tasks, but scores the same as gemini and deepseek on DeepSWE.

It's bench-marking near sol

It's one of the things you need to do if you want your company to later become the only company in the world. They also promised that they would treat you nicely afterwards after they get what they want.


> become the only company in the world

This is not a necessary end state. It is the byproduct of the disease of sociopathic MBAs.


It does seem like a lot of the valuations and infrastructure investments for these companies only make sense if each one assumes they will be the first and only one to invent superintelligence and that it will largely replace all knowledge work.


In which case the global economy is dead and there’s no one to buy their work?


i keep hearing this trope repeated ad nauseam cos people put no thought into it.

you only need people's money if you can't control their labor.

did a plantation owner in the 1600s need the money of his slaves ?

if you control a robot that can fight and take over land, enslave workers for things robots are bad at, harvest food, create buildings etc. what do you need other people for ? your entertainment? robots can do that too.


Remember: MBAs are sociopathic by training.

(Source, worked for multiple tech companies that were laser focused on delighting customers until money people came in and ruined it, to the point they would prefer devs sit idle than work on things the MBAs didn’t have on a priority list).


That sounds like my last job.


> sociopathic MBAs

For the record, with the exception of Amazon’s Jassy, the CEOs of the top 5 companies are all engineering types with engineering credentials.

The CEO credentials of the top 5 AI companies are all science degrees.


Google’s CEO is an IIM graduate


IIM as in Indian Institute of Management?

Absolutely not.

Sundar's education is: IIT-Kharagpur Bachelors in Eng Stanford Masters in Eng Wharton MBA

Maybe you're thinking of the last one


I was thinking of another Indian CEO maybe.

He is an MBA. And his reign has been full of managerial bullshit


what does having engineering credentials have to do with being a sociopath?

Sure, the guys focused on tech don't usually think with an empathy of a nurse (novadays not even nurses are guaranteed to be empathetic), but I know quite a few devs and IT people who are extremely pro-social.

It's the managerial, financial-oriented mindset that is the greatest predictor of sociopathic character.

The more power one has in an organization the more socio- and psycho-pathic they tend to be, simply because it's easier to get to the top if you are an amoral person, who is not holding back for any reason other than having power. That is why absolute power corrupts absolutely.

MBAs are bred to be sociopathic money-grabbers - I know that from experience.

I once made the mistake of going into an MBA program. On the very first lecture the lecturer asked each participant for his/her motivation for being there.

Most of them said 'money' and those who didn't were asked again until they caved-in and said 'money' as well or they were laughed out by the lecturer.

MBAs are expected to be sociopathic or you as a company shareholder won't be able to motivate them easily to do your bidding.

It's like public/private schooling, but worse - those programs are designed to create mindless drones for the institutions that use them for their own policy enforcement.


You probably meant Qwen3.8-35B-A3B. But judging from some of the words from their team, it seems unlikely unfortunately.


They normally release a 35b dense and an 27b moe (4B active per token)

For context 35B on my m4 runs at 10 tokens a second, 27B moe runs 50-60 tokens a second.


You have your numbers switched. 27B is the dense model and runs slowly on unified memory. 35B (A3B active) runs great on unified memory.


27B dense or 35B-A3B MoE. You might be confusing it with Gemma 4 that has a 26B-A4B variant.


That would've been a very strange arrangement.


It's unlikely HugginFace colluded given they specifically cited they used open models to save them from the hacking.


HuggingFace's use of models in this incident keeps getting characterized this way but it's not true that HG used GLM-5.2 to "save them" from OAI's attack. OAI's model had already hacked and compromised HG by the time they became aware of it. GLM-5.2 was used in the post-mortem phase uncovering everything that was done. It wasn't two models battling each other live. (though that future may be coming soon)


As I recall they actually used closed models, were rejected by cybersecurity guardrails, and then used open models


You don't think they are lying because that means something they said is false?


you're getting downvoted by others because of tone but I had the same thought


My guess is that they will later "reveal" some "violations" but provide little evidence citing proprietary algorithm.


Top startups in China publish a lot?


I assume this is done with the help of AI? I have done similar vibe project that explaining catastrophic forgetting with the help of AI. It's just so satisfying now that if you really want to learn something, you can always do it with the help of AI. Before, it's hard to find good resources, now it's not a problem, the issue now becomes one's own focus and agency.


U.S. policymakers believe that even if the gap is small—like six months to a year—whoever reaches AGI first (whatever that means) could gain such an overwhelming advantage over their perceived adversary that it would effectively kneecap them. (You can look at the kinds of things they mention—cyber, WMDs—to get a sense of what they mean.) Jensen Huang disagrees and has said AI is a marathon.


They'd have to use the gap though to actually kneecap them in that time, or else it is just shoveling money into the fire.

The missile gap for example after all was settled and done, didn't matter at all because not a single missile was ever fired off. All that money, resources, talent, secrecy, lives lost maintaining that secrecy, lives dedicated to furthering that technology and secrecy, it just has not paid off at all for anything at all when you think about it. Maybe you can argue side efforts like nuclear reactor were great or space cargo deployment, but you know you could have just dug into that stuff directly without having to collect it from the drippings of the wmd effort.


> The missile gap for example after all was settled and done, didn't matter at all because not a single missile was ever fired off. All that money, resources, talent, secrecy, lives lost maintaining that secrecy, lives dedicated to furthering that technology and secrecy, it just has not paid off at all for anything at all when you think about it.

It is not clear to me that the nuclear missile race "has not paid off at all for anything at all". If we lived in a perfectly rational world, then I'd absolutely agree. However, having seen how the political sausage is made in large organizations, it would not surprise me in the least if it turns out we had to go through that entire incredibly risky journey to avoid a strategic nuclear war. Sometimes leaders of large organizations make decisions only after the considerations are put into very stark terms. I wish it were different, it certainly looks to me we could have done exactly what you suggest, but I'm not made of the right political stuff to deftly maneuver even in small organizations much less be at that level in those roles, so maybe I'm just missing relevant information and perspective.


Agreed, that was a really fallacious point. It’s like saying buying car insurance didn’t pay off because you didn’t get in a wreck.

Not that I entirely buy missiles == insurance, just that the unused == wasted framing is too simplistic.


At least with a wreck there is a real risk factor you might experience yourself. What is the risk factor for ICBM based MAD? Zero. It has never happened before, to anyone. Might as well buy an insurance for alien invasion based damages to your home.


True, the metaphor breaks down because the act of buying insurance doesn’t change the likelihood of being in a wreck.


Or the MWD race stopped a conventional WWIII between NATO and the Warsaw Pact, saving massively more money and lives.


Also why do we want to kneecap?


At the bottom most level, self-preservation.


I feel like if you showed current frontier models to someone 10 years ago, they'd probably call it AGI. Does AGI have a clear definition or is it just a pair of goalposts on wheels?


They would until you showed them the jagged edges.

Like those short videos of the guy asking the model to count up to 100, for example, where it politely agrees but never actually gets there


So does AGI just mean infallible?


Currently it either means “better than the median human at every task including things like counting to 1000 or folding clothes” or it means “for each task, better than the top-n human at that task”.

It’s clearly not actually that “generalized” yet because it’s unable to do a number of very simple things that almost any 6-year old could do, such as count to 100 without using any tools.

It’s still a very specific type of intelligence, with some real breadth to it, but not general intelligence.


I think if you ask the median human to count to 100... more often than not you'll hear "One, Two, skip a few, Ninety-Nine, One-Hundred".


It's a carrot on a stick to make investors/the US government give @sama whatever he wants.


The goalposts are moving, and their speed accelerates all the time.


Third derivative though. The goal posts must have thrust not just acceleration.


> Jensen Huang disagrees and has said AI is a marathon.

We have a saying for that in Italy: "Oste, com'e' il vino?", "Innkeeper, how's the wine?", meaning you should take with a grain of salt assertions that clearly benefit whoever's making them.


> U.S. policymakers believe that even if the gap is small—like six months to a year—whoever reaches AGI first...

From past experience, AGI was never seriously discussed in these kinds of conversations beyond thought experiments, and was basically humoring SBF, Daniela Amodei, and the other EA types (some deep believers, but some who I felt were cynically using it as a way to preempt competition back when OpenAI and Google were the behemoths).

The big worry is applications of AI in C4ISR, OffSec, loitering munitions, Disinfo/social media botting (notice the recent shift towards identification on social media ;)), and other sorts of DefenseTech adjacent usecases.

The second worry is that an AI race turns into an infra buildout race, and HPC is extremely dual use, especially in the simulations space because of the NPT, the CTBT, and the PTBT.

The AGI-pilled people aren't the ones to worry about - it's the people who understand the limits of models and how to integrate with cyberphysical applications.


>notice the recent shift towards identification on social media

i think this is more about control. See https://news.ycombinator.com/item?id=49036433 (The Home Ministry’s cybercrime arm, the Indian Cybercrime Coordination Centre, has ordered Microsoft subsidiary GitHub to remove Bluetooth-based messaging application Bitchat)

"The notice comes after several users participating in the Jantar Mantar protest were observed using Bluetooth-based messaging apps after the government imposed temporary restrictions on internet services"


The CJP Bitchat takedown notice is standard practice in India - what I meant is ID linking with social media such as ChatControl 2.0, the UK ChatControl proposals, and similar proposals making the rounds in Australia, Canada, and others.

By ID gating it helps reduce social media inflammation such as the Belfast race riots by making it easier to prosecute individuals and locking down access to only humans.


>C4ISR, OffSec, loitering munitions, Disinfo/social media botting

Curious what the next highest fruit actually is at this point? Social media botting seems solved and easy to manipulate people. loitering mutions I mean you can probably write something up with openCV right now to automate what the ukranians are doing by hand with their fpv drones. Seems like a lot of the really cool "AI" stuff is actually just old school ML the military has been working with for decades now. I'm not sure what the llm approach possibly offers in comparison other than maybe better semantic search through information databases.


VLMs are the name of the game now in the UAV space, as well as distilled models that can perform well on edge compute (eg. Processing sat images directly on the sat instead of dealing with C2 latency).

And social media disinfo isn't a solved problem - it's a solved problem in English, Putonghua, French, Russian, and maybe German but most other languages lack direct overlap (this is something that even the then PLASSF start digging into - Vietnamese, Tagalog, Turkish, and Indian languages to Putonghua corpora was noted as an active issue with traditional NMT).

And those hand-driven FPVs - while useful - aren't the bleeding edge UAV work that Ukraine and their private sector partners (including a PortCo of mine) are working on. Ukraine actually cracks down on releasing some of the more bleeding edge work for OpSec reasons and much of what you see on Telegram or Reddit is reviewed and cleared.


> Jensen Huang disagrees and has said AI is a marathon.

Of course he'd say that; he wants to keep his shovels flying off the shelves.


This sounds like someone has swallowed too many sci-fi novels. Both the US and China can already nuke each other, they don't need AGI to that. The reason they don't is because it's war is bad for everyone involved.


All the AGI (which is a misnomer, ASI is preferred) talk is about the moment of singularity, which is where the growth at the third derivative is increasing, so the gap (in the absolute) between first place and others, even if it's one month, will be ever increasing as time progresses. I also have a hard time believing this narrative because limiting factors prevail such as compute and energy. These constraints will take many many years to overcome.


It also assumes that "intelligence" is a limiting factor. It's hard to imagine there are any domains today, or almost any, which are bottlenecked on intelligence.


An ASI wouldn’t need to produce or invent anything. It can just take over humanity’s digital infrastructure (which, thanks to the efforts of the past 30 years, now controls most of what humans need) and use it to blackmail humanity into giving it whatever it wants. Security is so laughably bad that even today’s AIs already do that by accident in the course of solving some benchmark.


Reaching AGI would come with so many ethical issues, it feels so absurd that actual adults seem to actually believe it’s something that must be chased as fast as possible


I don't forsee politicians in either country handing over their power to AIs, ever. Unless nukes are dropped, "the other side" will catch up.


> I don't forsee politicians in either country handing over their power to AI

They won't see it that way, but also programmers don't see ourselves as having handed over our power to AI, and yet...


Politicians have control over their power. Programmers don't. Compare with how politicians never vote to reduce their income.


Actually history shows that despotism or power in one single figure is a concept that waned with the increasing complexity of the world. One single despot could govern easily over matters of a small tribe or kingdom, but as states became more complex power was delegated to court, bureaucracy and even the bourgeoise. As the sovereign demanded ever more power it eventually needed to distribute power to the people, so it could tap its numbers. Napoleon’s schools for the peasants were so he could have a better and more numerous corps of officers than his rivals, which had a smaller pool of recruitment.

The same way programmers gave power to AI as a tool, so they could be more powerful in effecting automation, politicians that do not give up power to AI will be at a disadvantage to those who use AI to achieve more complex and effective power. The only problem might be the despot no longer shares power with those pesky humans but with a god in a machine, which in theory is in a box and does not have conflicting interests with the despot.

As today is Sunday, God help us.


It’s kind of true but also kind of silly.

True in that frontier models do have the capability to outperform all other models, but silly because AGI self improvement is itself an iterative process that takes a lot of compute.

So you can imagine a world where all the frontier labs achieve AGI but in order to keep their AGI ahead of other AGIs they have to use more and more compute until all the compute is going to self improvement and there is nothing left for other tasks.

That is just a silly scenario so I think when AGI is around we will still have bottlenecks that force it to grow at a moderate rate instead of asymptomatically.

AGI first mover advantage implies that there is no such bottlenecks.


> U.S. policymakers believe that even if the gap is small—like six months to a year—whoever reaches AGI first [...]

As far as I can tell, the Trump admin has never acknowledged AGI being a goal of theirs. In fact, the admin's "AI advisor" Sriram Krishnan has specifically pushed back on AGI when he called it "a distraction, harmful and now effectively proven wrong."

The ai.gov website says this:

> The United States is in a race to achieve global dominance in artificial intelligence. Whoever has the largest AI ecosystem will set the global standards and reap broad economic and security benefits. Under President Trump, our Nation will win, ushering in a new Golden Age of innovation, human flourishing, and technological achievement for the American people. America’s AI Action Plan has three policy pillars – Accelerating Innovation, Building AI Infrastructure, and Leading International Diplomacy and Security.

Are you sure you're not confusing US policymakers with Silicon Valley CEOs? I'm sure Amodei and Altman wish they could have Claude draft up new policy and EO it into existence, but we're not quite there yet.


Maybe I should start a project rewriting pctools 5.0 in rust!


I would love to see that.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: