Hacker Newsnew | past | comments | ask | show | jobs | submit | cacio-e-pepe's commentslogin

Could you expand? Curious.

sorry for the late reply, see answer above


Key quote : "Solving the problem by purely AI-powered methods [would be a] net negative for the progress of mathematics."

> So we did mechanistic studies on small models, Gemma 4 particularly, and found the hidden state for different layers carry meaningful self-awareness signal for various situations.

Neat! Just to make sure I understand - you trained your probe layer to take this hidden state and predict p(wrong)?

Curious to learn more. Any more info on your approach (esp the mechanistic study)?


Correct, the study is verbose, we will compile into a neat shareable report and publish once we solve the pending caveats. Interesting username btw haha.


Nice, looking forward to the report.

And thanks, huge pasta fan :)


anytime!


Honestly, stellar performance by the model at the capability being measured.


Okay so the diffusion model generates image assets, and you iterate on those, and then another model turns your favorites into code, is that right?


Exactly right. You design with the diffusion model, then hand those designs off to an agent to implement.

It's a lot like having an architect create plans for you before handing it off to a builder. In my (obv biased) experience, you end up getting better/more creative results with this approach. You're using the best model for the job at each specific task, ie a diffusion model as the designer, and a LLM as the engineer.


Looks interesting. I'm a bit too busy to read the docs, but I'm curious - how does it work?


What are the core knobs for a RAG pipeline? See any interesting patterns? Eg in practice, what did the agents tend to tune? Did domain matter?


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: