Hacker Newsnew | past | comments | ask | show | jobs | submit | nasutton12's commentslogin

fair longer is always helpful. on a 24gb mac i'm extremely constrained, but all these exercism tasks were plausible to pass at this context length.

this is in the repo in benchmarks/matrix for most of these metrics.

uv run python benchmarks/matrix/run.py setup. i'd happily take your hardware's numbers!


sounds fun, i'll try and add it.

the early 1.x versions of chad were tied to ornith. i still miss the speed of that moe. https://huggingface.co/nathansutton/Ornith-1.0-35B-UD-Q2_K_X...

I skimmed chad the other day. There's very little to it (by design). afaics you should be able to replicate what chad does by copying the chad system prompt into a SYSTEM.md for Pi. Then you could use it with whatever model/provider you want.

pi is a fantastic harness! they are a good default in the same way llama.cpp is. it works with everything and that is the point.

i was steering chad in the opposite direction. one model & one set of silicon -> taken to the max. swap out your CHAD_MODEL and it still runs, you just leave the drafter and the kernels behind.


Love the idea…

Except the "one model" is too small for a 64GB Mac much less 128GB, sad since the Q3 is proven less competent.

Offering a Q boost (with no leave behinds) on first run would be a bump worth some vibe coding while.


There are a comical amount of harness design features in the early 1.x release. A full LSP integration, find symbols, batch edits, etc. Terminal bench isn't everything but I spent a couple of weeks testing different combinations to no statistical effect greater than a bare loop. It was like running uphill against what the underlying LLM wanted to do.

sorry its a notion.

sorry, its a notion not a real site.

The annoyance is real. Let's do it.

I'll try it.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: