Hacker Newsnew | past | comments | ask | show | jobs | submit | smcleod's commentslogin

How is that not a native thing with Actions? GitLab had multi runners since early on.

That's not really how it works. I'm able to get 5-10x more work done, perhaps more, it's a significant accelerant to those that know how to be productive. The problem is more for those who don't - and we bring them up to speed or help them pivot to other work.

How does that shake out for your company? Are you growing 5-10x in sales? I'd guess not and that is sort of the rub with all this. Yes we can push more frequently to the git repo but that wasn't the limiting factor of scaling the business.

I work across multiple companies, but in general company profits are up across the board. Salaries are not. The larger enterprises are the ones that move the slowest and are the furthest behind. It's not just one thing (productivity), there's all of capitalism at play.

The new releases and breakthroughs do the opposite for me - I feel energised by them. I felt like nothing truly that interesting had happened in tech for quite some time, now it's like the space race (except there is no one moon to reach).

I appreciate boring tech as much as the next well worn engineer and I'm not saying this is all positive but it's so sure as hell thrilling and you don't have to be an astronaut to immediately benefit (or suffer I guess) from it.


It all smells pretty load-bearing to me.

But is it byte for byte identical ?

You're right to push back.

I wish someone would do this for Apple Music. Their app is dreadful!

Bandcamp, and I have on several occasions emailed the band directly and asked if I could purchase directly from them.

I'm running the UD-Q4_K_S on my 128GB M5 Max with 180K context, it uses around 100GB~

Not really compatible on all fronts, it's very capable especially with tool calling, workflows, logic and its base coding ability, but it's only a 27b model so it does not have anywhere near the level of knowledge baked in as larger models. This does not mean that it's not a good or useful model - it is on both accounts and very efficient, but it's not similar to a large model generally speaking.

50-70tk/s is what I get on my m5 max on a 5-6bit Qwen 3.8 27B?


I don't know what black magic you're up to but I see more like 30-35t/s on a 16" M5 Max using 3.8:27b Q4, regardless of whether it's mlx or gguf.

qwen3.5:122b-a10b is significantly faster at around 60-65.


No magic, just oMLX with MTP. You can look through the speed the community is getting here: https://omlx.ai/benchmarks/performance?model=qwen3.8&chip=&c...


With MTP? I get 25-30 TPS on a strix halo. 50+ on a M5 max should very doable. Dflash (2) will push your TG even further


It's a bit deceptive to state inference speeds without mentioning the additional things you're doing to achieve them


I tried 8-bit, perhaps I should try 6-bit.


That was mainly before the M4 generation when they didn't have matmul instructions.


M5 prefill is much faster than M4.

I've seen benchmarks that show 4-5x faster of M5 Max vs. M4 Max.

For local models you're likely using M5 Max, prefill is low thousands of tokens per second, as opposed to, say high hundreds with M4 Max.

For larger dense models, some fraction of that, but similar multiple.


Yes, I have the M5 Max. But there was no matmul acceleration before the M4 which made things a lot slower.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: