> “Rivers and birds” highlights the resilience and diversity of Europe’s natural ecosystems by showcasing different stages of rivers and various bird species, emphasising the importance of nature and environmental protection. The European institutions featured on the banknotes remind us of the fundamental values of the European project, which also embraces environmental protection.
For my AI Agent it sometimes detects if I manually modified the file contents or git state. And it always assumes it must have made a mistake. It's sort of annoying actually.
Yeah, I suspect RLHF conditioning heavily discourages models from ever implying that the user could be in the wrong (or, rather, to assume that they are in the wrong by default, since editing a file isn't really "wrong" per se). Though looking at the reactions to Opus 4.8, which has a more contrarian nature and caught a lot of flak as a result, that's probably for a reason.
It's also the reason why I ran the two tests on open weights models with unredacted thinking traces. Gemma never flagged anything in its response either, only in its thinking. Without knowing how the summarizer models are prompted, it's impossible to tell whether it was a genuine miss or just something the summarizer decided to omit.
DS4-Flash definitely stands its ground when I'm obviously wrong (i.e. me reading ifneq as ifeq for several minutes straight), and I've seen at least once a "thinking" trace that was almost verbatim "the user has changed this". That's local, so thinking traces are raw. Pretty sure the more powerful models (500+GB weights, closed SOTA, etc) are even better at this - haven't had GPT5.5 with codex sugar coat things for me.
I recently made an AI Agent and surprisingly coding with DeepSeek V4 Flash is quite cheap. It probably has to do with the aggressive prompt caching. I'm using OpenRouter with Novita AI as the preferred provider.
I’m using zen because I have a Claude subscription and just like dabbling with the other models and I was shocked at how little flash cost but it was noticeably not at the level I’d like my model to be.
For me MiniMax 3 has really hit the sweet spot of being very cheap, though more than flash, but I’d also very capable.
Not to be that guy, but the correct term is Open Weight LLM. And I’d argue it already has. Many open models are already very competitive with closed models at a fraction of the cost.
First of all the May jobs report was mostly in temporary workers possibly due to the World Cup. Second as already noted the jobs are mostly in healthcare. Third job openings does not equal employment. These numbers have been diverging for a while likely due to people holding multiple jobs. Also I believe the evidence suggests the job crisis is due to WFH and not AI https://news.ycombinator.com/item?id=48326721
They’d have to keep this up for 12 months which is not guaranteed as you have to factor in hardware depreciation and other costs. Also they’d have to issue more shares to get over the liquidity requirement. And looting may be too strong of a word. They are nudging your returns down a little bit. Although SpaceX could go up too, crazier things have happened.
Yeah, while the world is mostly clickbait at this point, "magically makes them eligible for inclusion in the S&P500" is probably some of the best (i.e. worst).