Hacker Newsnew | past | comments | ask | show | jobs | submit | RussianCow's commentslogin

Which doesn't matter if they're still not profitable.

I don't see any kind of input cache discount listed on your pricing page. Do you offer that, or is all input priced the same?


The issue is that all input (including context) counts towards that limit. So 10 requests with 50k of context will blow through the limit, even if little to no output was generated, which is incredibly easy to do with agentic workloads.


This is pretty terrible advice when there are dozens of AI inference providers out there serving great models with significantly more cost effectiveness than you'd get from buying your own hardware.


I don't trust any company with their word on anything. Luckily, privacy policies are legally binding.


> Luckily, privacy policies are legally binding.

Companies violate them all the time and massive leaks happen a lot.

The punishments are trivial.


GPDR fines in the EU are NOT trivial.


> Luckily, privacy policies are legally binding.

Laws are violated all the time. The graveyard is full of people who had the right-of-way at a crosswalk ...


Why does that matter? Legally binding just means "slightly more expensive when we get caught"


Presumably the number that OpenRouter shows is averaged across all requests.


Most providers do what's called "prefix caching", where each turn in a session is cached such that sending new messages with the exact same "prefix" (set of previous messages) gives you the cache read price on that input instead of the full price. As long as you're not changing your system prompt, available tools, etc mid-session, you automatically benefit from this.


But presumably everyone in your company/team is using Jira, so it's not an "ad" because it's a product already used internally. Claude is appending these links to all commits by default, whether or not others on the team use Claude. Those are very different things.


How else would you expect them to calculate it?


Do you really think they're docking points because cache invalidation due to provider switching? Seriously llms are frying ya'lls brain.


They're not "docking points", they're calculating it in the most straightforward way. If I start a session and the majority of requests are sent to Provider A, and my last request gets routed to Provider B, I have a 0% cache hit rate with Provider B. I'm very curious how else you expect this to be calculated? Do you think they're completely omitting requests that switch providers mid-session?

FWIW, I get significantly higher than listed cache hit rates when I pin my session to a specific provider, which is further evidence of the above.


Why wouldn't you only calculate consecutive requests with the same provider....


This is very much NOT my experience in practice, even though it's how I would expect it to work. OpenRouter will happily bounce you between several providers (none of which have downtime) even within the same session. Requesting specific providers is the only way I've been able to hit a cache rate above 90%.


Same experience here. It would choose 2 providers and then bounce between the two every 5 requests or so.

I don't know why there's no "Pick the cheapest provider above nTPS on first request and stick until cache bust" setting.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: