Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I don’t think they will until they change the architecture.

They don’t have a prefix cache like other providers, or at least don’t have a discount in their billing structure. Each message charges for the whole context window. It’s wildly more expensive for long multi turn scenarios with lots of tool calls (coding). It’s better for short few turn tasks.

Edit: I don’t know if they actually have a proper cache. This could just be a billing artifact.



Cerebras supports prompt caching and has a doc about it. A fairly standard automatic prefix-based implementation with 5min expiry.

They do not seem to discount cached input for the self-serve Developer tier. Maybe they do for enterprise rate cards?

https://inference-docs.cerebras.ai/capabilities/prompt-cachi...


ah, that makes it feasible! Okay, glad it's not technical limit. They should fix the pricing...




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: