Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Cerebras supports prompt caching and has a doc about it. A fairly standard automatic prefix-based implementation with 5min expiry.

They do not seem to discount cached input for the self-serve Developer tier. Maybe they do for enterprise rate cards?

https://inference-docs.cerebras.ai/capabilities/prompt-cachi...



ah, that makes it feasible! Okay, glad it's not technical limit. They should fix the pricing...




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: