All stories
PricingOpenAICaching

New Models, New List Items

OpenAI used to write to the token cache for free. With the 5.6 models that changed — and the new cache-write fee quietly reshapes what efficient usage looks like.

Richard WangEngineering, Ramp2 min read
Model Manager
FileViewToolsHelp
ModelStatusType
Claude SonnetNEWFrontier
GPT-4oGeneral
Gemini FlashFast
Open-weightLocal

OpenAI models were more efficient than Anthropic ones in previous generations for one simple reason: OpenAI would write to the token cache for free. With the release of the 5.6 models, OpenAI has made a decision to start charging for cache writes. Here's how the new pricing lands across the lineup.

ModelInputCached inputCache writesOutput
gpt-5.6-sol$5.00$0.50$6.25$30.00
gpt-5.6-terra$2.50$0.25$3.125$15.00
gpt-5.6-luna$1.00$0.10$1.25$6.00
gpt-5.5$5.00$0.50$30.00
gpt-5.5-pro$30.00$180.00
gpt-5.4$2.50$0.25$15.00
gpt-5.4-mini$0.75$0.075$4.50
gpt-5.4-nano$0.20$0.02$1.25
gpt-5.4-pro$30.00$180.00

The cache-write line is new, and it changes the math in three ways.

  1. Long workflows barely notice. Workflows that require repeated back-and-forth with the model should see only a modest difference in price. Our data shows that cached input tokens, in most cases, account for more than 90% of token volume. As long as you keep a consistent "hot" session, your prices will differ — but not by a large amount.

  2. Newer, short sessions just got more expensive. Every time you introduce new context into your workflow, the harness you are working with (Codex, ChatGPT, …) decides whether to cache it or not. If you repeatedly "start new sessions," the prompt caching will kick in, and suddenly you're likely to be paying for unnecessary caching — increasing your costs by up to 1.25x.

  3. Don't take breaks — literally. The cache defaults to a 20m minimum. That means if you take repeated breaks longer than 30m away from your session, you can actually end up having to recache a large portion of the chat once the cache expires, incurring an extra 25% premium on any tokens that need to be freshly cached.

The takeaway isn't "5.6 is expensive." It's that the shape of your session — how hot you keep it, how often you restart, how long you step away — now shows up directly on the invoice.