More on the topic...
Generating detailed summary...
Failed to generate summary. Please try again.
Anthropic cut the Claude Code prompt cache time-to-live from one hour down to five minutes in early March after rolling out a one-hour TTL on February 1. That shift hurts long coding sessions: every time the cache expires, Claude has to re-process your entire context, driving up token usage. Writing to the five-minute cache costs 25 percent more per token than a cache hit, and misses with a 1 million-token context window cost even more. Users on the $20/month Pro plan report burning through their quotas so quickly that they get only two prompts in five hours.
Sean Swanson, the developer who spotted the change, says his $200/month plan never ran dry until March. Jarred Sumner from Anthropic argues the five-minute cache is cheaper overall because many requests never revisit their context. He also says the client auto-selects TTL and there won’t be a global switch. Meanwhile, Claude Code creator Boris Cherny admits that full cache misses on large contexts are costly and says Anthropic is testing a 400,000-token default window with an option to go up to one million. Users running multiple agents and long background instructions have seen their costs spike.
Complaints go beyond caching. Multiple developers report that Claude’s performance has slipped: sessions get stuck in loops, repeat the same point, or churn out lengthy “but wait, actually” passages. An AMD AI director noted similar “overthinking” since the late-March update. Some bugs in the caching code could be skewing the numbers, but many believe that underlying compute quotas now buy less processing than before.
Questions about this article
No questions yet.