More on the topic...
Generating detailed summary...
Failed to generate summary. Please try again.
Boris Cherny, who built Claude Code at Anthropic, has pinpointed nine common habits that burn through roughly 73 percent of your Claude tokens before the model even tackles your latest prompt. The biggest chunk—14 percent—vanishes in that initial CLAUDE.md system prompt every session. Another 13 percent goes to Claude re-reading old chat history you didn’t need it to revisit. Hooks or plugins you installed weeks ago quietly nibble away at 11 percent more. Cherny argues that most of the “Claude got dumber” gripes really stem from these hidden costs, not a sudden drop in model quality.
He lays out the remaining six token drains in his podcast episode. A few highlights: overly verbose context-loading steps, redundant API calls buried in your automation, unnecessary safety checks you never switch off, and half-forgotten debug instructions you left active. Cherny notes that if you’re hitting Claude’s maximum token limit more than once each week, you’re likely guilty of at least four of these wasteful patterns—and probably closer to seven.
Over 400 hours of hands-on Claude use lets Cherny tie each pattern to a clear percentage hit on your monthly token budget. He’s also cataloged simple countermeasures: trim your system prompts, prune chat history when it’s no longer relevant, review your active hooks, and audit each tool you’ve connected. Instead of rebooting Claude or switching models, you can reclaim thousands of tokens every month just by cleaning up behind the scenes.
Questions about this article
No questions yet.