Hidden Claude Nerfs and What You're Actually Paying $100 For ✂️

As I recently wrote, the era of cheap AI agents is over. But how exactly companies are trying to balance unit economics deserves a special place in textbooks on audacity.

Right now, Reddit is in an uproar. Users of Claude premium plans (including the top-tier Max 20x at ~$200/month) have discovered that their limits were quietly reduced right in the middle of a paid month.

If before a 5-hour window was enough for an intense coding session via Claude Code, now a couple of prompts on Opus 4.6-4.7 consume 100% of the limit in 10 minutes. Support has gone into full defense mode, feeding people canned responses about "dynamic tokenization complexity" and "context size," refusing to escalate tickets to real humans and ignoring the real problem.

Obviously, Anthropic simply can't handle the inference cost of agentic workflows. But instead of honestly raising prices (or at least providing a transparent consumption metric), they've rolled out A/B testing of "throttling" users. Today it's your turn, tomorrow your neighbor's.

And now — watch closely 🤡

Against the backdrop of this acute shortage of GPU power for those paying real money, leaked data shows that Anthropic is preparing to launch Orbit — a proactive background assistant.

This thing is supposed to sit in the background, vacuuming your GitHub, messengers, email, etc., to "generate personal insights." The announcement will likely happen today at the Code with Claude conference in San Francisco.

Funny, right? The company physically lacks the computing power to process a direct request from a programmer trying to close a task and paying a hefty price for it. Yet they're building a feature that will burn tokens 24/7 in the background, reading your conversations, only to hit you with unsolicited advice later.

This is the kind of paradise we've landed in:
1️⃣ Warm up developers with cheap unlimited access at the start.
2️⃣ Get everyone hooked on agentic workflows.
3️⃣ Quietly cut everyone's limits, freeing up hardware for fat enterprise features like Orbit.

🐲 Even the Chinese are doing the same. I recently learned that in the GLM Coding Plan, during peak hours, requests are multiplied by a factor of 2-3. You can only find out about this in some tooltip on the website. Prices have effectively increased tenfold everywhere.

Now you also need to learn to use LLMs as economically as possible to get any benefit from them.