Note · Cost Reduction Through Automation

Your AI Plan Got Smaller. Five Stepsto Stop Depending on One Model

David He, FounderOctober 1, 20265 min read

AI plans now buy less of the best models. Five steps I took to make my context and tools portable, so switching models does not mean starting over.

The $200 AI plans now buy less of the best models, and the generous era for those models is probably ending. The fix I found was not a cheaper model - it was making my context and tools portable, so switching models does not mean starting over.

The plans got smaller

If you hit an AI usage limit this month, it wasn't you.

On September 29, OpenAI brought back its $200 Pro plan after pausing sales on September 10. The new plan buys about half as much of its top model, GPT-6 Astra, as the old one did. The same day it added a $500 tier with a faster mode priced at six times regular Astra. In July, Anthropic took Fable, its top model, out of the $20 plan and capped Max users at half their limit on it.

Alex Grankin walks through the numbers in The New Era of AI Has Just Begun. He cites a June SemiAnalysis estimate that the $200 Claude Max plan held about $8,000 of usage at API prices, and ChatGPT Pro about $14,000. His guess at why it is ending now is plain: both companies filed to go public in June.

He also makes the point that matters for budgets. The price of the very best model is not falling. The price of ordinary work is. By the chart he uses, work that cost about $9 on Fable 5 in June costs roughly 30 cents on a smaller model today.

What I did when I hit the wall

I hit the wall in early September, when Fable 5.1 and Astra were new. Fable would hit its limit within about an hour on a hard task, and then I waited four hours for it to reset. A few days later I was out of usage on Astra too. A week of usage went in about a day.

I didn't buy another $200 plan. I moved most of my work to cheaper models, and for most of it the results were close. (I wrote up the first weekend of that in An $80-a-Month Coding Tool Matched What I Was Paying $200 a Month For.)

The models weren't what made that possible. Five things I'd set up for other reasons were. Here they are, so you can do them on purpose.

Five steps to stop depending on one model

Step 1: Get your context out of the vendor

Put who you are, your goals, your projects and your preferences in plain files on your own disk. Then add one instruction file that tells any agent where to look.

Mine is a second brain I call OpenBrain. Its instruction file says, in effect: if someone asks about David, read the profile first. On September 24 I opened Kimi Code in that folder and typed "ok what do you know about me?" It read my profile and answered. I did not have to explain myself to a new tool.

Step 2: Make your skills portable, then trim them

Most coding tools now read a folder of skills - small instruction files for jobs you do often. Claude Code's is ~/.claude/skills. ZCode's is ~/.zcode/skills. My agent copied mine across in one pass: 223 in one folder, 207 in the other.

Then the catch. ZCode loaded every skill and tool I own on every call, about 130,000 tokens before it did anything, until I cut it to 24,000. I hit the same thing with a local model and ended up archiving most skills, because loading all of them filled the context and most were not needed. Copy everything; load only what the job needs.

Step 3: Do a week of real work in a second tool

Not a benchmark. The tasks you would otherwise do in your main tool.

I used ZCode with GLM 5.3, Kimi Code with Kimi K3, and DeepSeek V4.1 Flash. Some of the client hours I logged in September were worked in ZCode. That is the test that counts: work someone is paying for, done in the other tool.

Step 4: Find where the cheap model breaks before you depend on it

Give it a job where you already know the answer.

I asked Kimi K3 whether 12 AI-generated photos were real. It said yes to all 12, so that job stays on a frontier model. Grankin's video points the same way from the other side: on a private government test, one open model that looked 5 points behind on a public coding benchmark was 16 points behind. Public scores flatter these models. Your own known-answer test does not.

Step 5: Run the biggest local model your machine holds, and time it

Mine is Qwen 3.8 27B on a 36 GB Mac. It writes about 8 tokens a second. Before it writes anything, it spends close to a minute reading the prompt each turn, because the tool sends it over 8,000 tokens of instructions every time.

Grankin found the same model beating Opus 4.6 on a landing page, at about an hour per try. That makes it my floor, not my daily driver. If you are weighing hardware for this, I wrote up the tradeoffs behind my own order in Why I Ordered a $10,000 Mac Studio for Local AI.

Cloud or local?

My order today:

  1. A frontier subscription for the hard problems.
  2. Hosted open models for daily work. GLM 5.3 is my favorite right now.
  3. Local when everything else is gone.

Grankin's own advice is not to buy a Mac to replace your subscription yet, and to treat it as a long-term investment. I agree on the timing. I expect Astra-level models on a decent desk machine eventually. I can't prove when.

If you run a team on AI subscriptions, the same five steps apply to the team.

Anyway.

The plans will keep changing. Your context is the part you own.

Which AI tool would be hardest for you to leave tomorrow, and what's keeping you there?

More notesnewest first

Working on something like this?

Bring the app or the process to a free 15-minute call. I will tell you what I would look at first, and whether I am the right person for it.