The journal · Entry 24
Don't marry the model
You don't need the fastest horse in the paddock if you're not going to the race. A new frontier model every 16 days, a lot of loud posts about the newest one, and a lot of people wondering if they're behind. You're not. Here's how to think about which model does your work, how to test when you're curious, and how to keep up in twenty minutes a month if you want to.

The question I keep hearing this month comes in a few different versions.
“We’re six weeks into adoption with Claude, but OpenAI just came out with Astra and it’s all the rave. Do I need to switch my game plan?”
“Am I going to fall behind if I stick with Claude?”
“Am I missing out on the best model by not making a switch?”
My answer is a resounding “No.”
That’s the short version. The rest of this episode is the (probably too) long one, because “no” on its own sounds like loyalty, and I promise you, it’s not. I’m going to tell you exactly where I think Astra is better. I’m going to tell you why Astra being better doesn’t matter for most of the work you’re doing. And then I’m going to show you how to decide.
The race is on
The calendar is the whole argument, so let’s start there.

Between February 5th and September 3rd of this year, the big labs shipped thirteen frontier models. One every 16 days, on average. The last two, Claude Fable 5.1 and GPT-6 Astra, landed only 48 hours apart. Anthropic alone put out four Claude 5 models in under 60 days. The Arena leaderboard, where people vote blind on which LLM answer was better, has changed its number one spot 19 times since 2023.
We’re in our horse race era
The first thing we all need to understand is that we are moving at light speed. Humans have never seen such major advancements in such a short period of time. Literally. The horse race is happening right now, and it’s going to keep happening for months, probably years, until a true “winner” is declared.
Quite a few large creators are posting about how good Astra is right now, and they’re right. It is phenomenal. The thing everyone’s talking about is computer use, which is the model working your actual screen, clicking and typing through apps the way a person would. OpenAI shipped Astra able to fill out forms, work a CRM, and run quality checks on a website by itself. Anthropic put background computer use into Cowork and Claude Code the day before Astra dropped. I’ve run both. In my opinion, ChatGPT’s computer use is better than Claude’s. It’s faster, it’s more accurate, and it uses fewer tokens getting there. But that answer could change any day. I use Claude every single day and I’ll say that out loud, because if I can’t, none of the rest of this is worth reading.
So the creators aren’t wrong. They’re describing one horse, on one stretch of one track, during this month’s race.
Next month, there will be a new track, and all new horses.
You’re not going to the race
If your work with AI is the same route every day (the follow-up email, the call summary, the numbers out of the PDF, the doc cleanup, the inbox), you do not need the fastest horse in the paddock. You need the one that pulls the cart to town without any drama. Same route every time. Not a lot of change-ups, not a lot of planning. You put the equipment on and you go.

The middle tier and smaller models do that work, and they do it well. On the Claude side that’s Sonnet 5 and Haiku 4.5. On the OpenAI side it’s GPT-5.6 Terra and GPT-5.6 Luna, their own models page points you at them for everyday and cost-sensitive work. Google’s Flash models are the same idea. They’re cheaper, they’re faster, and for tilling the field they are not measurably worse than the model everybody’s screenshotting right now.
The people posting about the biggest, baddest model are doing one of two things. Either they’re stress-testing it on purpose, throwing the hardest cases they can find at it to see where it breaks, or they have those cases for real. Code across dozens of files. Processes routed across a bunch of systems and multiple data sources. That’s a real job, and the top tier exists for it. But it’s not most people’s job.
Nobody looks at a scientist who published fifteen papers on fifteen topics in the last three months and thinks, “I haven’t even started looking into those topics, I’m so behind.” Because of course you’re not. You’re not a scientist. You have a different job. Unless your job is literally figuring out which model fits which use case for a company, you are comparing yourself to professionals at a sport you’ve never even played.
And the vast, vast, vast majority of people are not using AI for heavy coding problems. About half of U.S. adults say they never use an AI chatbot at all, for anything. Yes, the engineers are, but I’m not talking to engineers here. I’m talking to everyone using AI for what they need it for, for what these tools do that fits their life and their work. If that’s you, you’re not falling behind. You’re simply using the tool for what you use the tool for.

Don’t change your bet mid-race
Back to the question. You’re six weeks into learning one model. Should you switch because another one caught up?
No. You should not. You should keep going on the model you’re learning. You’re building instructions, habits, a feel for how it answers, a sense of when it’s wrong, judgement. That compounds, and it’s worth far more than the gap between this month’s leader and next month’s. Switching now trades what you’ve already learned for a lead that won’t hold.
Walk the paddock when you have time
When you do have the time to check out other models, this is how I’d suggest you spend it.

Grade the task first. Four questions:
- How wrong can it be?
- How many steps will this take?
- How much context does it need?
- How many times will it run?
Low on all four is small-model work. Most daily drafting, summarizing, and editing is the middle tier. Multi-file builds and anything where a wrong answer costs more than the tokens spent building it is the top tier. Anthropic’s choosing a model page and OpenAI’s model selection guide are the vendors’ own versions of this. In Claude Code, switching tiers is /model and a name, or a few clicks.
I built a whole tool that does exactly this, called Groundwork. Check it out if you want, but that’s not the point of this post.
When you’re curious, test fair. One prompt, word for word, into every model you’re trying. Decide how you’ll score before you read a single answer. Hide the names with labeled outputs: A, B, C. Score them, then look. Within a point is a tie, and the tie always goes to the cheaper one.
Where to run the same prompt side by side:
- Arena side-by-side, formerly LMArena. Free. Pick any two models, and battle mode hides the names for you. For anyone.
- OpenRouter Chat. Pay per token, pennies for a test. For people who want every model in one place and don’t mind a few knobs.
- ChatHub. Two at once on the free plan, with a browser extension. For people who want it next to their work.
If you don’t want to pick at all:
- Perplexity. Several vendors’ models behind one friendly UI. For people who want one subscription and no ceremony. The easiest on-ramp on this list.
- Poe. Same idea, more models, more settings. For people who like to tinker.
Switching is a skill you have to practice. It gets easier every time, and it’s the whole reason to keep your instructions and context in files you own. But one skill is not the whole game. You do not need the top subscription to every model. Pausing one for a month to try another is normal and encouraged, but it’s also not required of anybody who uses AI.
Watch the race if you want to
If you want to keep up, the vendor pages say what shipped and what it costs.
- Anthropic: models overview and release notes. Claude Code’s changelog moves almost daily.
- OpenAI: models and the ChatGPT model release notes.
- Google: Gemini models and the Gemini API changelog.
Five minutes, pick one digest. Simon Willison for the densest record, Latent Space’s AINews for everything daily, The Neuron if you’re newer to this. Start with one, not all three.
One more board with a warning label: OpenRouter’s rankings show which models people are paying to route traffic through. That’s adoption scores, not quality.
And if none of that sounds like a good use of a Saturday morning, skip it. The horses will still be there next month, with new names and faster feet.
The skill
I am one of those people who gets paid to know about AI and which model is best for what. So I built a skill that uses AI to help me make those calls: it grades the task, tells me what a model can and can’t do for it, and keeps the scorecard. It’s called Which model, and it’s free.
If you get paid to make these choices, but don’t get paid to understand the difference, give mine a try. Maybe it helps.
What I’d tell you
- You’re not behind. If your work is the same route every day, the middle and small models do it well and cheaper. The people posting about the biggest model are stress-testing it or have brutal problems. Most of us are tilling the field.
- Don’t change your bet mid-race. If you’re deep in learning one model, stay. The lead will flip again before what you learned stops being useful.
- Walk the paddock between races. Grade the task, test blind, write it down. Switching is a skill you practice when you have the time, not a subscription you owe anyone.
- Watch the race if you want to. And if you simply can’t be bothered, that’s allowed too.
Nobody knows which model they’ll be running the most in October. Not even the people claiming they do know. But you can know exactly how you’ll decide when it does end up mattering for you.
More from the middle of the work soon.

Sources
- OpenAI unveils GPT-6 Astra amid rising scrutiny and safety concerns, Al Jazeera, September 4, 2026
- GPT-6 Astra system card, OpenAI, September 2026
- Introducing Claude Fable 5.1 and Claude Mythos 5.1, Anthropic, September 1, 2026
- Official Claude Cowork breakdown, Vellum, September 2026
- Anthropic releases new model, Opus 5, Axios, July 24, 2026
- Frontier AI Model Releases 2026, Job Security Meter, 2026
- LLM Leaderboard History: Arena Rankings 2023 to 2026, BenchLM, September 2026
- Models overview and Choosing the right model, Anthropic
- Model configuration, Claude Code docs
- Models and Model selection guide, OpenAI
- Gemini models, Google
- Americans and AI 2026: Chatbots, Smart Devices and Views on Impact, Pew Research Center, June 17, 2026 (51 percent of U.S. adults say they never use AI chatbots; survey fielded February 17 to 23, 2026)
- 2025: The State of Generative AI in the Enterprise, Menlo Ventures, December 2025
- Avoid AI Vendor Lock-In: A Multi Model AI Strategy, AvePoint, 2026
Stay in the loop
Get the next entry.
New notes land in your inbox when there is something worth reading. No cadence, no filler.
Powered by Buttondown. One click to unsubscribe, always.