Back to the journal

The journal · Entry 24

Don't marry the model

You don't need the fastest horse in the paddock if you're not going to the race. A new frontier model every 16 days, a lot of loud posts about the newest one, and a lot of people wondering if they're behind. You're not. Here's how to think about which model does your work, how to test when you're curious, and how to keep up in twenty minutes a month if you want to.

Illustration of a racetrack starting gate at sunrise with a row of small robots wearing numbered saddle cloths ready to run, while a woman sits on the fence rail in the foreground with her back to them, reading a newspaper with a coffee

The question I keep hearing this month comes in a few different versions.

“We’re six weeks into adoption with Claude, but OpenAI just came out with Astra and it’s all the rave. Do I need to switch my game plan?”

“Am I going to fall behind if I stick with Claude?”

“Am I missing out on the best model by not making a switch?”

My answer is a resounding “No.”

That’s the short version. The rest of this episode is the (probably too) long one, because “no” on its own sounds like loyalty, and I promise you, it’s not. I’m going to tell you exactly where I think Astra is better. I’m going to tell you why Astra being better doesn’t matter for most of the work you’re doing. And then I’m going to show you how to decide.

The race is on

The calendar is the whole argument, so let’s start there.

Timeline of thirteen frontier model releases from February 5 to September 3, 2026, one dot per release colored by vendor, with the last two, Fable 5.1 and GPT-6 Astra, marked as 48 hours apart

Between February 5th and September 3rd of this year, the big labs shipped thirteen frontier models. One every 16 days, on average. The last two, Claude Fable 5.1 and GPT-6 Astra, landed only 48 hours apart. Anthropic alone put out four Claude 5 models in under 60 days. The Arena leaderboard, where people vote blind on which LLM answer was better, has changed its number one spot 19 times since 2023.

We’re in our horse race era

The first thing we all need to understand is that we are moving at light speed. Humans have never seen such major advancements in such a short period of time. Literally. The horse race is happening right now, and it’s going to keep happening for months, probably years, until a true “winner” is declared.

Quite a few large creators are posting about how good Astra is right now, and they’re right. It is phenomenal. The thing everyone’s talking about is computer use, which is the model working your actual screen, clicking and typing through apps the way a person would. OpenAI shipped Astra able to fill out forms, work a CRM, and run quality checks on a website by itself. Anthropic put background computer use into Cowork and Claude Code the day before Astra dropped. I’ve run both. In my opinion, ChatGPT’s computer use is better than Claude’s. It’s faster, it’s more accurate, and it uses fewer tokens getting there. But that answer could change any day. I use Claude every single day and I’ll say that out loud, because if I can’t, none of the rest of this is worth reading.

So the creators aren’t wrong. They’re describing one horse, on one stretch of one track, during this month’s race.

Next month, there will be a new track, and all new horses.

You’re not going to the race

If your work with AI is the same route every day (the follow-up email, the call summary, the numbers out of the PDF, the doc cleanup, the inbox), you do not need the fastest horse in the paddock. You need the one that pulls the cart to town without any drama. Same route every time. Not a lot of change-ups, not a lot of planning. You put the equipment on and you go.

Illustration of a small robot pulling a wooden cart of vegetables and a milk can down a country lane lined with lavender, a woman in a straw hat walking beside it, while far behind them tiny robots race around a track in front of a crowded grandstand

The middle tier and smaller models do that work, and they do it well. On the Claude side that’s Sonnet 5 and Haiku 4.5. On the OpenAI side it’s GPT-5.6 Terra and GPT-5.6 Luna, their own models page points you at them for everyday and cost-sensitive work. Google’s Flash models are the same idea. They’re cheaper, they’re faster, and for tilling the field they are not measurably worse than the model everybody’s screenshotting right now.

The people posting about the biggest, baddest model are doing one of two things. Either they’re stress-testing it on purpose, throwing the hardest cases they can find at it to see where it breaks, or they have those cases for real. Code across dozens of files. Processes routed across a bunch of systems and multiple data sources. That’s a real job, and the top tier exists for it. But it’s not most people’s job.

Nobody looks at a scientist who published fifteen papers on fifteen topics in the last three months and thinks, “I haven’t even started looking into those topics, I’m so behind.” Because of course you’re not. You’re not a scientist. You have a different job. Unless your job is literally figuring out which model fits which use case for a company, you are comparing yourself to professionals at a sport you’ve never even played.

And the vast, vast, vast majority of people are not using AI for heavy coding problems. About half of U.S. adults say they never use an AI chatbot at all, for anything. Yes, the engineers are, but I’m not talking to engineers here. I’m talking to everyone using AI for what they need it for, for what these tools do that fits their life and their work. If that’s you, you’re not falling behind. You’re simply using the tool for what you use the tool for.

Stat card reading 51 percent of U.S. adults say they never use an AI chatbot, not ChatGPT, not Gemini, not Copilot, not Claude, with the line You are not behind, sourced to Pew Research Center, February 2026

Don’t change your bet mid-race

Back to the question. You’re six weeks into learning one model. Should you switch because another one caught up?

No. You should not. You should keep going on the model you’re learning. You’re building instructions, habits, a feel for how it answers, a sense of when it’s wrong, judgement. That compounds, and it’s worth far more than the gap between this month’s leader and next month’s. Switching now trades what you’ve already learned for a lead that won’t hold.

Walk the paddock when you have time

When you do have the time to check out other models, this is how I’d suggest you spend it.

Illustration of a horse barn aisle with a small robot looking out over each stall door, blank name plates on the doors, and a woman walking down the middle with a clipboard and pencil

Grade the task first. Four questions:

  • How wrong can it be?
  • How many steps will this take?
  • How much context does it need?
  • How many times will it run?

Low on all four is small-model work. Most daily drafting, summarizing, and editing is the middle tier. Multi-file builds and anything where a wrong answer costs more than the tokens spent building it is the top tier. Anthropic’s choosing a model page and OpenAI’s model selection guide are the vendors’ own versions of this. In Claude Code, switching tiers is /model and a name, or a few clicks.

I built a whole tool that does exactly this, called Groundwork. Check it out if you want, but that’s not the point of this post.

When you’re curious, test fair. One prompt, word for word, into every model you’re trying. Decide how you’ll score before you read a single answer. Hide the names with labeled outputs: A, B, C. Score them, then look. Within a point is a tie, and the tie always goes to the cheaper one.

Where to run the same prompt side by side:

  • Arena side-by-side, formerly LMArena. Free. Pick any two models, and battle mode hides the names for you. For anyone.
  • OpenRouter Chat. Pay per token, pennies for a test. For people who want every model in one place and don’t mind a few knobs.
  • ChatHub. Two at once on the free plan, with a browser extension. For people who want it next to their work.

If you don’t want to pick at all:

  • Perplexity. Several vendors’ models behind one friendly UI. For people who want one subscription and no ceremony. The easiest on-ramp on this list.
  • Poe. Same idea, more models, more settings. For people who like to tinker.

Switching is a skill you have to practice. It gets easier every time, and it’s the whole reason to keep your instructions and context in files you own. But one skill is not the whole game. You do not need the top subscription to every model. Pausing one for a month to try another is normal and encouraged, but it’s also not required of anybody who uses AI.

Watch the race if you want to

If you want to keep up, the vendor pages say what shipped and what it costs.

Five minutes, pick one digest. Simon Willison for the densest record, Latent Space’s AINews for everything daily, The Neuron if you’re newer to this. Start with one, not all three.

One more board with a warning label: OpenRouter’s rankings show which models people are paying to route traffic through. That’s adoption scores, not quality.

And if none of that sounds like a good use of a Saturday morning, skip it. The horses will still be there next month, with new names and faster feet.

The skill

I am one of those people who gets paid to know about AI and which model is best for what. So I built a skill that uses AI to help me make those calls: it grades the task, tells me what a model can and can’t do for it, and keeps the scorecard. It’s called Which model, and it’s free.

If you get paid to make these choices, but don’t get paid to understand the difference, give mine a try. Maybe it helps.

What I’d tell you

  1. You’re not behind. If your work is the same route every day, the middle and small models do it well and cheaper. The people posting about the biggest model are stress-testing it or have brutal problems. Most of us are tilling the field.
  2. Don’t change your bet mid-race. If you’re deep in learning one model, stay. The lead will flip again before what you learned stops being useful.
  3. Walk the paddock between races. Grade the task, test blind, write it down. Switching is a skill you practice when you have the time, not a subscription you owe anyone.
  4. Watch the race if you want to. And if you simply can’t be bothered, that’s allowed too.

Nobody knows which model they’ll be running the most in October. Not even the people claiming they do know. But you can know exactly how you’ll decide when it does end up mattering for you.

More from the middle of the work soon.

Illustration of a tilled field at golden hour, a woman kneeling to plant seedlings from a tray, a small robot standing beside her with a feed bucket on its head, and far in the background an empty racetrack grandstand with one lone robot still running the far turn

Sources

Signed,

Stay in the loop

Get the next entry.

New notes land in your inbox when there is something worth reading. No cadence, no filler.