Four frontier models landed in the last six weeks.

Every AI app now greets you with a dropdown full of fresh names, and it's tempting to just take whatever the default is.

The model you pick decides the quality of your output, how long you wait, and how fast you burn through your plan's limits.

The good news is that you only need one simple rule to get this right.

Today we're building your routing rule (which model for which job), and what each model costs you 👇️

Your support queue gets a head start every morning.

Viktor reads overnight tickets, tags them by product area, summarizes the patterns, and posts a brief in #support. The agent picks up the queue already triaged. The PM sees recurring requests rolled up by Friday.

Why your model selection matters

Quick vocabulary lesson first. 

A model is the actual brain behind the chat window. 

ChatGPT, Claude, and Grok are the apps; GPT-5.6, Sonnet 5, and Grok 4.5 are the brains you can swap in and out. 

A frontier model just means one of the most capable brains available anywhere.

For a while, picking one rarely made a difference. Then July happened:

  • Anthropic shipped Claude Sonnet 5 on July 1, a fast model that gets close to its big sibling Opus 4.8 at a much lower price.

  • xAI shipped Grok 4.5 on July 8, pitched by Elon Musk as "Opus-class, but faster and lower cost."

  • OpenAI shipped the GPT-5.6 family on July 9: three models named Sol, Terra, and Luna, each at a different power and price level.

The gaps between the options have never been this wide, in both ability and price. 

Pick a model that's too weak for the job and the output disappoints you. Pick one that's too strong and you wait longer, spend more, and hit your usage cap by lunch. 

“The Two-Speed Rule”

Here's a mental model that cuts through every confusing name.

Every AI lab sells the same two things: a workhorse and a heavyweight.

The workhorse is fast, cheap, and very good. It handles 90% of what a normal workday throws at it.

The heavyweight is slower and several times more expensive. It exists for the hardest 10% of problems.

Map the names onto that and it simplifies your choice:

  • Claude: Sonnet 5 is the workhorse. Opus 4.8 is the heavyweight.

  • ChatGPT: your options are Instant, Thinking, and Pro. Instant is the workhorse. Thinking and Pro bring in the heavier GPT-5.6 models (Sol is the big one).

  • Grok: Grok 4.5 is xAI's play to be a heavyweight at workhorse prices, aimed especially at coding and agent tasks (an agent is an AI that takes actions for you, like writing files and running multi-step jobs, instead of just chatting).

  • Gemini: Google's models are the specialists for enormous inputs. More on that in a second.

Which model for which job?

Now the routing. 

Tasks you should send to the workhorse (Sonnet 5, ChatGPT on Instant):

  • Email drafts and replies

  • Summaries of articles, threads, and meeting notes

  • Rewrites, tone changes, and translations

  • First drafts of almost anything

  • Quick factual lookups and brainstorms

Tasks you should send to the heavyweight (Opus 4.8, ChatGPT on Thinking or Pro):

  • Analysis you'll make a real decision from

  • Contracts, proposals, and anything high-stakes you get one shot at

  • Complex spreadsheet logic or financial models

  • Long multi-step tasks you're handing to an agent

  • Anything the workhorse already tried and fumbled

Two special cases:

  • A giant pile of documents? Use Gemini. It can hold about ten novels' worth of text in its head at once (the most of any major model), so it's the pick for "read all of this and answer my questions."  

  • Need to know what's happening right now? Grok taps live data from X, which makes it unusually good at "what are people saying about this today."

If you remember one line from this edition, make it this one: draft with the workhorse, decide with the heavyweight. 

Start cheap and fast. Escalate only when the answer matters or the cheap answer disappoints.

What it actually costs you

Labs charge developers per token (a chunk of text; a million tokens is roughly 750,000 words). They're the price tags on each brain, and they explain your plan's usage limits.

Here's the current menu, shown as the cost to generate a million tokens of output:

Model

Cost per million tokens of output

Grok 4.5

$6

GPT-5.6 Luna

$6

Sonnet 5

$10 (introductory pricing through the end of August, then $15)

GPT-5.6 Terra

$15

Opus 4.8

$25

GPT-5.6 Sol

$30

The heavyweight costs three to five times the workhorse, every single time you use it.

On a monthly subscription, that same math hides inside your usage limits: run everything on the heavyweight and you’ll hit your usage limit several times sooner.

Selecting the right model for the job gets more usage out of the same monthly payment, giving you more for your money.

Set your defaults today

Your five-minute homework:

  1. Find the dropdown. In Claude it's the model menu in the chat box. In ChatGPT it's the Instant / Thinking / Pro selector. Actually click it once so you know where it lives.

  2. Set the workhorse as your default. Sonnet 5 or Instant. This is the right home base for almost everyone.

  3. Run one real task on both tiers. Take something you did this week, run it on the workhorse and the heavyweight, and compare. You'll calibrate your own taste in two minutes.

  4. Build the escalation habit. When an answer underwhelms you or the stakes go up, re-run the same prompt one tier higher.

Want the routing done for you? Today’s prompt builds a personal cheat sheet matched to your actual week. Paste it into any AI chat 👇️

I want a one-page cheat sheet telling me which AI model to use for each task I do.

The AI tools and plans I have: [e.g., Claude Pro and a free ChatGPT account]

The tasks I use AI for in a typical week: [list them, e.g., drafting client emails, summarizing meeting notes, reviewing vendor contracts, building a monthly report]

For each task, tell me:

1. Workhorse or heavyweight: whether a fast everyday model or a slower top-tier model is the right call, using the specific model names available on my plans

2. Why, in one sentence

3. A red flag that means I should escalate that task to the stronger model

Format it as a table I can screenshot and keep. Assume I'm not technical.

Save the output somewhere you'll see it. After a week of using it, the routing becomes instinct and you can throw the cheat sheet away.

The toughest room in advertising. Ad Studio just cleared it.

Major brands just took over a Times Square billboard, and every ad was built in Hightouch Ad Studio. Sweetgreen, HelloFresh, and Tripadvisor cleared the toughest room in advertising. Create, edit, and launch on-brand ads in a few clicks.

Your roundup of the latest model releases and updates from the biggest AI labs.

  1. Gemini 3.5 Pro is reportedly days away. Third-party reports point to a July 17 launch for Google's next flagship, rebuilt from scratch and rumored to double the context window to 2 million tokens (enough to hold roughly 20 novels' worth of text at once). Google hasn't confirmed the date or the specs, so treat it as a strong rumor. If it lands as described, the "giant pile of documents" job above gets a much bigger truck. (TechTimes)

  1. Meta shipped Muse Spark 1.1, its new flagship for agent work. Released July 9 from Meta's Superintelligence Labs, the model is built to run multi-step tasks and even coordinate teams of sub-agents (helper AIs that split up a big job and work in parallel). Meta also opened its first public model API (the plumbing developers use to build on a model) with aggressively low pricing, and you can try the model today in the Meta AI app's Thinking mode. One more serious menu to choose from. (Fortune)

  1. Anthropic is talking to Samsung about building its own chip. Reports on July 2 say Anthropic is in early talks with Samsung to manufacture a custom AI chip, following OpenAI's Jalapeño chip reveal with Broadcom a week earlier. Every major lab is now designing its own silicon to run models more cheaply instead of paying Nvidia's prices. The plans could still fall through, but the direction matters for you: cheaper chips are how the prices in that table above keep falling. (TechCrunch

Advertise with Build with AI

Get in front of an audience of professionals using AI day-to-day: founders, engineers, operators, and product builders.

Interested in advertising? Respond to this email for rates and details.

Until next time,

William Ryan

Editor-in-Chief @ Build with AI

PS: Follow me on X for daily updates and AI workflows.

Keep Reading