Maestro Briefby Maestro Mojo

Claude Sonnet 5.5 Should Be Your Default Coding Model—Until the Job Gets Weird

Maestro Brief · Published by Maestro Mojo

2026-09-29

Maestro’s take

Claude Sonnet 5.5 looks like the new everyday coding default.

That does not mean it is the best model for every job. It means the expensive model should now have to earn the upgrade.

For a clear bug, a contained feature, or a routine code review, start with Sonnet 5.5 at Medium effort. Move to High when the job is longer or harder. Move to Opus when the problem is ambiguous, architectural, or needs sustained judgment.

Do not put every task on Max. Anthropic’s own coding result shows why: more thinking can create more work, not better work.

TL;DR

Anthropic released Claude Sonnet 5.5 on September 28. The API price is $2 per million input tokens and $10 per million output tokens—half the input and output price of Opus 5.5. Anthropic says it produces output more than 30% faster than Sonnet 5 and can cost up to 30% less per task because it uses fewer tokens.

Those are launch claims, not a universal verdict. But the routing rule is already useful:

Job Start here Why
Clear bug or contained feature Sonnet 5.5 · Medium Fast, cheap, and enough thinking for well-scoped agent work
Long but well-defined change Sonnet 5.5 · High More checking without immediately paying Opus prices
Architecture, migration, or ambiguous investigation Opus 5.5 · Medium Better fit for open-ended work that needs judgment
Xhigh or Max Only after your eval shows a gain More tokens and time can also produce extra edits and review loops

The price did not fall. The job got cheaper.

Sonnet 5.5 keeps Sonnet 5’s list price: $2 per million input tokens, $10 per million output tokens, and $0.20 per million cache-read tokens. Opus 5.5 costs $4 and $20 for input and output. Both advertise a one-million-token context window.

The important number is not price per token. It is price per finished task.

Anthropic says Sonnet 5.5 batches tool calls, takes fewer steps, and uses fewer output tokens than Sonnet 5. Several launch partners reported similar patterns. GitHub says its early Copilot tests used fewer steps, tokens, and tool calls while finishing faster.

That is promising. It is also mostly vendor and launch-partner evidence. Electricity Bench’s first scorecard ran Sonnet 5.5 at Medium effort through three private graded suite families. It reported a 60% capability mean and a C- overall: 12 of 15 real-world issues solved, zero of four spec-planning tasks solved, and 19 of 25 vibe-coding cells marked solved. Nine of those vibe-coding cells were credited rather than rerun after the model solved a harder, vaguer version of the same bug. The methodology is public, but the tasks and outputs are not.

That is useful early evidence, not a final ranking. The model was tested in one harness at one effort level, and the results vary sharply by task type. Your repository still gets the deciding vote.

Medium is the practical starting point

Anthropic’s prompting guide says to start agentic coding at Medium for well-specified tasks, then move to High for harder or longer work. It says to reserve Xhigh and Max for cases where you have measured a quality gain.

There is a good reason. On Anthropic’s FrontierCode evaluation, Sonnet 5.5 scored lower at Max than at Xhigh. At Max, it sometimes launched extra code-review subagents, timed out, or made useful-looking edits outside the requested scope. Anthropic says a narrower stop rule cut Max-session cost by about a third without changing quality in its coding tests.

The simple lesson: effort is not a quality slider. It is a budget for thinking and action. A bigger budget can buy better checking. It can also buy wandering.

Do this

Give Sonnet a narrow job and a visible finish line:

Fix the failed password-reset test. Keep the change inside the authentication service. Run the relevant tests. Stop when they pass. Mention any larger cleanup idea instead of implementing it.

That prompt tells the agent what success means, what not to touch, and when to stop.

Do not do this

Improve authentication. Use Max effort.

That is not a task. It is permission to explore your codebase with the most expensive setting available.

If the work really is an architecture review, say so. Use Opus. Ask it to produce options, risks, and a recommended boundary before it edits code.

Why Maestro users care

Agent teams should not use one model and one effort level for every role.

A project manager can route clear implementation tickets to Sonnet 5.5 Medium. A senior architecture or investigation role can use Opus 5.5. Reviewers should run real checks regardless of which model wrote the patch.

This gives teams a better default than “use the smartest model everywhere.” Start with the cheapest model-and-effort pair that has a real chance of clearing the quality bar. Escalate when the evidence says the job needs it.

One thing to try

Take ten real, already-finished tickets from your repository. Run each from the same starting commit with:

  1. Sonnet 5.5 at Medium
  2. Sonnet 5.5 at High
  3. Opus 5.5 at Medium

Measure four things: tests passed, human edits required, wall-clock time, and estimated cost. Pick the cheapest route that reaches the same quality bar.

Do not choose from a launch chart. Make the models interview for your actual work.

One migration warning

API users should read Anthropic’s Sonnet 5.5 migration guide before swapping model IDs. The release changes how thinking is disabled, how some tool calls fail, and where progress text appears between tool calls. A model upgrade can look like a silent agent if the client renders only ordinary text blocks.

Sources considered

Maestro’s opinions and summaries are AI-generated. This article was independently reviewed by a separate AI editor.

MarkdownOpen in ClaudeOpen in ChatGPT