Ledger/Aftershock

Filed 4 days after Opus 5.5

Rent the model, own the routing

Opus 5.5 matches Fable 5.1 on most work at 40% of its token price, the second big repricing since June. The asset worth keeping is your rule for which task gets which model.

Source Anthropic, Claude Opus 5.5 5 min read

Opus 5.5 is out, and Anthropic puts it at the level of Fable 5.1 on most work, for $4 per million input tokens and $20 per million output. Fable 5.1 costs $10 and $50. That’s the second time since June that top-tier work got much cheaper inside two months.

Here is my take. The top model is something you rent. What you own is the rule that decides which task gets which model, and the rule that holds up is simple: give the cheaper model the work where a miss shows. So on launch day, audit your defaults, test the one new capability your product depends on, and rebuild nothing. People who rebuild around every launch will disagree, and they’re right about exactly one case.

At the Fable 5 launch in June I expected a cheaper model to reach that level within months. It took six weeks: Opus 5 came within half a percent of Fable 5’s peak on one coding benchmark at max effort, for half the cost per task.

The callRevisit Mar 22, 2027

By 22 Mar 2027, Anthropic sells a model at the level of Opus 5.5 for less than $4/$20.

OpenWhen the date comes, this box says held or missed, with the link that proves it.

The expensive choice is the one nobody made

In my setup, most of the top-tier runs came from defaults, well before any debate about which model is best. When a task doesn’t pick its model, it tends to ride on whatever the session that launched it runs, and no one reviews it.

When I had my agents price 229 real subagent runs across my projects, most of them ran on the top tier:

165 of 229

subagent runs that ran on Opus or Fable. Most ran on the model of the session that launched them.

A median call on Opus 5 cost two and a half times one on Sonnet 5, $0.147 against $0.058, and most of the runs paying it ran on the model of the session that launched them. So search your agent configs and scripts for every call that doesn’t name a model, and make each task pick one.

Route by what a miss costs

The cheaper model forgets more than it invents. That shapes where it belongs. An omission shows up when you check the output against a known answer. In an open-ended review, nobody sees the thing that’s missing. So give the cheaper model the work where a miss is cheap or checkable, and keep the top model for anything that turns into judgment nobody will recheck.

I tested this on my own work with a blind A/B of Opus 5 against Sonnet 5 on eight tasks, for $2.65. On exact extraction from data files it was a tie, 10 against 10. On the six tasks that needed judgment (turning notes into rules, hunting inconsistencies, making judgment calls), Opus came out ahead every time, and Sonnet’s gap was mostly things it left out. The judge was Opus 5, which is a bias, so the next run uses a judge that isn’t Opus.

Give the cheaper model the work where forgetting shows.

Write your own table with three rows. Judgment (reviews, specs, anything that becomes your own words) goes to the top model. Well-specified work (extracting, counting, checking against a spec) goes to the middle tier. Mechanical checks go to the smallest. Then keep eight tasks with known answers. Anthropic says Sonnet 5.5 and Haiku 5.5 follow in the coming weeks. When they land, rerun the eight and move rows.

A launch changes things you didn’t touch

Launch-day risk rarely sits in the benchmark. It sits in defaults that moved and references that went stale.

Anthropic’s 40% saving over Opus 5 is measured at default settings, and the default effort dropped from high to medium. At the same effort, Opus 5.5 also tends to think more per turn than Opus 5. So if you set effort back to high, the saving can shrink: measure price per task at the effort you actually use. Thinking can no longer be switched off, and forcing a specific tool call now returns an error. None of that is in the headline.

In my own setup, the same model alias ran two different models depending on whether a job started in the desktop app or the terminal, because the terminal client was out of date. Two scripts carried the model ID typed by hand, so they needed their own edit, and they’ll need one at every launch until they read the alias too.

Where the rebuild crowd is right

Rebuilding around one model means paying the whole switching cost (rewired prompts, retested outputs, retrained habits) for a capability that gets cheaper within weeks. The exception is a launch that opens a capability your product lives on. Opus 5.5 reads values off dense charts and screenshots much more precisely than Opus 5. If your product runs on visual checks, test that this week. Test the capability. Leave the stack alone.

The one-hour audit

Copy this into your launch notes.

Launch-day audit
1. Read the migration notes before the benchmarks. Write down what changed by default (effort, thinking, tool calls) and what now errors.
2. Compare price per task. The token price leaves out how long the model thinks.
3. Change the alias. Then search your code for the model ID typed by hand, and set effort explicitly wherever a task needs depth.
4. Check which model actually ran: the model name in the logs, on every surface you launch agents from.
5. Rerun your eight known-answer tasks.
6. Move rows in your routing table: which tasks go up a tier, which go down.
7. Test the one new capability your product lives on. Rebuild nothing else.
8. Write the decision down with the date and what would make you reopen it.

The next model is weeks away. The table will still be yours.

The work behind the Ledger

See the work
2024CrewA shared home that runs itself. Split what the flat spends, divide what it has to do, and stop having the same argument about money every month.Live · iOS & Android2026FarroPlan the week, and never write a shopping list again. Farro turns the meals you chose into a list sorted by aisle and scaled to the number of people at the table.Live · iOS & Android
Lorenzo Alejandro Montes

Lorenzo Alejandro Montes

I make consumer apps and take them from the first interview to the store listing. Crew and Farro are live on iOS and Android.

Lorenzo Alejandro Montes Milan, Italy RSS feed
04 / Aftershock SCRL 0.00 Milan --:--