
The date everyone budgeted for, and the number that did not move.
If you built a 2027 budget in July, there is a line in it you can delete.
Claude Sonnet 5 launched at US$2 per million input tokens and US$10 per million output tokens, described as introductory pricing that would end on 31 August 2026. On 1 September the rate was scheduled to rise to US$3 and US$15, a 50% increase three weeks away.
That increase has been cancelled. From Anthropic's pricing documentation, retrieved 12 August 2026:
"The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur."
Nobody has to switch anything. The rate is the one you have been paying since launch. Your invoice is unchanged; your forecast is now wrong. And the decisions people made in anticipation of that rise are worth revisiting, because a few of them were wrong in an expensive direction.
Sonnet 5 is now cheaper than Sonnet 4.6
That is the part with a routing decision attached to it.
Claude Sonnet 4.6 still costs US$3 per million input tokens and US$15 output. Sonnet 5 is US$2 and US$10 — now the standard rate, with no end date attached.
The newer model is one-third cheaper per token than the one it replaced. Per-token price tells you half of it; jobs are what you actually pay for, and the next section cuts that gap to about 13% once the tokenizer is accounted for. The direction still holds for most workloads: a stack still pinned to Sonnet 4.6 because it was the known quantity is now paying more for the older model.
The tokenizer catch, and why it no longer wipes out the gain
Anyone who has read our earlier Sonnet 5 write-up knows the complication. Claude 4.7 and later models use a newer tokenizer, and Anthropic's model documentation, retrieved 12 August 2026, states plainly that it "produces approximately 30% more tokens for the same text."
So a per-token price cut is not the whole story. You are paying less per token, across more tokens. Here is the arithmetic on an identical document:
- Sonnet 4.6: US$3 per million tokens × 1.0× token count → US$3.00 per unit of work
- Sonnet 5: US$2 per million tokens × roughly 1.3× token count → about US$2.60 per unit of work
("Unit of work" means one fixed document put through each model. The multiplier is the token count each model turns that document into, so the two figures are directly comparable.)
Sonnet 5 comes out around 13% cheaper for the same work, once the tokenizer is accounted for.
Now run it at the cancelled September rate: US$3 across about 1.3× tokens is US$3.90 per unit of work, roughly 30% more expensive than the model it replaced. That was the future everyone was budgeting for. It is not happening.
Two honest caveats, because this is a number people will paste into spreadsheets. Anthropic says approximately 30%, and the increase depends on content and workload shape — dense structured text behaves differently from prose. So treat 13% as a direction, not a figure. And measure your own workload before planning around either number.
What this changes in ringgit and Singapore dollars
Anthropic bills in US dollars, so none of this is denominated in ringgit. The rate is settled; your exposure still moves with the currency.
At the mid-market rate of RM4.09 to the dollar on 11 August 2026, Sonnet 5's US$2 and US$10 work out to roughly RM8.20 per million input tokens and RM41 per million output. The cancelled September rate would have put those at about RM12.30 and RM61. For Singapore readers, at S$1.28 to the dollar on the same date, the same rates are about S$2.56 and S$12.80 per million, against S$3.84 and S$19.20 under the increase.
Put that on a real workload. Take a Klang Valley retailer running 10,000 WhatsApp and Shopee support conversations a month. Assume 3,000 input tokens and 700 output tokens per conversation — that is our assumption for illustration, not a vendor figure, so substitute your own. It gives 30 million input and 7 million output tokens: US$60 plus US$70, about US$130 a month, or roughly RM530.
Under the September rate the same volume would have cost US$195, about RM800. The cancellation is worth around RM265 a month to that retailer, every month, with nothing required of them.
The Singapore shape is different but the direction is identical. A professional-services firm putting 4,000 client email threads a month through Sonnet 5 — assume 5,000 input and 1,200 output tokens each, again our illustration rather than a measured figure — uses 20 million input and 4.8 million output tokens. That is US$40 plus US$48, about S$113 a month, against roughly S$169 under the cancelled rate.
Two levers cut it further, though rarely on the same workload — batch suits anything asynchronous, caching suits anything live and repetitive:
- The Batch API halves it: US$1 input and US$5 output for anything that does not need an answer this second. Overnight document processing, bulk classification, report generation.
- Prompt caching cuts repeated reads to a tenth, but the write costs extra. Reads cost US$0.20 per million against US$2 standard; writes cost US$2.50 (five-minute cache) or US$4 (one-hour cache).
That write premium decides whether caching is worth it, and the answer differs by cache lifetime. On the five-minute cache, one read back already pays for the write: US$2.70 against US$4.00 uncached. On the one-hour cache you need two, because the write costs double. A cache nobody reads is the only guaranteed loss.
Three things worth rechecking this week
1. Anything still routed to Sonnet 4.6. It is now the more expensive option, and the older one. There may be a good reason to stay: a test suite you have already validated against 4.6, or a workload the new tokenizer inflates unusually. But "it was cheaper" is no longer one of them.
2. Any business case that modelled a September step-up. If you wrote headroom into a 2027 forecast for a 50% rise, that headroom is now unallocated. Better to reclaim it deliberately than to discover it in a variance report.
3. Whether you are on the right model at all. A cheaper Sonnet does not make Sonnet correct for your task. The routing question is unchanged in shape, but the range above Sonnet has moved: Haiku 4.5 at US$1/US$5 for volume — and on the older tokenizer, so its per-token rate is the like-for-like one — Sonnet 5 for production workloads, Opus 5 at US$5/US$25 for hard reasoning, and Fable 5 at US$10/US$50 now sitting where Opus used to, at the top of the lineup.
The wider point about pricing announcements
A cancelled increase is not a discount. Nothing arrived; something merely stopped being scheduled. It is worth naming that clearly, because "prices cut" is how this will be reported, and it is not what happened.
What did happen is more useful to know: the rate you are already paying is now the rate you can plan on. If you deferred a Claude deployment to see where pricing settled, you now have your answer.
The number worth watching next is not Sonnet's. It is the tier above it — Fable 5 at US$10/US$50 now sets the ceiling of the range, and that is the line most likely to move a 2027 forecast.
Want your own numbers? The Claude cost calculator takes a work profile and returns a monthly estimate in ringgit.
Choosing a model rather than budgeting one? The Opus 5 versus Sonnet 5 comparison covers the routing logic in depth.
Comparing vendors? Claude versus Gemini for business sets these rates against Google's, where output pricing bundles thinking tokens and the arithmetic works differently.
Free consultation
Not sure which model your workload should be on?
We size Claude deployments for Malaysian and Singaporean teams — model routing, caching strategy and a monthly cost you can put in a budget.
For plan-level pricing rather than API rates, the Claude pricing Malaysia guide covers Pro, Team and Enterprise tiers in ringgit.
Source of record: Anthropic pricing documentation, platform.claude.com/docs/en/about-claude/pricing, retrieved 12 August 2026. All figures verified against the vendor page, not the announcement email. FX at mid-market rates for 11 August 2026 (frankfurter.dev, open.er-api.com).

