For three years the assumption underneath every AI business case has been that the price only goes one way. Build it now, and it gets cheaper while you own it. In the first fortnight of September that assumption broke — in both directions at once.

01 — What happened

On 3 September 2026, OpenAI released GPT-6 Astra, described in its own changelog as “our most capable model, built for the hardest end-to-end work.” It is priced at $10 per million input tokens and $50 per million output tokens. The flagship it succeeds, GPT-5.6 Sol, is listed at $4 and $20. The new best model costs two and a half times the old best model.

That is unusual. Frontier list prices have generally held or fallen at each generation. This one rose, and the surcharges rose with it: prompts over 272,000 input tokens are billed at double the input and cache rates, and a “fast mode” that roughly doubles throughput costs double the standard price.

Three other things happened in the same fortnight, and they went the other way.

Anthropic cut the cost of repetition. Claude Fable 5.1, released 1 September, kept the same $10 and $50 headline rates as Fable 5 but cut cache reads by 75%, to $0.25 per million tokens — 2.5% of the input price, where every other model in the range charges 10%. Anthropic’s own estimate is that typical workloads become around 25% cheaper, and “complex coding and highly agentic tasks” up to around 45% cheaper, with no change to the headline number at all.

A scheduled increase was cancelled. Claude Sonnet 5 launched at an introductory $2 and $10, with a rise to $3 and $15 scheduled for 1 September. The documentation now states that the increase will not occur and the introductory price is the standard price.

The cheap tier got cheaper and started pricing like a utility. DeepSeek released V4.1-Flash on 10 September, replacing two earlier models, with the stated reasoning that the model “lets us serve more users at a lower cost. We’re passing the savings on to you.” Its flash tier is listed at $0.30 per million input tokens and $1.20 output at peak, and exactly half that off-peak — peak defined as two fixed windows on weekdays, UTC.

So in fourteen days: the ceiling rose sharply, the floor fell, an announced rise was withdrawn, and one vendor began charging by time of day.

02 — Why it matters

The headline everyone will take away is “AI got more expensive.” That reading is wrong, and acting on it is expensive in a different way.

What actually happened is that the price of a token and the price of a job came apart. They used to move together, which is why a per-million-token figure was a usable planning number. It no longer is. Anthropic’s change lowered the real cost of a class of work by up to 45% while leaving the advertised price untouched. OpenAI raised its advertised price by 150% for a model most businesses will never need for the work they actually run.

There is a quieter version of the same problem in Anthropic’s documentation: models from Claude 4.7 onwards use a newer tokenizer that produces roughly 30% more tokens for the same text. The price per token is only half of a price. The other half is how many tokens your text becomes, and that now varies by model generation.

If you are comparing vendors on the number in the headline, you are comparing the wrong number.

03 — The economics

Take an illustrative case: a business handling 10,000 customer enquiries a month, each one read and answered by a model. Assume 3,000 input tokens per enquiry — the message, the customer record, the instructions — and 700 tokens of output. That is 30 million input and 7 million output tokens a month. At each vendor’s published list rates:

  • GPT-5.6 Luna ($0.20 / $1.20): about $14 a month.
  • Claude Haiku 4.5 ($1 / $5): about $65.
  • Claude Sonnet 5 ($2 / $10): about $130.
  • GPT-5.6 Sol ($4 / $20): about $260.
  • GPT-6 Astra or Claude Fable 5.1 ($10 / $50): about $650.

Identical work. A 45-fold spread. For corroboration on the order of magnitude, Anthropic’s own published worked example puts 10,000 support-ticket conversations on Haiku 4.5 at roughly $37.

Now apply the structural levers, which most businesses never touch.

Caching. Say 2,500 of those 3,000 input tokens are stable — the same instructions and the same product knowledge on every enquiry. On Fable 5.1, 25 million cached tokens read at $0.25 cost $6.25 instead of $250. The input line falls from $300 to about $56, and the monthly total from $650 to about $406.

Batching. Work that does not need an answer in the next second — overnight enrichment, classification, summarising yesterday — runs at 50% on both OpenAI’s and Anthropic’s batch APIs. That takes the same Astra workload to about $325.

Timing. On DeepSeek’s schedule, moving that batch job outside two weekday windows halves it again.

None of those three levers is a negotiation, a contract or a discount. They are consequences of how the system was built. A business that routes every enquiry to the most capable available model, uncached and in real time, is paying something like forty-five times a business doing the same job deliberately. That gap does not show up as an AI problem. It shows up as gross margin.

04 — What businesses should do

If you buy AI inside software you already pay for — your CRM’s summarising, your helpdesk’s draft replies — none of this changed anything for you this month. Your vendor absorbed it or will reprice you later. Do nothing.

If you run, or are costing, anything that calls a model per transaction, do three things this quarter.

First, work out your cost per completed job, not per million tokens. Take your real monthly volume, measure actual input and output tokens on twenty representative cases, and multiply. Most businesses have never done this arithmetic and are surprised in both directions.

Second, separate the work by difficulty rather than standardising on one model. Classifying an enquiry, extracting a name and address, or drafting a routine reply does not need a frontier model, and the published price gap between tiers is now wide enough that using one is a deliberate choice with a number attached. Reserve the expensive tier for the genuinely hard minority — and check whether the cheap tier fails on your actual cases before assuming it will.

Third, look at what repeats. If the same instructions and the same reference material are sent on every call, caching is the single largest lever available and it is now worth substantially more than it was a month ago.

If your volume is low — a few hundred operations a month — stop. At that scale the difference between the cheapest and dearest model is rounding. Use the best one and spend your attention somewhere it matters.

05 — Where technology fits

Honestly: for most businesses, nothing needs building because of this.

Token pricing only becomes a line on your profit and loss when AI runs inside a process at volume — every enquiry, every order, every document. Below that threshold this is interesting rather than actionable.

Where it does cross the threshold, the thing worth building is not “an AI system.” It is the routing and integration layer around the model: something that decides which tier of model handles which job, keeps the stable context stable so it can be cached, and pushes non-urgent work into a batch queue. That is ordinary systems work — integration, automation, a well-structured data layer — and it is the part that determines the bill. As we argued in where AI actually creates economic leverage, the value sits in the workflow, not the model. The cost now sits there too.

It also matters that this layer is yours. If model choice is hard-coded into one vendor’s SDK across your codebase, a 150% price rise is not a decision you get to make. If it sits behind a connected system you control, it is a configuration change. The same argument applied to the voice layer when answering the phone became a commodity: keep the expensive, fast-moving component replaceable.

06 — RevenueStack’s take

Our view: the important development is not that the best model got dearer. It is that capability and cost stopped being the same axis.

For three years “which model” was a proxy for “how good” and “how much” at the same time, so choosing the best one was defensible. That has ended. Vendors now compete on billing mechanics — cache multipliers, batch discounts, off-peak windows, tokenizers — as hard as they compete on benchmarks. The gap between a naive implementation and a deliberate one is wider than the gap between any two vendors.

Which means the AI cost question has quietly become an architecture question. The businesses that do well out of the next two years of price movement will not be the ones that picked the right model. They will be the ones whose systems were structured so that the model is a setting, not a foundation — the same dull discipline that decides whether automating follow-up pays for itself.

Prices will move again. They moved four times in a fortnight. Build for that, not against it.

Sources

  1. Changelog | OpenAI API — OpenAI, 2026-09-03
  2. GPT-6 Astra Model | OpenAI API — OpenAI, 2026-09-03
  3. Pricing | OpenAI API — OpenAI, 2026-09-17
  4. Introducing Claude Fable 5.1 and Claude Mythos 5.1 — Anthropic, 2026-09-01
  5. Pricing — Claude Platform Docs — Anthropic, 2026-09-17
  6. DeepSeek-V4.1-Flash: Smarter, Faster, More Efficient — DeepSeek, 2026-09-10
  7. Models & Pricing | DeepSeek API Docs — DeepSeek, 2026-09-10