HomeInsightsAI Tools
AI tools · 8 min read

One AI Model Got 80% Cheaper. Another Got 50% Dearer. Same Month

On 30 July 2026, OpenAI cut GPT-5.6 Luna from $1/$6 to $0.20/$1.20 per million tokens, an 80% reduction, and cut Terra 20%. On 1 September, Claude Sonnet 5 promotional pricing ends and rises from $2/$10 to $3/$15, a 50% increase on both. If you run AI in production, one of those is a windfall and the other is a bill you have not budgeted for.

Two announcements, five weeks apart, pointing in opposite directions.

On 30 July, OpenAI cut the price of GPT-5.6 Luna by 80%, from $1 and $6 per million input and output tokens down to $0.20 and $1.20. Terra came down 20% at the same time, from $2.50 and $15 to $2 and $12. The flagship Sol tier stayed where it was. This came three weeks after GPT-5.6 launched, which is an unusually short interval for a price cut of that size.

On 1 September, Claude Sonnet 5 goes the other way. Its promotional pricing of $2 and $10 per million tokens ends on 31 August, and standard pricing of $3 and $15 takes effect the next day. That is a 50% increase on both input and output, arriving on a Tuesday, for everyone who built something on it during the promotional period.

Neither company did anything unusual. Together the two moves explain almost everything difficult about budgeting for AI right now.

What actually moved

It is worth having the numbers in one place, because the direction is easier to reason about than the magnitude.

GPT-5.6 Luna, the fastest and cheapest tier, went from $1 input and $6 output per million tokens to $0.20 and $1.20. For a high-volume workload running on Luna, the monthly bill fell by four fifths overnight with no action required. OpenAI also gave 100,000 researchers free access to these models through 2027, which is a customer acquisition move aimed at the people who will be specifying tools inside companies for the next decade.

Claude Sonnet 5 moves from $2 and $10 to $3 and $15 on 1 September. A business spending €600 a month on Sonnet 5 today is looking at €900 for the same usage, and nothing about the model changes. This was always the plan, the promotional rate was clearly labelled as promotional, and it is still a 50% increase landing on a business that may not have registered the end date when it built the workflow.

For context on where the top of the market sits, Claude Opus 5 is at $5 and $25, and Grok 4.5 at $2 and $6. The spread between the cheapest and most expensive credible option is now more than twentyfold on output tokens, which is a wider range than most businesses realise they are choosing within.

Why prices moved in both directions

These are not contradictory strategies. They are the same market logic applied to different tiers, and understanding that helps you predict what happens next.

At the low end, the competitive pressure is brutal and getting worse. Cheap, fast, competent models are arriving from many directions at once, including open-weight releases that anyone can host. There is very little to defend at that tier, because if your budget model is not the cheapest adequate option, workloads migrate away with almost no friction. Cutting Luna 80% is not generosity, it is the price of staying in a segment where switching costs are near zero.

At the mid and upper tiers the dynamic reverses. Businesses running production workloads on a specific model have built prompts around its behaviour, tuned outputs to its quirks, and quietly accumulated a switching cost that has nothing to do with the API. Anthropic launching Sonnet 5 at a promotional rate and normalising it later is a standard and entirely legitimate way to acquire those workloads. The promotional rate did its job. Now it ends.

The pattern to expect going forward is more of exactly this: relentless deflation at the bottom, stable or rising prices in the middle where switching is painful, and premium pricing at the frontier. This is the same dynamic we traced through the AI price war, and the trajectory has not changed.

The promotional pricing trap

The Sonnet 5 increase is worth dwelling on, not because anyone was misled, but because of how reliably this catches businesses that were paying attention at the wrong moment.

The sequence is always the same. A model launches at an attractive introductory rate. You evaluate it against alternatives, and the price is part of why it wins. You build a workflow around it, spend a few weeks getting the prompts right, and it works. Months pass. The workflow becomes something the business quietly depends on. Then the promotional period ends, and you are paying 50% more for something you can no longer easily move, because moving means retesting everything you tuned.

Nobody did anything wrong here. The dates were public. But the decision to adopt was made at one price and the consequence arrives at another, and by then the cost of reversing the decision has grown considerably. This is the mechanism by which introductory pricing works, and it works because it is genuinely difficult to plan around.

The specific lesson is that when you evaluate a model on price, you should be evaluating the price you will pay in a year rather than the price you pay in month one. If a rate is labelled promotional, treat the standard rate as the real number in your comparison. That single adjustment would have made the Sonnet 5 change a non-event for anyone who applied it.

Not sure what your AI stack will cost after the next pricing change? A €49 audit reviews your models, workloads, and exposure.

What AI actually costs now

The honest picture for a small business is that the model bill is rarely the number that matters, and the obsession with per-token pricing is mostly misplaced.

Most small businesses running AI in production spend somewhere between €50 and €500 a month on model calls. At that scale, an 80% cut on one tier and a 50% rise on another largely cancel out into noise. The variable that actually dominates the bill is not price per token, it is how much unnecessary work you are asking for, which is why the effort settings on Claude Opus 5 matter more than any price announcement.

The bigger cost is almost always integration, maintenance, and the time somebody spends when a workflow breaks. A business paying €200 a month in model costs and four hours a month of somebody's attention is spending far more on the attention. Optimising the €200 while ignoring the four hours is a common and expensive mistake, and price announcements encourage it because the token number is the visible one.

Where price genuinely matters is at volume. If you process tens of thousands of items monthly, a twentyfold spread between the cheapest and most expensive credible model is real money and warrants proper evaluation. Below that threshold the correct answer is usually to use whatever works reliably and spend your attention elsewhere.

How to protect yourself from this

The structural protection is not choosing the right model. It is making sure that choosing wrong is cheap to fix.

The most valuable thing you can build is the ability to switch models without rewriting your business logic. In practice that means the model call sits behind a single point in your system rather than being scattered across twenty different automations, each with its own hardcoded prompt. When the model reference lives in one place, a price change becomes a configuration edit. When it lives in twenty places, a price change becomes a project you keep postponing while paying more each month.

The second protection is knowing your own numbers well enough to react. Most businesses cannot answer which workflow generates the majority of their token spend, which means they cannot tell whether a price change matters or not. That is a one-hour piece of work that turns every future pricing announcement from anxiety into arithmetic.

The third is treating promotional rates as temporary in your planning, always. Write the end date somewhere you will see it. When you evaluate on price, use the post-promotional number. This costs nothing and removes an entire recurring category of unpleasant surprise.

What to do in the next two weeks

If you use Claude Sonnet 5 in production, work out what a 50% increase does to your monthly bill before 1 September rather than after. For most small businesses the answer will be that it is annoying and affordable, and knowing that is worth the ten minutes. For a few it will be significant enough to justify evaluating alternatives while there is still time to do it calmly.

If you run high-volume work on GPT-5.6 Luna, you have simply become cheaper and there is nothing to do except notice. It is worth checking whether the new pricing makes something viable that you previously ruled out on cost, because an 80% reduction moves a lot of ideas from too expensive to obvious.

And regardless of which models you use, spend an hour finding out where your token spend actually goes. Not the total, the distribution. Almost every business discovers that one or two workflows account for most of it, and that at least one of those is doing more work than the task requires. Fixing that is worth more than any price change either company announced this summer, and unlike their pricing, it is entirely within your control.


Sources

Quick answers

Common questions.

Want this in your business?

The €49 audit shows you exactly which automations would pay back fastest in your specific operation.

€49 entryFull AI audit + strategy call included

Reserve your auditNo commitment. No contracts. Just clarity.