MaShop/Journal/Industry/The Labs Halved Their Prices. Your Invoice Did Not…
● IndustrySeptember 23, 2026
Read · 5 min
ai pricing · openai

The Labs Halved Their Prices. Your Invoice Did Not Notice

Token prices fell by half in one week. The software a shop actually buys is billed per seat, per resolution and per credit, so the saving stops upstream.

Key takeaways
  • Between 21 and 22 September 2026 three labs repriced flagship models, and two of them cut list prices by 20 to 58 percent.
  • Those ai price cuts apply to tokens, and almost nothing a small shop buys is billed in tokens.
  • Per seat pricing, per resolution pricing and monthly credit bundles absorb the saving before it reaches your invoice.
  • Cache reads on one model fell 60 percent while its headline input price fell 20 percent, so repeated context got cheaper faster than anything else.
  • Comparing cost per token across model generations is unsafe, because the number of tokens in the same paragraph changed.
  • The number worth asking a vendor for is cost per completed job, not cost per million tokens.

Your helpdesk tool did not get cheaper this week. The models underneath it did, by a lot, and the gap between those two sentences is the whole story.

On 21 and 22 September 2026, three of the largest model providers published new flagship prices within about thirty hours of each other. If you build software, that week mattered. If you run a shop with a support inbox, a product catalogue and an invoice from a SaaS vendor, the honest answer is that your costs are probably unchanged, and it is worth understanding exactly why before you go asking for a discount.

What actually changed in the last week of September?

Two providers cut published api pricing and one launched at rates in line with its previous generation. Anthropic moved Claude Opus 5.5 to $4 per million input tokens and $20 per million output tokens, down from $5 and $25 on Opus 5, according to the Claude platform pricing documentation. That is exactly 20 percent off both sides. TechCrunch reported the same output figure on the day of release, along with the company's claim that the model is faster to run because the compute required to serve it dropped.

OpenAI went further. VentureBeat's report on the GPT-6 Sol and Luna launch puts gpt-6 sol at $2 and $10 per million input and output tokens, against $4 and $20 for the model it replaces. gpt-6 luna landed at $0.10 and $0.50, against $0.20 and $1.20. That is a 50 percent cut on input and a 58.3 percent cut on Luna's output. An OpenAI representative told the publication these are permanent prices rather than promotional ones, which is a meaningful distinction given how often introductory rates quietly expire.

xAI shipped grok 4.7 a day earlier. The model sits at $2 per million input tokens and $6 per million output tokens with cached input at $0.50, on a 500,000 token context window, per the llm-stats model page.

ModelReleasedInput per millionOutput per millionReplaces
claude opus 5.522 Sep 2026$4.00$20.00Opus 5 at $5 and $25
gpt-6 sol22 Sep 2026$2.00$10.00GPT-5.6 Sol at $4 and $20
gpt-6 luna22 Sep 2026$0.10$0.50GPT-5.6 Luna at $0.20 and $1.20
grok 4.721 Sep 2026$2.00$6.00Launched alongside Grok 4.6

Read as a group, the week did something unusual. Price moves of this size normally arrive with a weaker model attached, the cheap tier you accept because it is cheap. Here the cuts came with the flagship tier, which is why the coverage treated it as competitive pressure rather than a product segmentation exercise.

Why did your invoice stay flat?

Because you are almost certainly not buying input tokens and output tokens. You are buying resolved conversations, seats, generated images or a monthly bundle of credits, and each of those wrappers sets its own price that has no contractual link to what the wrapper pays upstream.

The clearest public example is outcome pricing in customer support. Fin publishes a flat $0.99 per outcome with a fifty outcome monthly minimum on standalone helpdesk integrations, where an outcome means a resolution, a procedure handoff or a disqualification, and a qualification costs $9.99. Nothing in that page moves when a model provider reprices. The customer pays $0.99 whether the underlying generation cost a tenth of a cent or a full cent.

Diagram comparing falling token prices at the labs against the per seat, per resolution and credit bundle pricing a small shop is billed for

Put real numbers on it. Anthropic's own documentation works a support example: roughly 3,700 tokens per conversation across 10,000 tickets on Haiku 4.5 comes to about $37, which is $0.0037 a ticket. Take a shop handling 300 support conversations a month, of which 200 end without a human. At outcome pricing that is $198. The token cost underneath those 200 conversations is about 74 cents, or under half a percent of the bill. Halve the token price and the vendor's input cost falls by 37 cents. Nobody is going to rewrite a price list over 37 cents, and nobody should expect them to.

That arithmetic is assembled from two published sources rather than lifted from either, and the model in the example is a cheap one, so treat it as an order of magnitude rather than an audit of any specific vendor. The shape holds regardless. When inference is a rounding error inside a service price, inference price cuts are a rounding error inside your invoice.

Where the cut does reach a small shop

There are three places it genuinely lands, and it is worth knowing which apply to you.

The first is direct api use. If your site calls a model itself, for a product description generator or an internal catalogue tool, the new rate applies from the moment you switch model identifiers. Nothing else has to happen. This is the case for anyone whose storefront was generated with the code in their own repository rather than rented from a platform, which is the arrangement we describe on our own pricing page and the reason the distinction is worth caring about.

The second is prompt caching, and it moved more than the headline. Cache hits on Opus 5.5 are priced at 0.05 times the base input rate instead of the usual 0.1, which puts cached input at $0.20 per million against $0.50 on the previous model. That is a 60 percent reduction against a 20 percent headline cut. If you send the same 40,000 token catalogue or the same long policy document with every request, the cached portion of your bill fell three times faster than the uncached part. OpenAI's new models carry a 90 percent cached input discount on the same principle.

The third is batch work. Both providers keep a 50 percent discount for asynchronous batches, which on Opus 5.5 means $2 and $10 per million. Rewriting nine hundred product descriptions overnight is exactly the workload that qualifies, and it is the single cheapest way for a catalogue owner to buy generation. If you have been doing that job interactively, one at a time, in a chat window, you have been paying four times the necessary rate for a task with no deadline.

Is a token the same size it used to be?

No, and this is the trap in every price comparison table published this month. Anthropic's documentation states plainly that Claude 4.7 and later models use a newer tokenizer that produces approximately 30 percent more tokens for the same text, while Sonnet 4.6 and earlier use the previous one.

Work through what that does. A price quoted per million tokens is only comparable across models that chop text the same way. If an older model costs $3 per million input tokens and a newer one costs $2, the newer one looks 33 percent cheaper, but if the same product page becomes 30 percent more tokens under the new tokenizer, the real cost per page is $2.60 against $3.00. Still cheaper, by 13 percent rather than 33. Compare across that boundary without adjusting and you will overstate your saving by a factor of two and a half.

Within a single generation the comparison is safe. Opus 5 and Opus 5.5 share a tokenizer, so the 20 percent cut is a true 20 percent. The care is needed when you migrate across families, which is also when people are most tempted to quote the headline number to themselves. We went through the related problem of models disappearing under you in our piece on what happens when a vendor deprecates the model you built on.

Note

Data residency has its own price. Pinning inference to a single geography carries a 1.1 times multiplier on every token category, and regional endpoints on the partner clouds add a 10 percent premium. If you are an EU seller who has committed to keeping processing inside a region, half of a 20 percent cut is spent before you start.

Reading your own bill: the four shapes it can take

Nearly every AI line item a small business pays falls into one of four shapes, and the shape decides whether a price cut upstream can ever reach you.

Per seat is the most common and the most insulated. You pay a fixed monthly figure per person with usage included, so the vendor's margin expands when its costs fall and your bill does not notice. A copilot add-on at $35 per user per month is the pattern.

Per outcome ties the price to a business event rather than to compute, which is better for budgeting and equally insulated from token pricing. Per credit sits in between: credits usually map to work units, and a vendor that wants to compete can quietly make a credit buy more work. Watch for that rather than for a price drop, because it is the change vendors actually make.

Metered pass through is the only shape where the cut arrives automatically, and it is rare outside developer tooling. If your invoice shows input tokens and output tokens as separate lines, you are in this category and you already saved money this week without doing anything.

Card listing the four billing shapes an AI line item takes on a small business invoice and which of them react to falling token prices

How do you tell whether a vendor passed the cut on?

Ask for cost per completed job and compare it against the same figure from three months ago. That single question cuts through every pricing page, because it is the only unit that means the same thing before and after a repricing.

A few follow ups make the answer useful. Ask which model the product runs on today and whether it changed in the last quarter, because a vendor that silently moved to a cheaper tier has already taken the saving. Ask whether your plan's included volume went up, since raising the allowance is the most common way a vendor passes value along without touching the headline price. Ask whether caching is used on your account, given how much of the recent movement sits in cached input rather than in the base rate.

If you are running the numbers on whether an AI task pays for itself at all, the method matters more than the rate. We set out a way to test that on a single task in the piece on measuring return from one job rather than a whole strategy, and the underlying question of what owners actually end up paying for AI has not changed shape this week even though the input costs did.

What does a catalogue rewrite cost at the new rates?

Less than most owners assume, and the gap between the cheapest and dearest way to do the same job is now roughly forty times. Here is the arithmetic for a job almost every online seller eventually faces: rewriting 900 product descriptions.

Assume each one sends about 800 input tokens, which is the existing description plus the product attributes plus your instructions, and returns about 250 output tokens. That is 720,000 input tokens and 225,000 output tokens for the whole catalogue. Those are my assumptions, not a published figure, and your own numbers will shift with how much context you attach, but the ratios below hold for any reasonable variation.

How you run the jobInput costOutput costTotal for 900 descriptions
gpt-6 luna, interactive$0.07$0.11about $0.18
GPT-5.6 Luna, the old rate$0.14$0.27about $0.41
claude opus 5.5, interactive$2.88$4.50about $7.38
claude opus 5.5, batch api at half price$1.44$2.25about $3.69

Two things jump out of that table. The cheap model now does the whole catalogue for the price of a coffee, and the expensive model does it for the price of a sandwich. At these levels the cost of the tokens has stopped being the deciding factor for a one person business. What decides it is output quality, how many passes you need before the copy is usable, and how long you spend reviewing it.

That is the real consequence of the week. When generation cost falls below the noise floor of a small budget, the constraint moves to review time, and review time does not fall when token prices do. A shop that rewrites 900 descriptions and reads none of them has not saved anything, it has published 900 unverified pages. The cost that matters is the hour you spend checking, and it is the same hour it was in August.

What this week does not mean

It does not mean AI got cheap. The same documentation that lists a 20 percent cut also lists a fast serving mode at $8 and $40 per million, double the standard rate, and a web search tool billed at $10 per 1,000 searches. Capability tiers are still being added above the price line as fast as the line comes down. A shop that upgrades to the newest model, turns on search and pins inference to a region can easily spend more per task after a price cut than before it.

It also does not mean you should switch anything today. The cost of changing a working setup, retesting outputs and rewriting prompts is measured in your hours, and your hours are the scarcest input in a business of one or two people. A 37 cent saving does not buy an afternoon. The sensible response to ai price cuts of this size is to note them, check whether your own direct api calls are pointed at the new identifiers, and otherwise carry on.

The people who should act are the ones running batch jobs and the ones sending large repeated context. For them the saving is real, immediate and worth ten minutes of configuration. For everyone else, the useful takeaway is structural rather than financial: you now know why your ai software bill is insulated from the loudest numbers in the industry, and what to ask when you want a share of them.

"These are permanent prices, not promotional or introductory pricing."OpenAI representative, quoted by VentureBeat, 22 September 2026

That quote is the part worth keeping. Permanent list prices give a vendor no excuse to hold a rate, which means the pressure now sits on the layer between the labs and you. Whether it moves is a question of competition in your category, not of arithmetic in theirs. The practical test arrives at your next renewal. Vendors rarely cut a price mid term, and they very often improve an allowance at renewal to keep a customer who has just asked an awkward question. So the moment to raise cost per completed job is four to six weeks before your contract rolls, with last quarter figures in hand, rather than on the day a lab publishes a headline.

Discussion 0

0 / 4000Your email address is not displayed with your comment.
No comments are published yet.

Explore — related articles.

Build something. Move your work forward.

Start with a software project or an agent task. Describe the result you need, review the work and keep control of your connected accounts.

Open the workspace →