Over the past year you have probably watched AI prices fall several times without ever being told why.
Google cut Gemini Flash. OpenAI cut GPT-5.6 Luna by 80%. Anthropic cut cache reads by 75%. DeepSeek shipped a model at a fifteenth of the going rate. Each arrived as its own headline, usually framed as competitive pressure, and each was treated as a separate event.
They are mostly not separate events. They are the same event, arriving repeatedly, and the underlying cause is that the machines these models run on got dramatically cheaper to operate. Nvidia's Vera Rubin platform is the current instalment of that, and understanding it tells you more about your AI budget for the next two years than any individual pricing announcement will.
What Rubin actually is
Vera Rubin is Nvidia's current generation AI platform, now in full production, and it is not a single chip despite how these things get reported. It is seven components designed together: the Vera CPU, the Rubin GPU, an NVLink 6 switch, a ConnectX-9 network card, a BlueField-4 data processing unit, a Spectrum-6 Ethernet switch, and an integrated Groq 3 inference processor.
Nvidia calls the approach extreme codesign, meaning the components are engineered as one system rather than assembled from parts that happen to be compatible. That matters because the bottleneck in AI infrastructure has not been raw processing power for some time. It has been moving data between processors fast enough to keep them busy, which is a networking and memory problem more than a compute one.
The headline figures are up to a 10x reduction in inference cost per token compared with the previous Blackwell generation, up to 5x greater inference performance, and a 4x reduction in the number of GPUs needed to train mixture-of-experts models. Rubin-based products are shipping through the second half of 2026, with AWS, Google Cloud, Microsoft, and Oracle among the first cloud providers deploying them, alongside CoreWeave, Lambda, Nebius, and Nscale.
Why inference cost is the number that matters
There are two distinct costs in AI and the public conversation has consistently focused on the wrong one.
Training is what it costs to create a model, and it is enormous, one-off, and paid by the AI company rather than by you. It generates most of the headlines because the numbers are spectacular and the capital expenditure is visible in quarterly earnings. It has almost no direct bearing on what you pay.
Inference is what it costs to run a model each time somebody uses it. Every question your support assistant answers, every document your workflow processes, every classification your automation performs is an inference. That cost is paid every single time, forever, and it is the entire basis of what the provider charges you per token.
So a 10x reduction in inference cost per token is not an abstract infrastructure statistic. It is the input to your invoice getting ten times cheaper. To put concrete numbers on it, running a GPT-4 class model on H100 hardware cost somewhere between $8 and $15 per million output tokens. On Vera Rubin that falls to roughly $0.80 to $1.50. Providers do not pass all of that through, and they pass enough of it through that you have felt it five times this year.
How a chip becomes your invoice
The path from hardware to your bill is short, which is why the effect shows up faster than most people expect.
Nvidia sells the platform to cloud providers and AI labs. Those buyers can now serve the same number of requests using dramatically less power and hardware, which collapses their cost per request. In a market with genuine competition, that saving becomes a price cut, because the alternative is watching a competitor cut first and lose the volume.
That is exactly the dynamic we described in the AI price war and again in what AI models actually cost now. The competitive pressure is real and it is downstream of the hardware. Cheap hardware plus several well-funded competitors produces falling prices; cheap hardware plus a monopoly would not.
There is a second-order effect worth noting too. A 10x reduction against Blackwell compresses margins for cloud providers currently charging premium rates on Blackwell capacity, which forces repricing of existing inference contracts. If you are on a longer-term commitment with a provider, it is worth checking whether your rate still reflects the market, because contracts signed against previous-generation economics are now expensive by default.
What gets possible at these prices
The more interesting consequence is not that existing things get cheaper. It is that a category of thing that was never viable becomes ordinary.
At $8 to $15 per million output tokens, AI runs when someone asks it to. You send a request, you get an answer, you pay. That economic shape produces the AI most small businesses currently use: a tool that responds. At under $1.50 per million, continuous operation becomes affordable, and an AI that runs constantly rather than reactively is a different kind of product entirely.
Concretely, that means monitoring rather than answering. An agent watching your inbox continuously instead of processing it when triggered. Something reviewing every order as it arrives rather than in a nightly batch. Continuous reconciliation, always-on anomaly detection, a system that notices a customer has gone quiet without anyone running a report. Those were all technically possible before and economically silly.
This is the same shift we traced in agents moving inside the tools you already use, and the hardware is why it is happening now rather than in 2029. The products arriving in your existing software over the next year exist because somebody can finally afford to run them all day.
The catch nobody mentions
Cheaper per unit does not mean cheaper in total, and this is the part that catches businesses out with genuine regularity.
When something gets ten times cheaper, usage does not stay flat. It expands to fill the new economics, because things that were not worth automating suddenly are. The result is a bill that stays the same or grows while the unit price collapses, which feels contradictory until you look at the volume. This is a well-documented pattern in every commodity that has ever got cheaper, and AI is not going to be the exception.
The specific risk is the always-on agent. A monitoring agent that runs continuously is affordable per hour and relentless in aggregate, and the shape of its cost is completely different from a tool you invoke. We saw exactly this in the Notion credits article, where a team consumed $1,500 in a month from software doing precisely what it was configured to do.
There is also a resource question sitting underneath all of this that does not disappear with efficiency. More efficient hardware serving vastly more demand still consumes enormous amounts of power, which is the pressure we covered in AI data centre electricity costs. Efficiency gains at this scale tend to be absorbed by growth rather than banked.
How to plan around it
The practical implications are fewer than you might expect, and the main one is about posture rather than action.
Assume prices keep falling and plan accordingly. Anything that looks marginal on cost today is likely to be comfortable within a year, which argues against long fixed commitments at current rates and in favour of staying flexible. If a vendor offers a substantial discount for a two-year lock-in, weigh it against the fact that the market rate will almost certainly be below your locked rate before the term ends.
Revisit the ideas you rejected on cost. Most small businesses have a mental list of automations that were investigated and abandoned because the numbers did not work. A good portion of that list is now viable, and nobody sends you a notification when the economics of a rejected idea change. Going back through it once a year is a genuinely high-return habit.
And put ceilings on anything continuous before you enable it, not after. The cost structure of an always-on agent is the one thing here that can genuinely surprise you, and the surprise is entirely preventable with a spend cap and a calendar reminder to check the first full month. Cheap per unit multiplied by running forever is still a number worth knowing in advance.
Sources
- Nvidia Newsroom: NVIDIA Vera Rubin opens agentic AI frontier
- Nvidia Newsroom: NVIDIA kicks off the next generation of AI with Rubin, six new chips
- Tom's Hardware: Nvidia launches Vera Rubin NVL72, promises up to 5x greater inference performance and 10x lower cost per token
- TechRepublic: Nvidia introduces Vera Rubin, an AI platform designed to slash costs