A furniture shop switched on Meta Business Agent in July because it was free to build and test, and because the messages had genuinely become unmanageable. Two hundred a week across WhatsApp and Instagram, most of them asking the same four things: do you deliver to this postcode, is this in stock, how long is the lead time, can I see it in another colour.
The agent handled them. It was, by every measure the owner cared about, a success.
On 1 August the free window ended and every one of those replies started costing money. Nothing about the setup changed, no announcement arrived in the inbox on the day, and the first real signal will be an invoice that does not match anything the business budgeted for. This is not a large amount of money. It is an amount of money nobody has calculated, which is a different problem and a more common one.
What actually changed
Meta Business Agent launched globally on 3 June 2026 across WhatsApp, Instagram, and Messenger, offering AI that answers customer questions, recommends products from a catalogue, books appointments, qualifies leads, and completes transactions. A free build-and-test window ran from 1 July to 31 July 2026, which is when a great many small businesses set theirs up.
From 1 August, replies became chargeable and the billing model changed shape entirely. The old approach was per-message pricing. The new one is token-based, at $2.00 per one million tokens, where tokens are the units of text the AI processes and generates. Meta puts typical consumption at roughly 20,000 to 25,000 tokens per delivered AI message.
That works out at approximately 4 to 5 cents per reply. It is a small number and it is also a number with a multiplier attached, which is the thing worth paying attention to. A business handling a handful of enquiries a week will not notice this. A business that automated its messaging precisely because the volume was overwhelming has, by definition, volume.
What it costs in real numbers
The useful exercise is running your own numbers rather than reading a general reassurance, so here is the arithmetic in a form you can apply to your own business in about a minute.
At roughly 4.5 cents a reply, a hundred AI replies a month costs about $4.50, which is genuinely nothing. Five hundred replies a month is around $22.50, which is a line item you would notice but not one that changes any decision. Two thousand replies a month, which is where the furniture shop lands once you count both platforms and the fact that a single customer conversation is usually several replies rather than one, costs about $90 a month.
The multiplier people forget is that conversations are not messages. A customer asking about delivery, then about lead time, then about a colour option has generated three billable replies inside what feels like one interaction. When you estimate your volume, count exchanges rather than customers, and expect the real figure to be somewhere between three and five times your instinct.
Against that, the comparison worth making is not zero. It is what the same messages cost when a person answers them. Ninety dollars a month buys a small fraction of an hour a week of someone's time at any realistic wage, and the two thousand replies would have taken considerably more than that. The economics remain strongly favourable. The point is not that this is expensive, it is that it is no longer free and almost nobody has updated their mental model.
The bundling detail that matters
There is a genuine simplification buried in this change that got very little coverage, and it slightly improves the picture rather than worsening it.
Previously, businesses running AI on WhatsApp paid on two separate lines: a charge for the AI processing and a separate conversation or message delivery fee. The new token rate bundles both into one. You are no longer paying for the intelligence and the transport as distinct items, which makes the cost of a reply a single number you can actually reason about.
That matters more than it sounds for a small business, because the old two-line structure was the reason most owners could not answer the question "what does one automated reply cost me". The answer involved two rate cards, a conversation-window definition, and a category system that determined which rate applied. Now it is tokens times two dollars per million, and you can hold that in your head.
It is worth noting that this is the same direction the wider market has moved, and not always in your favour. Simpler pricing is easier to understand and also easier to raise, because a single number can be adjusted without renegotiating a structure. We went through this dynamic in what AI models actually cost now, where promotional rates ending is the most reliable pricing event in the industry.
Is it still worth running?
For most businesses using it seriously, yes, and the reasoning is not really about the four cents.
The value of an agent answering WhatsApp at eleven at night is not that it saves you the labour cost of answering it yourself. It is that the message gets answered at all. A customer asking whether you deliver to their postcode at eleven on a Sunday is a customer who will ask a competitor by Monday morning if nobody replies, and the cost of that is not measured in cents. This is the argument we made at length in WhatsApp automation for small business and it survives a price change of this size comfortably.
Where it becomes a genuine question is at high volume with low conversion. If you are handling several thousand replies a month and most of those conversations do not lead anywhere, you are paying to have a machine talk to people who were never going to buy. That is not a reason to switch it off, but it is a reason to look at what the agent is being asked and whether some of it should be answered before a conversation starts, by a clearer delivery page or a visible stock indicator.
The scenario where the answer genuinely changes is if you set this up during the free window without a real use case, because it was free and seemed worth trying. Plenty of businesses did exactly that. An agent running on autopilot answering questions nobody was struggling with is now a small recurring cost with no return attached, and the correct response is to turn it off rather than to optimise it.
How to control what it costs
Because billing follows token consumption rather than message count, the levers available to you are about length and necessity rather than about volume alone.
Shorter replies cost less, directly and proportionally. An agent configured to answer warmly and at length burns more tokens per message than one configured to answer the question and stop. There is a real tension here, because terse automated replies read as cold, and the balance worth striking is conversational but not padded. The instruction to keep answers brief and specific is both a cost measure and usually a customer experience improvement, which is a rare combination.
The larger lever is handling the highest-frequency questions before they become conversations. If a quarter of your agent traffic is people asking whether you deliver to a given area, a visible delivery checker on your product page removes those conversations rather than automating them. The cheapest reply remains the one that was never needed, and that principle does not change regardless of what a token costs.
Finally, know your baseline before you try to improve it. Check what your first full month at the new rate actually came to, because that number is the only one that matters and almost every business is estimating rather than reading it. A great deal of anxiety about AI costs dissolves on contact with an actual invoice, and the small number of cases where it does not are the ones genuinely worth acting on.
The pattern this fits into
The specific numbers here matter less than the shape, because this will happen again with something else you use.
The sequence is now familiar enough to predict. A large platform launches an AI capability, offers it free during a build-and-test window, acquires a base of businesses who wire it into their operations, and then begins charging once switching away has become inconvenient. Nobody was misled, the dates were published, and the free window was clearly labelled as a window. It still catches people, because adoption decisions get made under one set of costs and reviewed under another, if they get reviewed at all.
WhatsApp now has more than 200 million small business users, and Meta's paid messaging revenue reached a $2 billion annual run rate as of December. That is the scale at which a four cent charge becomes a serious business line, and it explains the timing without requiring anyone to have behaved badly.
The habit worth building from this is small: when you adopt something during a free period, write the end date somewhere you will actually see it, and put a reminder a fortnight before to check what it will cost. That takes a minute and it is the difference between choosing to keep paying for something and discovering that you already are.