HomeInsightsAI Tools
AI tools · 8 min read

A 2K AI Video Model Just Went Open Source. Here Is What Changes

MiniMax H3, also called Hailuo 3.0, launched on 31 July 2026 and generates 4 to 15 second videos at up to 2K resolution with native stereo audio in a single pass, priced at $0.13 per output second. On 3 August MiniMax released the model weights publicly, which is unusual in a field where the strongest video models have stayed closed. For a small business, the practical change is that product video stopped being a production budget and became a per-second cost.

A four-person homeware brand wants a fifteen-second clip of its new ceramic mug, turning slowly in morning light, with the sound of coffee being poured and a single line of voiceover. Two years ago that was a shoot: a photographer, a small studio, a half-day, a separate sound pass, and an invoice somewhere between eight hundred and three thousand euros depending on who picked up the phone.

On 31 July 2026, MiniMax released a model that does that for roughly two euros.

The model is called H3, or Hailuo 3.0, and the specification is worth stating plainly because the numbers are the story. It takes text, images, video, and audio together as a single context. It returns 4 to 15 seconds of video at up to 2K resolution. The audio is generated in the same pass as the picture, in stereo, including dialogue, effects, and room tone. And on 3 August, three days after launch, MiniMax published the weights.

What H3 actually does

The way to understand H3 is not as a text-to-video generator, which is what most people picture, but as a model that treats every kind of input as the same kind of thing. You can drop a product photograph, a short clip of the motion you want, and a recording of a voice into one prompt box, write a sentence describing the result, and get back a finished piece of video with sound.

That capability set includes some specific things small businesses have wanted for a long time and rarely got. Instruction-guided editing means you can describe a change to an existing clip rather than regenerating from scratch. Text and brand rendering means the model is built to put legible words and logos into a frame, which has historically been the single most reliable way to spot AI video, because the text came out as melted nonsense. Video-to-video motion transfer means you can take the camera movement from one clip and apply it to different subject matter.

None of that is magic and none of it removes the need for taste. What it removes is the fixed cost of the attempt. The reason most small businesses have almost no video is not that they never wanted any. It is that every idea carried a several-hundred-euro entry fee, so only the safest ideas ever got made, and the safe ideas are rarely the ones that work.

The single-pass audio matters more than the resolution

Most coverage of H3 leads with 2K. The resolution is not the interesting part. Plenty of tools produce sharp frames.

The interesting part is that the sound is generated together with the picture rather than bolted on afterwards. Anyone who has assembled AI video before knows the workflow this replaces: generate the visual, take it into an editor, find or generate audio separately, then spend an irritating stretch of time nudging the two into alignment so that the cup lands on the table at the same moment you hear it land. That alignment step is where most of the actual working time went, and it is where amateur results announced themselves.

Removing a synchronisation step sounds minor and is not. It is the difference between a task that requires an editing session and a task that returns a finished asset. For a business owner who is not going to open a video editor at nine on a Tuesday night, that difference determines whether the thing gets made at all. This is the same pattern we described in multimodal AI for small business workflows: the value of handling several media types at once is not the novelty, it is the elimination of the handoffs between them.

Why open weights change the calculation

On 3 August, MiniMax released the model weights publicly, under a planned community licence that permits commercial use for organisations under $20 million in revenue with attribution. Almost every small business reading this sits comfortably inside that threshold.

This matters for a reason that has nothing to do with running the model yourself, which most businesses should not attempt. It matters because an open-weight model cannot be taken away from you. We have written about vendor availability risk before, and the pattern is consistent: a business builds a process around a hosted AI tool, the vendor changes the pricing or the terms or the regional availability, and the process breaks with no warning and no recourse.

Published weights change that. Even if you never download them, their existence means multiple providers can host the same model, which creates competition on price and a genuine fallback if one host disappears. The practical effect for a small business is that a workflow built on H3 has more than one door out. That is worth real money over a two-year horizon, and it is invisible on any feature comparison chart.

What it actually costs to use

MiniMax lists 2K generation at $0.13 per output second. A 768p tier is listed at $0.09 per second but is currently in closed beta and requires contacting sales, so treat the 2K figure as the real number for now.

Run that arithmetic against something recognisable. A fifteen-second clip at 2K costs about $1.95. Ten variations of that clip, which is what you actually need because the first attempt is never the one you use, costs about $19.50. A month of weekly social video at four finished pieces, budgeting five attempts each, comes to roughly $39. Reporting around the launch put H3 at approximately 30% of the cost of Seedance 2.0, which is the comparison most people in the space were making.

The honest framing is that the generation cost has become the smallest line in the budget. What it costs you now is attention: deciding what the video should say, judging which of the ten attempts is actually good, and having a point of view worth fifteen seconds of someone else's time. Those were always the hard parts. They were just previously hidden behind a production invoice that made them look like the easy part.

Want AI content generation wired into your actual marketing workflow rather than sitting in a browser tab? Book a €49 audit and we will map where it fits.

Where this fits in a small business

The temptation with a new video model is to try to replace your brand film. Do not start there. The clip that carries your brand identity is the one place where the uncanny quality of generated video is most likely to be noticed and most expensive to get wrong.

Start instead with the volume tier, the video you currently do not make at all because it could never justify a shoot. Product variant clips are the clearest case: if you sell the same item in nine colours, you have nine pieces of video you were never going to commission. Seasonal and promotional cutdowns are another, where the message changes every three weeks and the production cost never made sense. Explainers for a specific recurring customer question sit in the same bracket, as does social content in the formats that expire in twenty-four hours.

The thread connecting all of those is that they are high-volume, short-lived, and low-risk. Nobody is going to build their impression of your company from a fifteen-second clip showing that the mug also comes in green. That is exactly why it is the right place to learn what the tool does well and where it falls apart, before you point it at anything that carries weight.

The honest limits

Fifteen seconds is a real ceiling, not a soft one. This is a model for clips, not for anything with a narrative arc. If you are imagining a two-minute explainer, you are imagining stitching together eight generations and fighting continuity between every one of them, which is a genuinely painful job and frequently worse than shooting it.

Consistency across generations remains the hardest unsolved problem in this category. The same product prompted twice will come back subtly different: the handle sits at a different angle, the material reads slightly differently, the light has moved. For a single standalone clip that is invisible. Across a set of nine colour variants meant to sit in a grid on a product page, it is the thing your customers notice without being able to say why.

There is also a disclosure question that is about to stop being optional in Europe. Under Article 50 of the EU AI Act, AI-generated or manipulated content must be clearly labelled and carry machine-readable marks, with the deadline for synthetic audio, image, video, and text content falling on 2 December 2026. If you are producing marketing video with a generative model and selling into the EU, labelling is a compliance requirement with a real date attached, not a courtesy. We covered the wider timeline in the EU AI Act deadline guide.

The last limit is the one nobody writes about. Cheap video means everyone gets cheap video, and a feed full of competently generated fifteen-second clips is a feed where competently generated fifteen-second clips stop working. The advantage here is not going to come from having access to the tool, because in six months everyone will. It comes from having something specific to say, which no model has ever been able to supply.


Sources

Quick answers

Common questions.

Want this in your business?

The €49 audit shows you exactly which automations would pay back fastest in your specific operation.

€49 entryFull AI audit + strategy call included

Reserve your auditNo commitment. No contracts. Just clarity.