MiniMax H3 vs MiniMax H3 Max: Both Take References
Most comparisons say H3 Max drops reference generation. It doesn't. We run both models in production, and the real split is speed against cost, not references.

If you have read anything comparing MiniMax H3 to MiniMax H3 Max in the past few weeks, you have probably seen some version of this: H3 takes reference images, video and audio, H3 Max doesn't, so pick based on whether your shot needs references.
That is the cleanest possible way to frame the choice. It is also wrong.
Both models accept reference generation, with the same limits: up to 9 reference images, 3 reference clips and 3 audio files in a single request. We run both on Veevid, and reference to video works on H3 Max today.
The confusion is easy to explain. H3 Max launched with text to video and image to video front and center, and most model pages list only those two. The reference endpoint exists but was never part of the announcement, so write-ups built on launch materials all inherited the same gap. If you only read the marketing page, you would conclude references were missing too.
So if references are not the dividing line, what is?
The Short Version
| MiniMax H3 | MiniMax H3 Max | |
|---|---|---|
| Output duration | 4 to 15 seconds | 5 to 15 seconds |
| Resolution | 768P, 2K | 480P, 768P, 1080P |
| Reference images | 9 | 9 |
| Reference clips | 3 | 3 |
| Reference audio | 3 | 3 |
| Native audio | Yes, stereo | Yes |
| First and last frame | Yes | Yes |
| Typical wait | Minutes | Seconds |
| 5s at 768P | 40 credits | 80 credits |
Two numbers in that table do most of the work: the wait and the price. Everything else is close enough that it rarely decides anything.
Speed Is What H3 Max Actually Sells
H3 Max is a post-trained build of MiniMax H3. The training added data aimed at prompt adherence and aesthetics, but the part you feel immediately is throughput.
Published benchmarks put a 5 second 768P render at roughly 2.5 seconds of inference, and a 15 second one at about 15 seconds. Those figures cover the denoising step alone. In our own request logs the full upstream call, queueing and encoding included, returns a 5 second 768P clip in under 4 seconds. On top of that sits callback handling and the transfer of the finished file, and H3 Max still lands in seconds where H3 takes minutes.
That gap changes how you work rather than what you get. A clip that renders while you are still looking at the prompt lets you run three takes and compare them. A render you wait minutes for turns every attempt into a decision you have to commit to before you see the result.
One setting affects this more than any other. Prompt expansion rewrites your brief before rendering, and it has four modes. Balanced, the default, adds about a second. Fast returns almost immediately. Quality can spend up to 30 seconds on the rewrite alone, so if you are iterating, turn it down.
Resolution: 2K Against 1080P
H3 tops out at 2K. H3 Max tops out at 1080P, and there is a detail worth knowing: 1080P on H3 Max is a latent refinement of a native 768P render rather than a native 1080P one. Treat it as a cleaner finish, not a jump in detail. If you need real resolution headroom for a deliverable, 2K on H3 is the one that gets you there.
H3 Max does have something H3 lacks at the other end: a 480P tier at 10 credits per second, which is the cheapest way to iterate on H3 Max before committing to a full render. H3 starts at 768P, though at 8 credits per second and a 4 second minimum, a throwaway H3 render is still the cheaper of the two.
What Each One Costs
Output is charged per second of video. H3 is 8 credits per second at 768P and 13 at 2K. H3 Max is 10 at 480P, 16 at 768P and 32 at 1080P.
Reference material is charged differently on each model, which is where the comparison gets less obvious. Here is what actual jobs come to:
| Job | MiniMax H3 | MiniMax H3 Max |
|---|---|---|
| 5s at 768P, prompt only | 40 | 80 |
| 15s at 768P, prompt only | 120 | 240 |
| 5s at 768P with 3 reference images | 40 | 80 |
| 5s at 768P with 9 reference images | 56 | 101 |
| 5s at 768P with a 5s reference clip | 80 | 213 |
| 15s at 768P with a 15s reference clip | 240 | 373 |
| 5s at top tier (2K vs 1080P) | 65 | 160 |
H3 is cheaper in every row. The multiple ranges from about 1.5x to 2.7x depending on the job, and reference clips are where it widens most: H3 folds the clip length into the billed duration, while H3 Max charges reference material out of a separate allowance that a single clip can exhaust.
One detail narrows that gap. Reference clips add nothing to the cost of an H3 Max render at 1080P, which removes the largest line item when you are conditioning on a long clip. It does not make H3 Max the cheaper option, it just stops the clip from compounding the difference. Reference images and audio still count against the shared allowance at every resolution, so this is specifically about clips.
The generator shows the exact total before you spend anything, so you never have to work this out by hand.
Which One to Use
Reach for H3 Max when the loop matters. You are exploring, the brief is not settled, and you want to see three versions of a shot before picking one. Paying double per render is worth it when it turns minutes of waiting into seconds, and the 480P tier makes the exploratory passes cheap.
Reach for MiniMax H3 when the render matters. You know what you want, the shot is going into something, and you would rather have 2K and stereo audio than speed. It is also the straightforward choice for anything reference heavy, since reference clips cost far less there.
If you are conditioning on a long clip, the gap narrows sharply. A 15 second render at H3 Max 1080P costs four times the H3 768P equivalent on a prompt alone, but only twice as much once both are conditioned on a 15 second clip, because the clip is free on H3 Max at that tier. H3 still comes out cheaper, so this is not a reason to switch. It is a reason not to rule out H3 Max on price if you already want it for the speed.
Both models generate audio in the same pass as the picture, both hold a character across a sequence, and both take first and last frames on image to video. Those are table stakes on either one now, which is precisely why the decision comes down to how fast you need it back and what you are willing to pay for that.
Both are live on Veevid: MiniMax H3 Max and MiniMax H3. Reference to video runs on either model.