TL;DR MiniMax H3 Max, released August 27, 2026, is marketed as producing a 5-second clip in under 3 seconds — a backend-inference figure. We timed 26 end-to-end runs instead: a median of 7.2 seconds for a 5.18-second clip with synchronized audio at 768p, but 192.6 seconds at 2K. Resolution, not the tier you choose, is what actually costs you time.
Something genuinely changed in AI video at the end of August 2026, and the industry immediately started describing it badly. The phrase doing the rounds is "faster than real time" — the idea that a model can now produce a video quicker than you can sit and watch it. The framing spread within days of launch, and it is doing a lot of work. The demos behind it are real: fal's launch note reports a 5-second video in under 3 seconds, roughly 35x the throughput of the official endpoint.
We wanted a number that a stranger could reproduce. So we ran a timed test — same prompt style, stated settings, eleven runs, a stopwatch on the whole round trip rather than the flattering middle of it. This post is that data, plus what it means for the way you actually produce video.

1.What is MiniMax H3 Max?
MiniMax H3 Max is a speed-tuned variant of MiniMax H3, post-trained by fal and announced on August 27, 2026. It generates 5 to 15 second clips at 480p or 768p with synchronized audio, and it trades resolution ceiling for throughput.
That trade is the whole design. The base model, MiniMax H3, reaches 2K and offers a reference-driven mode that accepts up to nine images, three video clips and three audio files in a single generation. H3 Max gives up the top of the resolution ladder — there is no 2K and no 4K rung, 768p is the ceiling — and gets dramatically faster in exchange. fal claims roughly 35x the throughput of the base hosted endpoint.
The mode coverage is not the dividing line, which is worth stating because several write-ups have it backwards: checked on 2026-08-30, the speed tier serves reference-to-video too, with the same nine images, three clips and three audio files. What it does not have is a custom-adapter path — there is no trainer for it, and its reference route has no LoRA variant. Resolution ceiling and adapters are the difference; input modes are not.
| MiniMax H3 | MiniMax H3 Max | |
|---|---|---|
| Released | July 31, 2026 | August 27, 2026 |
| Resolution ceiling | 2K (1440px short edge) | 768p |
| Clip length | 5–15 seconds | 5–15 seconds |
| Native audio | Yes, 32kHz stereo | Yes, synchronized |
| Modes | Text, image and reference to video | Text, image and reference to video |
| Custom adapters | LoRA trainers available | No trainer; no LoRA route |
| Built by | MiniMax | fal, post-trained from H3's open weights |
| Positioned for | Maximum quality and control | Maximum speed |
Worth noting for anyone reading vendor pages: H3 Max is fal's tuned edition of an open MiniMax model, not a MiniMax release. And the model is listed under several names across platforms — some catalogues file H3 under Hailuo 3.0 or Hailuo 03 branding, since MiniMax's consumer app serves it. MiniMax's own materials say MiniMax H3.
2.How fast is fast, once you measure the whole errand?
In eleven timed runs on a production platform's fast video tier, a 5-second 768p clip with audio took a median of 7.2 seconds end to end — from the moment the request left our machine to the moment a finished file was downloadable. The fastest run was 6.9 seconds and the slowest 8.2.
To be precise about what this is and is not: we did not benchmark MiniMax's hosted endpoint, and these numbers are not a measurement of it. We measured what a customer actually experiences on a commercial platform at the current fast tier — the whole errand, not the model in isolation. That is the number nobody publishes, and it is the one that decides how you work.
It also explains why a sub-3-second claim and a 7-second experience are both true. A vendor's inference figure measures the model producing frames. The errand additionally includes request, queue admission, audio muxing, storage write, signed URL and download. Nobody ships video from a GPU straight into a browser.
Here is the full run table. Every clip was 5 seconds requested, 768p, with audio, generated through a public dispatch path rather than a private benchmarking harness.
| Run | Scene | Aspect | Wall time |
|---|---|---|---|
| 1 | Marble rolling off a table | 16:9 | 8.20 s |
| 2 | Coastal cliff drone shot | 16:9 | 7.17 s |
| 3 | Two-person cafe dialogue | 16:9 | 7.21 s |
| 4 | Anime rainy street | 16:9 | 8.08 s |
| 5 | Barista pouring latte art | 9:16 | 6.98 s |
| 6 | Skatepark bowl run | 16:9 | 6.94 s |
| 7 | Watercolour fox in snow | 16:9 | 7.05 s |
| 8 | Watercolour harbour at sunset | 16:9 | 8.06 s |
| 9 | Street food market, multi-shot | 16:9 | 7.99 s |
| 10 | Marble scene, repeat | 16:9 | 6.90 s |
| 11 | Marble scene, repeat | 16:9 | 8.04 s |
Median 7.2 seconds. Mean 7.5. The three runs of the identical marble prompt landed at 8.20, 6.90 and 8.04 seconds — a 1.3 second spread on identical input, which tells you queue conditions move the number about as much as the prompt does.
3.Which choice actually costs you time: the tier, or the resolution?
The resolution, overwhelmingly. Moving from a fast tier to a quality tier at the same resolution cost about one second in our runs. Moving from 768p to 2K cost about three minutes.
We ran a second timed set to isolate that — same prompt, same 5-second request, same stopwatch method, five runs per row:
| Tier and resolution | Runs | Median wall time | Range | Credits | Audio in same pass |
|---|---|---|---|---|---|
| Fast tier, 768p | 5 | 8.1 s | 7.0–9.6 s | 60 | Yes |
| Quality tier, 768p | 5 | 9.1 s | 8.9–9.5 s | 60 | Yes |
| Quality tier, 2K | 5 | 192.6 s | 111.9–273.0 s | 90 | Yes |
Three things fall out of that table, and the third is the one that should change your workflow.
First, the tier choice is nearly free in time terms at a fixed resolution — one second between them, which is inside the run-to-run noise on the fast tier. Second, the 2K row is not just slower, it is far less predictable: a 161-second spread between the fastest and slowest run, against 2.6 seconds at 768p. If you are quoting a turnaround to someone, the 2K number you can promise is the slow end, not the median. Third, and most usefully, the whole "faster than real time" conversation is a 768p conversation. At 2K nothing in this class is close to playback speed, and the honest planning rule is to draft at 768p and spend 2K only on the shots that have already earned it.
The same disclaimer applies as above: these are measurements of a commercial platform's tiers end to end, not of any vendor's hosted endpoint in isolation.
The output files were 1344x768 H.264 with an AAC stereo track at 32kHz, running 5.184 seconds. So the honest ratio is 1.39x: it took about 39% longer to make the clip than the clip takes to play.
Run 2, generated in 7.17 seconds. Prompt: an aerial drone shot sweeping along a rugged coastal cliffside at golden hour, waves crashing against dark rocks below, seabirds crossing the frame, warm cinematic colour grade. Audio: wind and surf, generated in the same pass as the frames.
4.What does "faster than real time" actually mean?
"Faster than real time" means the model produces a second of video in less than a second of compute — a claim about the inference step, not about your experience of using the tool.
It is a legitimate engineering milestone and a misleading shopping metric. The gap between the two is everything a production system has to do around the model: admission control when a lot of people press the button at once, writing a file somewhere durable, and handing you a URL that will still work tomorrow.
Our advice for reading any speed claim in this category: ask what the clock started and stopped on. A number that starts at "GPU begins denoising" and stops at "last frame emitted" is a real number about a real thing, and it is not the number you will feel.

5.Is seven seconds actually a big deal?
Yes — because the threshold that matters is not playback speed, it is attention span. Roughly seven seconds is short enough to keep you in the loop rather than sending you off to do something else.
This is the real story that the "faster than real time" framing obscures. When a clip took four minutes, video generation was a batch job: write ten prompts, submit them, go make coffee, come back and triage. At seven seconds it becomes a conversation. You watch, you adjust one clause in the prompt, you go again. Ten iterations is a two-minute exercise rather than a lost afternoon.
That changes what the tool is for. Batch-job video pushes you toward getting it right first time, which means over-specified prompts and low willingness to explore. Conversational video pushes you toward trying the strange idea, because trying it costs seven seconds.

Run 4, generated in 8.08 seconds. Prompt: anime style scene of a girl with a clear umbrella walking through a neon-lit rainy street at night, reflections shimmering on wet pavement, gentle rain and a soft synth score. The rain audio and score were produced alongside the frames, not added afterwards.
6.Does the audio come out in the same pass?
Yes. Every clip in our test arrived with a synchronized AAC stereo track at 32kHz, generated with the frames rather than dubbed on afterwards.
This is the part of the 2026 shift that gets underplayed next to the speed numbers. Native audio removes an entire stage from the pipeline. The old shape was generate video, then source or generate sound, then sync the two and hope the footsteps land on the footfalls. The new shape is one request.
For the cafe dialogue clip, the model produced two speaking characters with room tone and background chatter underneath. It is not a finished mix and you would still ride the levels for anything client-facing. But it is a starting point that used to cost a second tool and a sync pass. If you are assembling scenes from existing footage, automatic sound effects solve the adjacent problem for clips that arrived silent.
Run 3, generated in 7.21 seconds. Prompt: two friends at a corner table in a cosy cafe, one says "You will not believe what happened this morning", the other laughs and replies "Tell me everything", with ambient chatter and clinking cups. Turn your sound on — the dialogue, the room tone and the cups are all in the single generated track.

7.Where the speed tier costs you something
H3 Max tops out at 768p, and there is no custom-adapter path on it: no trainer exists for the speed tier, and its reference route has no LoRA variant. If you need a 2K master, or you are running a character you trained yourself, the base model is the right tool. Reference-driven shots are not the dividing line — the speed tier does those too.
In practice that splits work cleanly:
- Exploration, social drafts, ad variants, storyboards. Speed tier. You are making many things and discarding most of them, so iteration rate dominates.
- Hero shots, client deliverables, anything that gets graded. Quality tier. You are making one thing carefully, so the resolution ceiling dominates.
- Anything vertical for short-form. Either, but note that 768p on the short edge of a 9:16 frame is a 768x1344 file, which is comfortably above what most social platforms re-encode to anyway.
A 768p clip that needs to finish at higher resolution is not a dead end, either — video enhancement is a separate step that upscales an existing file, so a fast draft can be promoted rather than regenerated.
8.How to measure this yourself
The method matters more than our specific numbers, because queue conditions differ by hour and platform. Ours was deliberately boring:
- Fix the settings and state them. Every run: 5 seconds, 768p, audio on, one aspect ratio change noted in the table.
- Start the clock before the request leaves. Not when the job is admitted.
- Stop it when a file exists on disk. Not when the API says "processing", and not when a preview appears.
- Poll at a stated interval. We polled once per second, which means every number carries up to one second of granularity error. Say so rather than quietly rounding down.
- Repeat one prompt at least three times. Single runs measure the queue, not the model. Our 1.3-second spread on identical input is the honest error bar.
- Report the median and the range. A single best-case number is marketing, not measurement.
Run that against any platform you are considering and you will have a comparable figure. We would rather you did that than took ours on trust.

9.Where the model class sits on quality
Speed only matters if the output is worth having. On the Artificial Analysis video leaderboards, this generation of MiniMax models sits at the top of the with-audio arenas rather than somewhere down the list — the base H3 led the video editing arena at launch, and the speed-tuned variant ranks first in image-to-video with audio.
That is unusual. The normal shape of a fast tier is a visible quality tax, and the interesting thing about the current crop is how small that tax has become. In our runs the physics scene held up under scrutiny — a glass marble rolling, catching light, casting a moving shadow — and the anime clip kept a coherent character across the full five seconds. Failure modes still exist, mostly in hands and in any text the model tries to render, which is why an honest workflow assumes some regeneration.
Run 1, generated in 8.20 seconds. Prompt: a glass marble rolls off the edge of a rustic wooden table and bounces across a sunlit tiled kitchen floor with realistic physics, the camera tracking low along the floor. This was the prompt we repeated three times to measure variance.
10.Doing this without the setup
Everything above ran through Imagera's video generator at its fast tier, using the same public dispatch path any account uses. No private endpoint, no preferential queue.
Two things are worth saying plainly about cost, because per-second vendor rates are a bad way to plan. First, you pay for the attempts that miss as well as the ones that land, so your real cost per usable clip is a function of how many times you regenerate — which is exactly why iteration speed and cost are the same conversation. Second, on Imagera the whole chain runs on one credit balance that does not expire, and the cost of each step is shown on the button before you commit, so the arithmetic happens before the spend rather than after it.
If you want the finished clip cut into vertical formats afterwards, video to reels handles that step, and pricing lays out what a credit buys.
11.The honest summary
AI video generation crossed an important threshold in August 2026, and the threshold is not the one in the headlines. The industry gets to say "faster than real time" about the inference step, and that is fair. What you will experience on a production platform is something like seven seconds for a five-second clip with sound — about 1.4x playback.
Seven seconds is the number that matters, because seven seconds is short enough to stay in the loop. That is a bigger change to how video gets made than any resolution bump this year.



