On 31 July 2026, while ByteDance was shipping its own closed Seedance 2.5 the same day, MiniMax released H3: a video model that treats generation, editing and audio as one problem instead of three separate products bolted together. Within hours, Artificial Analysis's independent leaderboard placed it above every other video model in the world for instruction-based editing - a category most competitors do not seriously compete in at all.
This review works through what H3 actually does, the real (not self-reported) benchmark numbers behind the "beat Google and ByteDance" headlines doing the rounds on AI YouTube, official pricing, the open-weights timeline MiniMax has promised, and the copyright litigation shadowing the platform H3 is built on.
A hands-on walkthrough of H3's editing mode, 12-reference input limit and native audio.
Executive Summary
MiniMax H3 is the third major release in MiniMax's Hailuo video lineage, and the first to be marketed as a genuinely omni-modal system rather than a text/image-to-video generator with audio layered on top. The model accepts text prompts up to roughly 7,000 characters, up to nine reference images, up to three reference videos and up to three reference audio clips in a single request - twelve inputs combined - and produces clips of 4 to 15 seconds at up to 2K resolution with native stereo sound baked into the same generation pass rather than added afterwards.
The headline result is independent, not self-reported: Artificial Analysis's Video Editing Leaderboard, built from blind human preference votes, ranks H3 first with an Elo of 1,130 from over 5,000 samples, ahead of Google's Gemini Omni Flash and Alibaba's HappyHorse-1.0. It also places top three on the separate Text-to-Video and Image-to-Video boards. That combination - competitive generation plus a genuine lead in editing - is what is driving the "beat the biggest AI video models" framing across AI YouTube this week.
- Best for: teams that need to edit or extend existing footage with natural-language instructions, not just generate new clips from a blank prompt - product video iteration, ad variant generation, motion transfer between reference clips.
- Main caveat: open weights are promised but not yet shipped, and the model sits on the same Hailuo platform currently being sued by three major studios over character reproduction.
What Is MiniMax H3?
MiniMax is a Shanghai-based AI lab probably best known outside China for its MiniMax M-series text and reasoning models and its Hailuo consumer video app. H3 - also referred to as Hailuo H3 in some regional listings - is positioned by the company as a "general-purpose omni-modal generation model" rather than a narrow video generator, built to "jointly understand multimodal contexts spanning text, images, video, and audio," in MiniMax's own words from its 31 July launch post.
That framing matters for what the model is actually for. Where most 2026-era video generators are optimised for a single workflow - prompt in, clip out - MiniMax is explicitly targeting commercial production pipelines: advertising and ad-variant generation, e-commerce and product listing videos, animated posters, film title sequences, website hero loops, character-consistent game cinematics, and video-to-video motion transfer where a performance from one clip is retargeted onto a different subject.
H3 went live simultaneously in MiniMax's platform API, under the model ID MiniMax-H3, and in the consumer Hailuo app, meaning both developers and everyday users had access from day one rather than a staggered rollout.
Architecture: Four Named Components
MiniMax has published more architectural detail for H3 than is typical for a video model launch, naming four distinct components in its technical write-up rather than describing a single black-box diffusion pipeline.
H3-Contextual Omni Representation
The core idea is using language as a generalisable bridge across modalities - text, image, video and audio inputs are mapped into a shared representation space anchored by language, rather than each modality having its own isolated encoder that gets fused only at the output stage. This is what allows a single request to combine "one image for identity, another for a product, a video for motion, and an audio clip for voice," per third-party technical breakdowns of the launch.
H3-VAE
An overhauled tokenizer that MiniMax says delivers a 4x gain in effective sequence length, which is what makes native 2K output computationally viable rather than an upscaled afterthought - the same problem DeepSeek and other labs solve with sparse attention for text, applied here to video tokens.
H3-Omni Transformer
Separates understanding workloads (parsing the prompt and reference inputs) from generation workloads (producing the output tokens), which MiniMax credits with a roughly 30% training throughput improvement over an entangled architecture.
H3-In-Context Regeneration
The base model first generates a lower-resolution draft, then regenerates it in-context for high-quality upscaling to the final 2K output - rather than relying on a bolt-on super-resolution model that has no awareness of what the base model was actually trying to produce.

Capabilities Deep Dive
Instruction-Based Video Editing
This is H3's standout feature and the reason it tops the editing leaderboard specifically: you can feed it an existing video plus a text instruction, and it will modify the footage rather than regenerating from nothing. That includes swapping elements within a scene, retiming action, and transferring motion from a reference clip onto a different subject or character - a workflow closer to an AI-assisted edit than a fresh render.
Native Stereo Audio
Audio is generated in the same pass as the video, without separating voice, sound effects and music into distinct post-production layers the way most competing pipelines do. MiniMax describes this as producing output with "no separation between voice, sound effects, music" baked in from generation.
Twelve-File Multimodal Input
| Input Type | Limit |
|---|---|
| Reference images | Up to 9 (first 5 free, then $0.04 each) |
| Reference videos | Up to 3, each 2-15s (15s combined max) |
| Reference audio | Up to 3 clips, free, 2-15s (15s combined max) |
| Text prompt | Up to ~7,000 characters |
| Combined reference files | Up to 12 |
That 9 images + 3 videos + 3 audio ceiling is not a MiniMax-published spec sheet figure but is consistent across multiple third-party technical trackers and matches the "12 reference files" figure creators have been quoting since launch. It mirrors, almost exactly, the reference-input ceiling ByteDance built into Seedance 2.0 earlier in 2026 - suggesting twelve combined references may be settling into an informal industry ceiling for this generation of models rather than a MiniMax-specific innovation.
Output Specifications
- Resolution: 2K generally available; a cheaper 768p tier exists but is currently closed beta, requiring direct contact with MiniMax sales
- Duration: 4 to 15 seconds, in integer-second increments
- Commercial licence: positioned explicitly for advertising, e-commerce and brand use cases from launch
Benchmarks: The Artificial Analysis Numbers
MiniMax's own launch page is unusually light on head-to-head benchmark tables - it talks about relative pricing rather than quality scores. The independent numbers come from Artificial Analysis, which runs blind human-preference voting across video models and converts the results into Elo ratings, the same methodology used by LMArena for text models.
| Rank | Model | Elo (Video Editing) |
|---|---|---|
| #1 | MiniMax H3 | 1,130 |
| #2 | Google Gemini Omni Flash | 1,122 |
| #3 | Alibaba HappyHorse-1.0 | 1,096 |
| #5 | ByteDance Dreamina Seedance 2.0 (720p) | 1,035 |
| #6 | Runway Aleph 2.0 | 1,011 |
| #7 | KlingAI Kling 3.0 Omni (1080p) | 1,000 |
The 1,130 Elo score is drawn from over 5,043 blind-preference samples as of 31 July 2026 - a meaningful sample size for this kind of leaderboard, though it is a snapshot taken on launch day and will move as more votes accumulate. Artificial Analysis separately places H3 top three on its Text-to-Video board (#2, behind one other model) and top three on Image-to-Video (#3), meaning the editing win is not offset by weak generation-from-scratch quality elsewhere.
The distinguishing factor Artificial Analysis and independent commentators both point to is architectural, not just a quality edge: H3 is one of very few models on the editing board that accepts a video as input at all, rather than only text or an image. Most of the models it out-scored are being judged on a task type - editing an existing clip - that they were not primarily built for, which is worth keeping in mind when reading the ranking as "H3 is the best video model" rather than the more precise "H3 is the best video editor."
Pricing and Access
| Item | Price |
|---|---|
| 2K generation (public, pay-as-you-go) | $0.13 per output second (~$1.95 for 15s) |
| 768p generation (closed beta) | $0.09 per output second, sales contact required |
| Reference audio | Free |
| Reference images (1st-5th) | Free |
| Reference images (6th-9th) | $0.04 each |
On MiniMax's own framing, the pitch is cost per unit of quality rather than an outright cheapest-in-market claim: the company says its 2K per-second price is "less than a third of mainstream models," and its (currently closed) 768p tier is "less than half the price of mainstream models' 720p" rate. Those are MiniMax's own comparative claims rather than an independently audited price comparison, so treat the multiplier as directional.
Access is live now through MiniMax's platform API under the model ID MiniMax-H3, and through the consumer Hailuo AI app, with no waitlist reported at launch for the 2K tier.
The Open-Weights Promise
MiniMax has stated it intends to release H3's weights, but has been deliberately non-specific about timing: its own blog post says only that this will happen "in the coming days, subject to applicable laws and regulations" - the regulatory caveat referring to China's AI export and content-review rules that any open release of this kind has to clear first.
Several Chinese-language outlets have reported a more specific target: weights going live on 3 August 2026 at 00:00 Beijing time. That date has not been confirmed directly by MiniMax and should be treated as a reported expectation rather than an announced commitment. As of the 31 July/1 August launch window, MiniMax's official Hugging Face organisation carried no H3 checkpoint and its GitHub contained no H3 repository - only its existing M-series text models (M1, M2, M3).
If the weights do land as reported, H3 would become by a wide margin the most capable openly-licensed video model available, since none of the models ranked above or near it on the Artificial Analysis boards (Gemini Omni Flash, HappyHorse-1.0, Seedance 2.0, Runway Aleph, Kling 3.0) are open-weight releases.
The Disney, Universal & WBD Lawsuit
H3 does not launch into a clean legal environment. Disney, Universal and Warner Bros. Discovery jointly filed a copyright infringement suit against MiniMax on 16 September 2025, alleging that its Hailuo video and image generation service was trained on unauthorised copies of their copyrighted characters and that the tool reproduces recognisable figures - Spider-Man, Darth Vader, Shrek and Wonder Woman among the examples cited - from ordinary user prompts. The studios also pointed to MiniMax's own promotional material, which reportedly featured generated clips of studio characters, and to sponsored tutorial content walking users through prompts for scenes like "Spider-Man and Supergirl kissing in the park."
On 26 May 2026, US District Judge Stanley Blumenfeld denied MiniMax's motion to dismiss, finding the studios' allegations plausible enough to proceed into full discovery. As of July 2026 the case remains active and unsettled - meaning H3 shipped on 31 July while its parent platform was mid-discovery in exactly the kind of character-reproduction lawsuit that forced ByteDance into emergency safety measures after Seedance 2.0's launch earlier in the year.
H3 is built on the same underlying Hailuo platform named in the suit, though the complaint predates H3 specifically and targets the broader product line rather than this release in isolation. MiniMax has not issued a public H3-specific statement on content moderation for copyrighted characters at the time of writing; anyone evaluating H3 for commercial or brand-facing work should treat that as an open question rather than an assumed safeguard.
Where H3 sits in the same week's wider AI video news cycle, including ByteDance's competing Seedance 2.5.
How H3 Compares
| Feature | MiniMax H3 | Seedance 2.0 | Sora 2 |
|---|---|---|---|
| Developer | MiniMax | ByteDance | OpenAI |
| Max resolution | 2K | 2K | 1080p |
| Max duration | 15s | 15s | 12s |
| Video-input editing | Yes - #1 on Artificial Analysis | No (V2V motion transfer only) | No |
| Native audio | Yes | Yes | Yes |
| Max combined references | 12 (9 img + 3 vid + 3 audio) | 12 (9 img + 3 vid + 3 audio) | 1 image |
| Open weights | Promised, not yet shipped | No | No |
| 2K price | $0.13/second | ~£2.05-£5.05 per 15s clip | Not offered at 2K |
| Active copyright litigation | Yes (Disney/Universal/WBD, filed Sep 2025) | Cease-and-desist letters (Feb 2026) | No known litigation |
Read together with our Seedance 2.0 review, the pattern across 2026's Chinese video model releases is consistent: technically the most feature-dense options on the market, priced well below Western alternatives, shipping alongside real and unresolved IP disputes over how their training data and character generation were handled.
Limitations and Open Questions
- Open weights not yet real: every capability discussed here reflects the hosted API and app; a self-hosted deployment does not exist yet, and the reported 3 August date is unconfirmed by MiniMax.
- 768p tier is gated: the cheaper resolution option requires a sales conversation rather than self-serve access, limiting how meaningfully "less than half the price" claims can be tested by an individual developer today.
- Editing-leaderboard framing can overstate the win: H3's #1 spot is specifically for video editing, a category few top-tier competitors seriously contest; its text-to-video and image-to-video rankings (#2 and #3) are strong but not first place.
- No independent replication of the architecture claims: the 4x sequence-length gain and 30% training throughput figures are MiniMax's own, with no third-party technical paper available at launch to verify them against.
- Unresolved litigation risk: brands considering H3 for commercial content should factor in that its parent platform is in active discovery over character-reproduction claims from three major studios.
Who Should Use It
Reach for H3 if your workflow genuinely involves editing existing footage rather than generating from a blank prompt - product video iteration, ad variant testing, or motion transfer between a reference performance and a different subject - since that is the specific task Artificial Analysis measured it winning at, and the $0.13/second 2K pricing is competitive for that use case.
Wait, or stick with an established option, if you need guaranteed IP-safe output for brand work with zero litigation exposure risk, if you need the 768p self-serve tier that is not yet generally available, or if you specifically need open weights today rather than a promised release - in which case Seedance 2.0 or Sora 2 remain the more established, if pricier, choices.
The Bottom Line
MiniMax H3's #1 ranking on Artificial Analysis's video-editing leaderboard is real, independently measured, and worth taking seriously - it is not another self-reported benchmark table from the model's own maker. The combination of instruction-based video editing, native audio, twelve-file multimodal input and a promised open-weights release genuinely differentiates it from both Sora 2 and Seedance 2.0, and its $0.13-per-second pricing undercuts most Western alternatives for equivalent output.
But the open-weights promise is still a promise, not a shipped artefact, and the model launches directly into the shadow of a live copyright suit against its parent platform - one where a federal judge has already found the studios' claims credible enough to proceed to discovery. For technical evaluation and non-brand-sensitive production work, H3 looks like the most capable editing-focused video model available today. For anything requiring airtight IP indemnification, the safer path remains a Western provider with a cleaner litigation record, at least until MiniMax's legal position and open-source release both become clearer.
Last updated: 2 August 2026. Based on MiniMax's official H3 launch blog, Artificial Analysis's public leaderboards, and contemporaneous reporting on the Disney/Universal/WBD litigation. Open-weights timing and 768p general availability are subject to change; check MiniMax's official channels for the latest status.





