AI Tools Review

Review

MiniMax H3: Open-Weights Video Model Explained

AI Tools Review Editorial Team2 August 2026
MiniMax H3: Open-Weights Video Model Explained
  • MiniMax
  • MiniMax H3
  • Hailuo
  • AI Video

On 31 July 2026, while ByteDance was shipping its own closed Seedance 2.5 the same day, MiniMax released H3: a video model that treats generation, editing and audio as one problem instead of three separate products bolted together. Within hours, Artificial Analysis's independent leaderboard placed it above every other video model in the world for instruction-based editing - a category most competitors do not seriously compete in at all.

This review works through what H3 actually does, the real (not self-reported) benchmark numbers behind the "beat Google and ByteDance" headlines doing the rounds on AI YouTube, official pricing, the open-weights timeline MiniMax has promised, and the copyright litigation shadowing the platform H3 is built on.

A hands-on walkthrough of H3's editing mode, 12-reference input limit and native audio.

Executive Summary

MiniMax H3 is the third major release in MiniMax's Hailuo video lineage, and the first to be marketed as a genuinely omni-modal system rather than a text/image-to-video generator with audio layered on top. The model accepts text prompts up to roughly 7,000 characters, up to nine reference images, up to three reference videos and up to three reference audio clips in a single request - twelve inputs combined - and produces clips of 4 to 15 seconds at up to 2K resolution with native stereo sound baked into the same generation pass rather than added afterwards.

The headline result is independent, not self-reported: Artificial Analysis's Video Editing Leaderboard, built from blind human preference votes, ranks H3 first with an Elo of 1,130 from over 5,000 samples, ahead of Google's Gemini Omni Flash and Alibaba's HappyHorse-1.0. It also places top three on the separate Text-to-Video and Image-to-Video boards. That combination - competitive generation plus a genuine lead in editing - is what is driving the "beat the biggest AI video models" framing across AI YouTube this week.

  • Best for: teams that need to edit or extend existing footage with natural-language instructions, not just generate new clips from a blank prompt - product video iteration, ad variant generation, motion transfer between reference clips.
  • Main caveat: open weights are promised but not yet shipped, and the model sits on the same Hailuo platform currently being sued by three major studios over character reproduction.

What Is MiniMax H3?

MiniMax is a Shanghai-based AI lab probably best known outside China for its MiniMax M-series text and reasoning models and its Hailuo consumer video app. H3 - also referred to as Hailuo H3 in some regional listings - is positioned by the company as a "general-purpose omni-modal generation model" rather than a narrow video generator, built to "jointly understand multimodal contexts spanning text, images, video, and audio," in MiniMax's own words from its 31 July launch post.

That framing matters for what the model is actually for. Where most 2026-era video generators are optimised for a single workflow - prompt in, clip out - MiniMax is explicitly targeting commercial production pipelines: advertising and ad-variant generation, e-commerce and product listing videos, animated posters, film title sequences, website hero loops, character-consistent game cinematics, and video-to-video motion transfer where a performance from one clip is retargeted onto a different subject.

H3 went live simultaneously in MiniMax's platform API, under the model ID MiniMax-H3, and in the consumer Hailuo app, meaning both developers and everyday users had access from day one rather than a staggered rollout.

Architecture: Four Named Components

MiniMax has published more architectural detail for H3 than is typical for a video model launch, naming four distinct components in its technical write-up rather than describing a single black-box diffusion pipeline.

H3-Contextual Omni Representation

The core idea is using language as a generalisable bridge across modalities - text, image, video and audio inputs are mapped into a shared representation space anchored by language, rather than each modality having its own isolated encoder that gets fused only at the output stage. This is what allows a single request to combine "one image for identity, another for a product, a video for motion, and an audio clip for voice," per third-party technical breakdowns of the launch.

H3-VAE

An overhauled tokenizer that MiniMax says delivers a 4x gain in effective sequence length, which is what makes native 2K output computationally viable rather than an upscaled afterthought - the same problem DeepSeek and other labs solve with sparse attention for text, applied here to video tokens.

H3-Omni Transformer

Separates understanding workloads (parsing the prompt and reference inputs) from generation workloads (producing the output tokens), which MiniMax credits with a roughly 30% training throughput improvement over an entangled architecture.

H3-In-Context Regeneration

The base model first generates a lower-resolution draft, then regenerates it in-context for high-quality upscaling to the final 2K output - rather than relying on a bolt-on super-resolution model that has no awareness of what the base model was actually trying to produce.

A photorealistic frame generated by MiniMax H3 showing a young woman with curly hair drinking from a small white cup at an outdoor cafe table, with blurred pedestrians and a 'CAFÉ' sign in the background.
A sample frame from MiniMax's official H3 launch materials, showing the model's photorealistic output at street-level detail. Source: MiniMax.

Capabilities Deep Dive

Instruction-Based Video Editing

This is H3's standout feature and the reason it tops the editing leaderboard specifically: you can feed it an existing video plus a text instruction, and it will modify the footage rather than regenerating from nothing. That includes swapping elements within a scene, retiming action, and transferring motion from a reference clip onto a different subject or character - a workflow closer to an AI-assisted edit than a fresh render.

Native Stereo Audio

Audio is generated in the same pass as the video, without separating voice, sound effects and music into distinct post-production layers the way most competing pipelines do. MiniMax describes this as producing output with "no separation between voice, sound effects, music" baked in from generation.

Twelve-File Multimodal Input

Input TypeLimit
Reference imagesUp to 9 (first 5 free, then $0.04 each)
Reference videosUp to 3, each 2-15s (15s combined max)
Reference audioUp to 3 clips, free, 2-15s (15s combined max)
Text promptUp to ~7,000 characters
Combined reference filesUp to 12

That 9 images + 3 videos + 3 audio ceiling is not a MiniMax-published spec sheet figure but is consistent across multiple third-party technical trackers and matches the "12 reference files" figure creators have been quoting since launch. It mirrors, almost exactly, the reference-input ceiling ByteDance built into Seedance 2.0 earlier in 2026 - suggesting twelve combined references may be settling into an informal industry ceiling for this generation of models rather than a MiniMax-specific innovation.

Output Specifications

  • Resolution: 2K generally available; a cheaper 768p tier exists but is currently closed beta, requiring direct contact with MiniMax sales
  • Duration: 4 to 15 seconds, in integer-second increments
  • Commercial licence: positioned explicitly for advertising, e-commerce and brand use cases from launch

Benchmarks: The Artificial Analysis Numbers

MiniMax's own launch page is unusually light on head-to-head benchmark tables - it talks about relative pricing rather than quality scores. The independent numbers come from Artificial Analysis, which runs blind human-preference voting across video models and converts the results into Elo ratings, the same methodology used by LMArena for text models.

RankModelElo (Video Editing)
#1MiniMax H31,130
#2Google Gemini Omni Flash1,122
#3Alibaba HappyHorse-1.01,096
#5ByteDance Dreamina Seedance 2.0 (720p)1,035
#6Runway Aleph 2.01,011
#7KlingAI Kling 3.0 Omni (1080p)1,000

The 1,130 Elo score is drawn from over 5,043 blind-preference samples as of 31 July 2026 - a meaningful sample size for this kind of leaderboard, though it is a snapshot taken on launch day and will move as more votes accumulate. Artificial Analysis separately places H3 top three on its Text-to-Video board (#2, behind one other model) and top three on Image-to-Video (#3), meaning the editing win is not offset by weak generation-from-scratch quality elsewhere.

The distinguishing factor Artificial Analysis and independent commentators both point to is architectural, not just a quality edge: H3 is one of very few models on the editing board that accepts a video as input at all, rather than only text or an image. Most of the models it out-scored are being judged on a task type - editing an existing clip - that they were not primarily built for, which is worth keeping in mind when reading the ranking as "H3 is the best video model" rather than the more precise "H3 is the best video editor."

Pricing and Access

ItemPrice
2K generation (public, pay-as-you-go)$0.13 per output second (~$1.95 for 15s)
768p generation (closed beta)$0.09 per output second, sales contact required
Reference audioFree
Reference images (1st-5th)Free
Reference images (6th-9th)$0.04 each

On MiniMax's own framing, the pitch is cost per unit of quality rather than an outright cheapest-in-market claim: the company says its 2K per-second price is "less than a third of mainstream models," and its (currently closed) 768p tier is "less than half the price of mainstream models' 720p" rate. Those are MiniMax's own comparative claims rather than an independently audited price comparison, so treat the multiplier as directional.

Access is live now through MiniMax's platform API under the model ID MiniMax-H3, and through the consumer Hailuo AI app, with no waitlist reported at launch for the 2K tier.

The Open-Weights Promise

MiniMax has stated it intends to release H3's weights, but has been deliberately non-specific about timing: its own blog post says only that this will happen "in the coming days, subject to applicable laws and regulations" - the regulatory caveat referring to China's AI export and content-review rules that any open release of this kind has to clear first.

Several Chinese-language outlets have reported a more specific target: weights going live on 3 August 2026 at 00:00 Beijing time. That date has not been confirmed directly by MiniMax and should be treated as a reported expectation rather than an announced commitment. As of the 31 July/1 August launch window, MiniMax's official Hugging Face organisation carried no H3 checkpoint and its GitHub contained no H3 repository - only its existing M-series text models (M1, M2, M3).

If the weights do land as reported, H3 would become by a wide margin the most capable openly-licensed video model available, since none of the models ranked above or near it on the Artificial Analysis boards (Gemini Omni Flash, HappyHorse-1.0, Seedance 2.0, Runway Aleph, Kling 3.0) are open-weight releases.

The Disney, Universal & WBD Lawsuit

H3 does not launch into a clean legal environment. Disney, Universal and Warner Bros. Discovery jointly filed a copyright infringement suit against MiniMax on 16 September 2025, alleging that its Hailuo video and image generation service was trained on unauthorised copies of their copyrighted characters and that the tool reproduces recognisable figures - Spider-Man, Darth Vader, Shrek and Wonder Woman among the examples cited - from ordinary user prompts. The studios also pointed to MiniMax's own promotional material, which reportedly featured generated clips of studio characters, and to sponsored tutorial content walking users through prompts for scenes like "Spider-Man and Supergirl kissing in the park."

On 26 May 2026, US District Judge Stanley Blumenfeld denied MiniMax's motion to dismiss, finding the studios' allegations plausible enough to proceed into full discovery. As of July 2026 the case remains active and unsettled - meaning H3 shipped on 31 July while its parent platform was mid-discovery in exactly the kind of character-reproduction lawsuit that forced ByteDance into emergency safety measures after Seedance 2.0's launch earlier in the year.

H3 is built on the same underlying Hailuo platform named in the suit, though the complaint predates H3 specifically and targets the broader product line rather than this release in isolation. MiniMax has not issued a public H3-specific statement on content moderation for copyrighted characters at the time of writing; anyone evaluating H3 for commercial or brand-facing work should treat that as an open question rather than an assumed safeguard.

Where H3 sits in the same week's wider AI video news cycle, including ByteDance's competing Seedance 2.5.

How H3 Compares

FeatureMiniMax H3Seedance 2.0Sora 2
DeveloperMiniMaxByteDanceOpenAI
Max resolution2K2K1080p
Max duration15s15s12s
Video-input editingYes - #1 on Artificial AnalysisNo (V2V motion transfer only)No
Native audioYesYesYes
Max combined references12 (9 img + 3 vid + 3 audio)12 (9 img + 3 vid + 3 audio)1 image
Open weightsPromised, not yet shippedNoNo
2K price$0.13/second~£2.05-£5.05 per 15s clipNot offered at 2K
Active copyright litigationYes (Disney/Universal/WBD, filed Sep 2025)Cease-and-desist letters (Feb 2026)No known litigation

Read together with our Seedance 2.0 review, the pattern across 2026's Chinese video model releases is consistent: technically the most feature-dense options on the market, priced well below Western alternatives, shipping alongside real and unresolved IP disputes over how their training data and character generation were handled.

Limitations and Open Questions

  • Open weights not yet real: every capability discussed here reflects the hosted API and app; a self-hosted deployment does not exist yet, and the reported 3 August date is unconfirmed by MiniMax.
  • 768p tier is gated: the cheaper resolution option requires a sales conversation rather than self-serve access, limiting how meaningfully "less than half the price" claims can be tested by an individual developer today.
  • Editing-leaderboard framing can overstate the win: H3's #1 spot is specifically for video editing, a category few top-tier competitors seriously contest; its text-to-video and image-to-video rankings (#2 and #3) are strong but not first place.
  • No independent replication of the architecture claims: the 4x sequence-length gain and 30% training throughput figures are MiniMax's own, with no third-party technical paper available at launch to verify them against.
  • Unresolved litigation risk: brands considering H3 for commercial content should factor in that its parent platform is in active discovery over character-reproduction claims from three major studios.

Who Should Use It

Reach for H3 if your workflow genuinely involves editing existing footage rather than generating from a blank prompt - product video iteration, ad variant testing, or motion transfer between a reference performance and a different subject - since that is the specific task Artificial Analysis measured it winning at, and the $0.13/second 2K pricing is competitive for that use case.

Wait, or stick with an established option, if you need guaranteed IP-safe output for brand work with zero litigation exposure risk, if you need the 768p self-serve tier that is not yet generally available, or if you specifically need open weights today rather than a promised release - in which case Seedance 2.0 or Sora 2 remain the more established, if pricier, choices.

The Bottom Line

MiniMax H3's #1 ranking on Artificial Analysis's video-editing leaderboard is real, independently measured, and worth taking seriously - it is not another self-reported benchmark table from the model's own maker. The combination of instruction-based video editing, native audio, twelve-file multimodal input and a promised open-weights release genuinely differentiates it from both Sora 2 and Seedance 2.0, and its $0.13-per-second pricing undercuts most Western alternatives for equivalent output.

But the open-weights promise is still a promise, not a shipped artefact, and the model launches directly into the shadow of a live copyright suit against its parent platform - one where a federal judge has already found the studios' claims credible enough to proceed to discovery. For technical evaluation and non-brand-sensitive production work, H3 looks like the most capable editing-focused video model available today. For anything requiring airtight IP indemnification, the safer path remains a Western provider with a cleaner litigation record, at least until MiniMax's legal position and open-source release both become clearer.

Last updated: 2 August 2026. Based on MiniMax's official H3 launch blog, Artificial Analysis's public leaderboards, and contemporaneous reporting on the Disney/Universal/WBD litigation. Open-weights timing and 768p general availability are subject to change; check MiniMax's official channels for the latest status.

Frequently Asked Questions

What is MiniMax H3?
MiniMax H3 is a general-purpose omni-modal generation model launched by the Shanghai AI firm MiniMax on 31 July 2026. It jointly understands text, images, video and audio in a single context, and generates video clips up to 15 seconds long at up to 2K resolution with native stereo audio, as well as editing existing footage from an instruction and a source clip.
Is MiniMax H3 better than Sora 2 or Veo 3?
On Artificial Analysis's independent leaderboards, H3 ranks #1 for video editing (Elo 1,130, ahead of Google's Gemini Omni Flash and Alibaba's HappyHorse-1.0) and top three for text-to-video and image-to-video. It is the only model in the top tier that accepts a source video as input for instruction-based editing, which neither Sora 2 nor Veo 3 currently offers as a first-class feature. Raw single-shot generation quality is closely contested, and MiniMax has not published a head-to-head against either model.
How much does MiniMax H3 cost?
MiniMax lists 2K generation at $0.13 per output second through its platform API, working out to roughly $1.95 for a full 15-second clip. A 768p tier at $0.09 per second exists but is in closed beta and requires contacting MiniMax sales. Reference audio is free, and the first five reference images in a generation are free, with each additional image (up to nine total) costing $0.04.
Are MiniMax H3's weights open source?
Not yet, as of the model's 31 July 2026 launch. MiniMax says it plans to release the weights 'in the coming days, subject to applicable laws and regulations,' and several Chinese outlets have reported a specific target of 3 August 2026 at 00:00 Beijing time, though MiniMax itself has not confirmed an exact date. At the time of writing, no H3 checkpoint has appeared on MiniMax's Hugging Face organisation.
Is MiniMax involved in a copyright lawsuit?
Yes. Disney, Universal and Warner Bros. Discovery jointly sued MiniMax in September 2025 over its Hailuo video and image generator - the same underlying platform H3 extends - alleging it was trained on and reproduces copyrighted characters including Spider-Man, Darth Vader, Shrek and Wonder Woman. A federal judge denied MiniMax's motion to dismiss in May 2026, and the case remains in active discovery as H3 ships.

Key takeaways

A real #1, not a marketing claim

Artificial Analysis - not MiniMax - ranks H3 first on its Video Editing leaderboard with a measured Elo of 1,130 from over 5,000 blind-preference votes.

Editing is the actual differentiator

H3 is one of the few top-tier models that accepts a source video as input for instruction-based edits, not just text-to-video generation from scratch.

Ships under active litigation

MiniMax lost its bid to dismiss a Disney/Universal/WBD copyright suit over Hailuo in May 2026. H3 extends that same platform while the case proceeds.

Explore more AI tool comparisons

In-depth reviews, benchmarks and guides to help you choose the right AI tools.

Browse all reviews
AI Tools Review Editorial Team

AI Tools Review Editorial Team Expert verified

Our editorial team consists of veteran AI researchers, software engineers, and industry analysts. We spend hundreds of hours benchmarking frontier models natively to provide you with objective, actionable intelligence on agentic AI capabilities and cybersecurity landscapes.