MiniMax H3 Deep Dive: From “Specialised Toolkit” to “General‑Purpose Intelligence” – An Omni‑Modal Revolution
On July 31, 2026, MiniMax officially released its next‑generation general‑purpose omni‑modal generative model, MiniMax H3. Three days later, on August 3, the model weights were made open‑source. This 33B‑parameter omni‑modal model topped the Artificial Analysis video editing leaderboard with an Elo score of 1130, securing the world’s No. 1 spot, while ranking second in text‑to‑video and third in image‑to‑video generation. It is not only MiniMax’s latest flagship in video generation but also repr
Published August 26, 2026
On July 31, 2026, MiniMax officially released its next‑generation general‑purpose omni‑modal generative model, MiniMax H3. Three days later, on August 3, the model weights were made open‑source. This 33B‑parameter omni‑modal model topped the Artificial Analysis video editing leaderboard with an Elo score of 1130, securing the world’s No. 1 spot, while ranking second in text‑to‑video and third in image‑to‑video generation. It is not only MiniMax’s latest flagship in video generation but also represents a paradigm shift in design philosophy – from a “collection of specialised models” to “general‑purpose task generalisation.”
I. Core Features: Omni‑Modal, Native Audio, 2K Straight‑Out Unified Omni‑Modal Understanding and Generation
H3’s most defining tag is “omni‑modal.” It can uniformly understand multimodal contexts composed of text, images, video, and audio, and directly generate videos with native stereo audio. Users can mix different types of assets – up to 9 images, 3 video clips, and 3 audio clips – within a single context, totalling up to 12 input files. H3 automatically comprehends the relationships among these assets and completes coherent content generation or editing.
Native Stereo Audio with Synchronised Audio‑Visual Generation
Unlike most models on the market that first generate video and then dub it separately, H3’s video and audio are jointly generated. The output is 24 FPS video with 32 kHz stereo audio – dialogue, sound effects, ambient sounds, and background music are produced simultaneously during generation, ensuring better synchronisation.
Three Resolutions and Two Generation Modes
By default, H3 outputs videos with a short side of 768 pixels. Through the H3‑Regenerate‑2K module, it can achieve 2K resolution (2560×1440). The model provides two core modes:
H3‑Base‑FL2VA (first‑last frame mode): supports 0, 1, or 2 input images, corresponding to text‑to‑video, first‑frame‑to‑video, last‑frame‑to‑video, and first‑and‑last‑frame‑to‑video generation.
H3‑Base‑Ref2VA (omni‑modal reference mode): supports mixed input of up to 9 images, 3 videos, and 3 audio clips, enabling multimodal reference generation.
The output duration ranges from 4 to 15 seconds, supporting multiple aspect ratios including 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16.
Three Core Technologies
H3’s system consists of three modules:
H3‑Context‑IR: a managed preprocessing and orchestration system that understands relationships among multimodal inputs and converts them into structured representations understandable by H3‑Base.
H3‑Base: the actual generation module, outputting 768p audio‑video.
H3‑Regenerate‑2K: takes the 768p result along with the original context and regenerates it at 2K resolution – not through traditional upscaling, but through a full regeneration pass.
Architecturally, H3 adopts a dense single‑stream Transformer (DiT) with 33B parameters, completely breaking down modality barriers. Approximately 13B of these parameters reside in AdaLN‑related branches, which can be removed at inference via pre‑computation techniques, leaving roughly 20.11B parameters that need to stay resident.
Pricing and Open‑Source
H3’s API pricing is 0.8 RMB per second for 2K resolution, roughly one‑third of comparable flagship models in the industry. At 768P, the price is about half of mainstream 720P models. The model weights have been open‑sourced on platforms including Hugging Face. On the very first day of open‑sourcing, 16 chip manufacturers and platforms – including Huawei Ascend, Moore Threads, AMD, and Intel – completed adaptation. However, the open‑source license explicitly excludes usage rights in the United States, the European Union, the United Kingdom, and South Korea.
II. Vertical Comparison: From Hailuo 01/02 to H3 – A Paradigm Revolution H3 is not a simple upgrade of Hailuo 02; it represents a complete architectural overhaul.
The “Specialised Model Toolkit” Approach of Hailuo 01 and 02
Earlier, MiniMax built the Hailuo 01 system from scratch, while Hailuo 02 focused on improving architecture efficiency, data quality, and scale. Their philosophy was: make each specialised model better and faster – image generation was split into separate experts like T2I, subject reference, and motion reference; voice generation treated vocals, sound effects, and music as independent domains; and video generation was even further subdivided into text‑to‑video, image‑to‑video, first‑last frame, subject reference, motion reference, and more. These tasks, capabilities, and modalities had clear boundaries.
H3’s “Task Generalisation” Revolution
The H3 design team explicitly recognised that such compartmentalisation limited flexibility and constrained the model’s generalisation ability. Therefore, the first guiding principle for building H3 was unified task generalisation.
The specific changes include:
Complete Tokenizer overhaul: H3‑VAE achieves comprehensive improvements in reconstruction quality and learnability, with a high compression rate delivering a 4× sequence‑length gain.
Unified audio modelling: no longer distinguishing vocals, sound effects, or music – all are jointly modelled.
Generalised reference and editing: natural language is used to express reference and editing relationships, rather than being confined to a limited set of task types.
Training strategy: mixing various data types and tasks as early as possible.
The team even stated openly that H3 deliberately abandoned the Hailuo 02 architecture, because the latter “would introduce unnecessary complexity.”
Specification Comparison
Dimension Hailuo 01 Hailuo 02 Hailuo 2.3 MiniMax H3 Release date Early June 2025 October 2025 July 31, 2026 Max resolution — 1080p — 2K Max duration — 10s — 15s Audio Post‑dub Post‑dub Post‑dub Native stereo Architecture Specialised toolkit Specialised toolkit Specialised toolkit Unified omni‑modal Open‑source No No No Yes From Hailuo 01 to H3, MiniMax has transitioned from a “collection of specialised tools” to a “general‑purpose intelligent system.” H3 is no longer a stack of isolated functions, but an omni‑modal system that can understand complex multimodal instructions and perform unified generation and editing.
III. Horizontal Comparison: A Comprehensive Contest with Industry Rivals Artificial Analysis Leaderboard: Video Editing World No. 1
On the authoritative third‑party evaluation platform Artificial Analysis:
Video editing: World No. 1 (Elo 1130), ahead of Gemini Omni Flash, HappyHorse‑1.0, Wan 2.7, and others.
Text‑to‑video: World No. 2
Image‑to‑video: World No. 3
On the Arena image‑to‑video leaderboard, H3 also ranks first globally.
vs. ByteDance Seedance 2.5
Seedance 2.5 is another leading domestic video generation product. The two have distinctly different positioning:
Seedance 2.5 is more like a “virtual camera”: it pursues precise camera movement, framing, and motion control, supports up to 30‑second generation, and is more stable in complex prompt understanding and long video tasks.
H3 is more like a “post‑production team that understands the final cut”: its strengths lie in visual completion, native audio, and video editing capabilities. The 2K straight‑out clarity meets the needs of general creative production.
Some analysts believe that together they have raised both the lower and upper limits of the AI video赛道: Seedance 2.5 explores high‑value film‑grade applications, while H3, with its cost‑performance and open‑source strategy, replaces a large volume of short‑form drama production needs at the lower end.
vs. Kling 3.0
In side‑by‑side comparisons conducted by independent YouTube creators, H3 produced cleaner details and more convincing character movements, outperforming Kling 3.0. Community feedback also indicates that users are “significantly more satisfied” with H3’s results compared to Kling AI’s free tier.
vs. Runway Gen‑3
Runway Gen‑3 still leads H3 in overall image fidelity and the coherence of complex physical motions. However, H3’s advantages lie in:
Cost: generating a 15‑second 2K video with H3 costs about $1.5, while Runway uses a subscription model (starting at $12/month).
Native audio: Runway requires post‑dubbing; H3 supports it natively.
Open‑source deployability: H3 can be run locally; Runway is a closed‑source service.
vs. Sora 2
Compared to OpenAI’s Sora 2, the biggest differences are:
Audio: Sora requires “generate video first, then dub separately”; H3 generates video and audio simultaneously.
Editing capability: H3 supports contextual video editing; Sora focuses more on narrative text‑to‑video.
Price: H3 is priced at roughly one‑third of closed‑source models like Sora.
Open‑source: H3 weights are open and can be deployed locally.
Cost‑Performance Summary
Model Max Resolution Max Duration Native Audio Price per second (2K) Open‑source MiniMax H3 2K 15s ✅ 0.8 RMB ✅ Seedance 2.5 2K 30s ❌ ~2.4 RMB ❌ Kling 3.0 2K 10s ❌ — ❌ Runway Gen‑3 4K 10s ❌ Subscription ❌ Sora 2 1080p 20s ❌ — ❌ H3’s 2K pricing is about one‑third of comparable flagship models. Some reviews note that H3 offers the highest picture quality per dollar at $0.13 per second (2K).
IV. Conclusion The release of MiniMax H3 marks a paradigm shift in video generation models – from “specialised tools” to “general‑purpose intelligence.” It deliberately abandoned the mature Hailuo 02 architecture in favour of a more difficult but more promising path of “task generalisation.” In performance, H3 has joined the top tier with world No. 1 in video editing and top‑three rankings in text‑to‑video and image‑to‑video. In pricing, it challenges closed‑source giants at one‑third of the industry average. In ecosystem, it pushes open‑source adoption, with 16 chip vendors adapting on day one.
As industry observers have noted, H3 may be bringing a “DeepSeek moment” to the video model industry – opening up generation and editing capabilities that were previously controlled by a few closed‑source models, giving enterprises and developers, for the first time, a video model that approaches the top tier while being deployable on their own infrastructure. Of course, H3 still lags behind top‑tier products like Runway Gen‑3 in complex physics simulation and long‑motion coherence, and there is room for improvement in image detail and model scale. But this 33B‑parameter omni‑modal model has already proven that open‑source and high performance, low cost and high quality, can coexist.
