# MiniMax H3

> Source: https://aiwiki.ai/wiki/minimax_h3
> Updated: 2026-08-04
> Fact-checked: 2026-08-04
> Categories: AI Models, Chinese AI, Open Source AI, Video Generation
> License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - attribute to "AI Wiki (aiwiki.ai)"
> Cite as: AI Wiki. "MiniMax H3." aiwiki.ai, 4 Aug 2026. https://aiwiki.ai/wiki/minimax_h3
> From AI Wiki (https://aiwiki.ai), the free encyclopedia of artificial intelligence. Reuse freely with attribution.

**MiniMax H3** (also written MiniMax-H3) is a multimodal video generation model developed by [MiniMax](https://aiwiki.ai/wiki/minimax), the Shanghai-based AI company that operates the [Hailuo AI](https://aiwiki.ai/wiki/hailuo) video service. Announced on July 31, 2026 and released as an [open-weight](https://aiwiki.ai/wiki/open_weights) model on Hugging Face on August 3, 2026, H3 generates video with natively synchronized stereo audio at up to 2K resolution and up to 15 seconds, from a context that can mix text, images, video clips and audio clips.[1][2][3] MiniMax positions it as a general-purpose "omni-modal" system rather than a task-specific video model: text-to-video, image-to-video, video editing, reference-based generation and audio-driven generation are all expressed as natural-language descriptions over a shared multimodal context instead of separate expert models.[1][3] At the time of the weights release, [Artificial Analysis](https://aiwiki.ai/wiki/artificial_analysis) ranked H3 first on its Video Arena video editing leaderboard, which the trade publication The Decoder reported as the first time an open model topped an AI video ranking.[11][12]

The release is open weights with significant strings attached. The MiniMax H3 Community License Agreement, dated August 2, 2026, defines its "Applicable Territory" as worldwide excluding the European Union, the United Kingdom, South Korea and the United States, meaning users in those jurisdictions are not licensed to run the weights locally without separately obtained authorization; MiniMax's hosted API remains available globally.[4][5] Two of the three modules of the full H3 system, the H3-Context-IR preprocessing pipeline and the H3-Regenerate-2K resolution stage, were withheld from the initial release, so local deployments top out at 768p output.[3]

## Key facts

| Item | Detail |
| --- | --- |
| Developer | [MiniMax](https://aiwiki.ai/wiki/minimax) (Hong Kong-listed, Shanghai-based; license issued by its entity Nanonoble Pte. Ltd.)[4] |
| Announced | July 31, 2026, via the blog post "MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities"[1] |
| Weights released | August 3, 2026, on Hugging Face (MiniMaxAI/MiniMax-H3) and ModelScope[2][6][8] |
| Generator | H3-Omni-Transformer, a 33 billion parameter dense single-stream transformer (roughly 13B of it in cacheable AdaLN branches)[3] |
| Text encoder | Qwen3-VL-32B (full pretrained weights, Apache 2.0), hidden states taken from its 50th layer[3][4] |
| Tasks | Text-to-video, first/last-frame image-to-video, reference-to-video, video-to-video editing, audio-driven generation, all with audio-video output[3][8] |
| Output | 4 to 15 seconds, 24 fps, 768p native (2K via the hosted Regenerate-2K stage), 32 kHz stereo audio, dialogue in 11 languages[3] |
| License | MiniMax H3 Community License Agreement, dated August 2, 2026; excludes the EU, UK, South Korea and US; revenue and attribution conditions on commercial use[4] |
| API pricing | $0.13 per second at 2K, $0.08 per second at 768p on MiniMax's Open Platform, as of August 5, 2026[13] |

## Announcement and release

MiniMax announced H3 on July 31, 2026, framing it as a break from the task-fragmented structure of the generative video field: instead of separate expert models for editing, subject reference and motion reference, H3 was pretrained to treat reference and editing relationships as language, on the argument that "language is the bridge to generalization".[1] The launch post said MiniMax planned "to open up the model weights in the coming days, subject to applicable laws and regulations", and claimed a per-second price at 2K of less than a third of mainstream models; both the openness commitment and the price positioning were company claims at that point.[1] On the consumer side, the model replaced Hailuo 2.3 as the current model of the [Hailuo AI](https://aiwiki.ai/wiki/hailuo) product, whose earlier Hailuo entries MiniMax's platform documentation now lists as legacy models,[17] and it launched the same day inside third-party creative tools: [Krea](https://aiwiki.ai/wiki/krea_ai) announced H3 in its video tool on July 31, telling users the model "matches seedance's features at a fraction of the cost", a marketing comparison with ByteDance's [Seedance](https://aiwiki.ai/wiki/seedance) that Krea did not support with published measurements.[15][16]

The weights followed three days later. On August 3, 2026, MiniMax published the model repository publicly on Hugging Face and ModelScope (the repository had been created privately on July 28) and announced the release under the title "Open General Intelligence: MiniMax H3 Is Now Open Source".[2][6][8] The Hailuo AI account described it as "an open-weight video model, built for the community and verified to run on GPUs including the RTX 5090 and RTX 6000", a hardware compatibility claim MiniMax made without publishing supporting benchmarks.[7] The release landed in a crowded news window: The Decoder noted that ByteDance shipped its competing Seedance 2.5 the same day, with support for longer 30-second clips.[12]

## Capabilities

H3 accepts a multimodal context and generates video and audio jointly. The documented output envelope is 4 to 15 second clips at 24 frames per second, in aspect ratios including 21:9, 16:9, 4:3, 1:1, 3:4 and 9:16, with 32 kHz stereo audio in which speech, sound effects and music are modeled together rather than as separate tracks.[1][3] The model card states stable dialogue support for 11 languages: Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian and Spanish.[3]

The open release ships as two task-specific checkpoints of the base model, both CFG-distilled and in BF16 precision:[3]

- **H3-Base-FL2VA** covers text-to-video and first/last-frame generation: no input image gives [text-to-video](https://aiwiki.ai/wiki/text_to_video), one image anchors the first or last frame, and two images anchor both.
- **H3-Base-Ref2VA** covers omni-reference generation: up to nine reference images, up to three reference video clips (2 to 15 seconds each, 15 seconds total) and up to three reference audio clips, capped at 12 files per request, with audio permitted only alongside image or video input. Reference video plus text instructions also covers video-to-video editing, and reference audio supports voice-timbre-driven speech in the output.

Native output resolution is 768 pixels on the short side. 2K output is produced by a separate stage, H3-Regenerate-2K, in which the base model regenerates its own low-resolution result with the original multimodal context still in hand, in place of a conventional super-resolution module; MiniMax argues this recovers detail such as small text that upscalers can only guess at.[1][3]

## Architecture

The complete H3 system consists of three modules, of which only the middle one is open:[3]

1. **H3-Context-IR**, a hosted preprocessing and orchestration system that parses free-form multimodal input, performs cross-modal association and temporal reasoning, and serializes the result into a "Context Intermediate Representation" consumed by the base model. MiniMax describes it as critical to output quality, relies on multiple hosted models to run it, and did not include it in the open release; it is exposed as an API, and prompting guides are provided for developers who want to approximate it.
2. **H3-Base**, the open-weights generator. It encodes text through the H3-Encoder (the full Qwen3-VL-32B vision-language model, with hidden states drawn from the 50th layer), visual input through both the encoder and a visual VAE, and audio through an audio VAE, then packs everything into a single multimodal token sequence processed by the H3-Omni-Transformer, a 33 billion parameter dense single-stream transformer that predicts video and audio latents jointly. Roughly 13 billion of those parameters sit in AdaLN modulation branches whose outputs can be precomputed and cached, so they need not be loaded for inference-only use. Positional structure uses three-dimensional multimodal rotary embeddings (MM-RoPE) over time, height and width. The model was trained with native sparse attention for long sequences, but the initial release supports full-attention inference only, with the sparse implementation promised later.
3. **H3-Regenerate-2K**, the in-context regeneration stage for 2K output, likewise not yet open-sourced and available only through MiniMax's API.

The H3-VisualVAE is a temporally causal video autoencoder with 16x spatial and 4x temporal compression and 24 latent channels; a further 1x2x2 patchification takes the effective spatial downsampling to 32x before the transformer. The H3-AudioVAE compresses each of the two stereo channels of 32 kHz audio into latent tokens at 40 Hz, processing the channels independently through shared encoder and decoder weights.[3] In the July 31 announcement MiniMax additionally described a captioning approach ("Contextual Omni Representation") that distills roughly 100,000 tokens of inference about source material into about 4,000 tokens of description, credited its rebuilt VAE with a fourfold gain in effective sequence length, and said separating understanding and generation workloads lifted end-to-end training throughput by nearly 30 percent; these are the company's own engineering claims, made ahead of any full technical report.[1]

## License

H3 is released under the MiniMax H3 Community License Agreement, a custom license dated August 2, 2026, issued by the MiniMax entity Nanonoble Pte. Ltd. and governed by Hong Kong law.[4] Its most unusual term is territorial: the license applies worldwide except in the "Excluded Territories" of the European Union, the United Kingdom, South Korea and the United States, and it states that use, reproduction, modification, distribution or display of the model or its outputs outside the licensed territory is not authorized.[4] Organizations in excluded regions can apply to MiniMax for a separate formal license, which the company says it grants after reviewing the deployment's compliance controls.[5]

Other notable terms:[4]

- Commercial use is permitted, but licensees whose commercial products and services exceed 20 million US dollars in yearly revenue must obtain separate prior written authorization from MiniMax.
- Commercial products or services using H3 must prominently display "MiniMax H3" on their user interface.
- The model and its outputs may not be used to improve any other AI model, other than H3 and its derivatives.
- Redistribution requires passing on the license and a copyright NOTICE file; a "Powered by MiniMax H3" notice and AI-generation identifiers on output files are encouraged but not required.
- MiniMax claims no rights over generated outputs.
- An acceptable use policy bans, among other things, military use, high-risk automated decision-making, impersonation without consent, and undisclosed machine-generated content in public channels.

In an accompanying Q&A, MiniMax explained the territorial carve-out as a response to regulatory exposure specific to video generation, citing the EU AI Act, comparable uncertainty in the UK and South Korea, and, for the US, the fact that the company "is also involved in ongoing copyright-related legal proceedings specifically concerning generative video AI", a reference to the copyright suit filed against MiniMax by Disney, NBCUniversal and Warner Bros. Discovery in September 2025 (covered in the [MiniMax](https://aiwiki.ai/wiki/minimax) article). The company characterized the limitation as meaning "not yet", not "not ever": the hosted API stays globally available, and the license scope is to be revisited as regulation settles.[5]

MiniMax's own release post calls H3 "open source", but the combination of territorial limits, the revenue threshold, use restrictions and withheld system components makes "[open weights](https://aiwiki.ai/wiki/open_weights)" the more precise description.[2][4]

## Benchmarks and reception

The most-cited independent yardstick is the [Artificial Analysis](https://aiwiki.ai/wiki/artificial_analysis) Video Arena, which ranks models by Elo scores from blind pairwise human votes. As of August 5, 2026, in the default with-audio views, H3 ranked first in video editing (Elo 1,130, ahead of Google's Gemini Omni Flash at 1,122), second in text-to-video (Elo 1,239, behind Gemini Omni Flash at 1,244), and third in image-to-video (Elo 1,187, behind ByteDance's Dreamina [Seedance 2.0](https://aiwiki.ai/wiki/seedance_2) 720p at 1,196 and Gemini Omni Flash at 1,194), with the small gaps at the top of each board leaving nearby models statistically close.[9][10][11] Artificial Analysis marks H3 as open weights on those boards, and The Decoder reported on August 3 that this made H3 "the first open model to top an AI video ranking".[11][12] On the text-to-video board, H3's normalized price of $7.80 per minute of 1080p-equivalent output was listed against $9.07 for Seedance 2.0 720p, $20.16 for [Kling](https://aiwiki.ai/wiki/kling) 3.0 Pro and $21.00 for SkyReels V4.[9]

Uptake of the open release was rapid: within two days of publication the Hugging Face repository had accumulated over 1,900 likes and more than 20 community Spaces built on the model.[8] Day-zero inference support was coordinated with the main open-source serving stacks: the model card links official deployment recipes for SGLang and vLLM, a dedicated diffusers pipeline (MiniMaxH3ModularPipeline), and native ComfyUI workflows for text-to-video, image-to-video and reference-to-video, which require ComfyUI 0.30.0 or later and cap the native canvas at a 768-pixel short edge (up to 768x1344).[3][8][14] The Decoder noted the practical limits of the release: because H3-Context-IR and the 2K stage remain closed, self-hosted deployments must build their own context preprocessing and cannot reproduce 2K output locally, though the open weights do allow fine-tuning on custom footage, characters or styles.[12]

## Availability and pricing

The full three-module system runs on MiniMax's own surfaces: the Hailuo AI web and mobile apps for consumers, and the MiniMax Open Platform API for developers, where H3 is invoked as model `MiniMax-H3` through the v2 video generation endpoint, with companion endpoints exposing H3-Context-IR and H3-Regenerate-2K so self-hosters can mix local base-model inference with the hosted stages.[3] As of August 5, 2026, MiniMax's pay-as-you-go pricing listed H3 output at $0.13 per second at 2K and $0.08 per second at 768p, with input audio free, the first five input images free ($0.04 per additional image), and input video billed by duration at the output resolution's per-second rate.[13]

Third-party availability began before the weights shipped: [Krea](https://aiwiki.ai/wiki/krea_ai) offered H3 in its video tool from July 31, 2026, listing 2K output, clip lengths selectable from 5 to 15 seconds and tagged multi-reference input.[15][16] For self-hosting, the weights are on Hugging Face and ModelScope in both the original checkpoint layout and diffusers format, in BF16, with MiniMax recommending SGLang, vLLM, diffusers or ComfyUI for serving; MiniMax's example SGLang configurations use four GPUs, while its RTX 5090 and RTX 6000 compatibility statement addresses the consumer and workstation end.[2][3][7]

H3 sits in the same competitive field as ByteDance's [Seedance](https://aiwiki.ai/wiki/seedance) line, Kuaishou's [Kling](https://aiwiki.ai/wiki/kling), Google's [Veo 3](https://aiwiki.ai/wiki/veo_3) and successors, OpenAI's [Sora](https://aiwiki.ai/wiki/sora), and the open-weight Chinese lines Wan ([Wan 2.5](https://aiwiki.ai/wiki/wan_2_5) and successors from Alibaba) and Tencent's [HunyuanVideo](https://aiwiki.ai/wiki/hunyuan_video); within the open-weight group, H3 is the first to lead an Artificial Analysis video board.[9][12]

## See also

- [MiniMax](https://aiwiki.ai/wiki/minimax)
- [Hailuo AI](https://aiwiki.ai/wiki/hailuo)
- [Open weights](https://aiwiki.ai/wiki/open_weights)
- [Artificial Analysis](https://aiwiki.ai/wiki/artificial_analysis)
- [AI video generation](https://aiwiki.ai/wiki/ai_video_generation)
- [Seedance](https://aiwiki.ai/wiki/seedance)

## References

1. MiniMax. "MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities." July 31, 2026. https://www.minimax.io/blog/minimax-h3
2. MiniMax. "Open General Intelligence: MiniMax H3 Is Now Open Source." August 3, 2026. https://www.minimax.io/news/minimax-h3-open-source
3. MiniMaxAI. "MiniMax-H3" model card. Hugging Face. Accessed August 5, 2026. https://huggingface.co/MiniMaxAI/MiniMax-H3
4. MiniMax (Nanonoble Pte. Ltd.). "MiniMax H3 Community License Agreement." August 2, 2026. https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE
5. MiniMax. "Q&A About License." MiniMax-H3 repository. Accessed August 5, 2026. https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/QA-about-License.md
6. MiniMax (@MiniMax_AI). "MiniMax-H3 Is Now Publicly Available." X, August 3, 2026. https://x.com/MiniMax_AI/status/2084106804032872591
7. Hailuo AI (@Hailuo_AI). "An open-weight video model, built for the community and verified to run on GPUs including the RTX 5090 and RTX 6000." X, August 3, 2026. https://x.com/Hailuo_AI/status/2084123156818915727
8. Hugging Face. API metadata for MiniMaxAI/MiniMax-H3. Accessed August 5, 2026. https://huggingface.co/api/models/MiniMaxAI/MiniMax-H3
9. Artificial Analysis. "Text to Video Leaderboard." Accessed August 5, 2026. https://artificialanalysis.ai/video/leaderboard/text-to-video
10. Artificial Analysis. "Image to Video Leaderboard." Accessed August 5, 2026. https://artificialanalysis.ai/video/leaderboard/image-to-video
11. Artificial Analysis. "Video Editing Leaderboard." Accessed August 5, 2026. https://artificialanalysis.ai/video/leaderboard/video-editing
12. Schreiner, Maximilian. "China's MiniMax H3 is the first open model to top an AI video ranking." The Decoder, August 3, 2026. https://the-decoder.com/chinas-minimax-h3-is-the-first-open-model-to-top-an-ai-video-ranking/
13. MiniMax API Docs. "Pay as You Go." Accessed August 5, 2026. https://platform.minimax.io/docs/guides/pricing-paygo
14. Comfy Org. "ComfyUI MiniMax H3 Native Workflow Tutorial." Accessed August 5, 2026. https://docs.comfy.org/tutorials/video/minimax/minimax-h3
15. Krea (@krea_ai). "introducing MiniMax H3." X, July 31, 2026. https://x.com/krea_ai/status/2083014535481598410
16. Krea. "MiniMax H3 by MiniMax, AI Video Generator." Accessed August 5, 2026. https://www.krea.ai/models/minimax-h3
17. MiniMax API Docs. "Models." Accessed August 5, 2026. https://platform.minimax.io/docs/guides/models-intro

