minimax h3
Updated 2026-10-04
MiniMax H3 is MiniMax’s open multimodal video model for generating and editing short clips from text, images, video, and audio. It supports 768P or 2K output lasting 4–15 seconds, and the easiest way to test it is through a browser-based MiniMax H3 tool without installing software or using a local GPU. Developers can also use the pay-as-you-go API or investigate local deployment separately.
What Is MiniMax H3?
MiniMax announced H3 on July 31, 2026, describing it as its first open multimodal generation model. Unlike a basic text-to-video generator, H3 can work with several input types in one context:
- Text prompts
- First- and last-frame images
- Character and style reference images
- Existing video clips
- Audio references
- Prompts for video editing and motion transfer
The official MiniMax API guide lists three principal creation modes: Text-to-Video, Image-to-Video, and Reference Generation. H3 outputs 768P or 2K video at an integer duration from 4 to 15 seconds.
There is also a separate MiniMax H3 Max model. It is documented as a faster option that outputs 480P or 768P and supports 5–15 second clips. Select plain MiniMax H3 when you need 2K output, four-second clips, reference creation, or H3’s broader editing functions.
How to Use MiniMax H3 Online
Use this browser workflow for a first test:
Open the MiniMax H3 online generator.
Sign in if the page requests it. An online interface avoids model downloads, Python environments, and local GPU configuration.Confirm the selected model.
Check that the model menu says MiniMax H3, not MiniMax H3 Max. The two models have different resolution and duration ranges.Choose a generation mode.
- Select Text-to-Video for a scene created from a written description.
- Select Image-to-Video when a first or last frame should control the composition.
- Select Reference Generation when images, clips, or audio should guide character appearance, motion, camera work, style, voice, or editing rhythm.
Add the prompt and reference files.
Enter one connected shot rather than several unrelated scenes. If the interface numbers uploaded files, identify each one directly in the prompt—for example, “Use Image 1 for the character, not the background.”Set the output controls.
Choose the required aspect ratio, 768P or 2K resolution, and a duration from 4 to 15 seconds. H3 supports common aspect ratios as well as adaptive output; exact API options are listed in the API reference.Review before creating another version.
Check motion, subject consistency, camera movement, dialogue, on-screen text, and audio. If the interface offers 768P and 2K separately, validating the motion at the lower resolution can prevent spending the higher-resolution setting on an unsuitable take.
Interface labels may change, so confirm the model, file count, resolution, aspect ratio, and duration immediately before submitting.
Which H3 Mode Should You Use?
Text-to-Video
Use a text-only prompt when you are exploring a new location, visual style, abstract concept, or cinematic shot. Be specific about the subject, action, camera movement, lighting, and sound.
A useful structure is:
Subject → action in order → camera movement → setting and lighting → dialogue or text → ambient sound and music
MiniMax H3 Image-to-Video
Image-to-Video is usually the most practical starting point when appearance matters. The API permits zero, one, or two images in the first/last-frame workflow:
- One image: Use it as the starting frame and describe what happens next.
- Two images: Use one as the first frame and one as the last frame.
- Zero images: The workflow becomes Text-to-Video.
Input images must be between 256 and 5760 pixels in width or height, with an aspect ratio between 2:5 and 5:2.
Reference Generation
Reference Generation is for combining controlled assets. According to the official specification, one request may include:
- Up to 9 images
- Up to 3 video clips
- Up to 3 audio clips
- No more than 12 files in total
Each video or audio clip must be 2–15 seconds long, and their combined duration cannot exceed 15 seconds. Supported formats include H.264/AVC and H.265/HEVC video, JPG, JPEG, PNG, WEBP, HEIC, and HEIF images, and WAV or MP3 audio. The documented file-size limits are 50 MB per video and 30 MB per image.
Do not simply write “use everything.” Explain what each file controls. For example, use Video 1 only for walking speed, Video 2 only for the orbiting camera, and Audio 1 only for pacing—not for copying its melody.
Prompting Tips for Better Results
- Describe one continuous action. A 4–15 second clip works better as one shot than as a rapid sequence of unrelated events.
- Separate identity from motion. State whether a video reference should transfer body movement, camera direction, pacing, or all three.
- Quote exact on-screen text. Spell out capitalization and punctuation, then inspect every frame because higher resolution does not guarantee perfect small text.
- Name exclusions. If a reference’s clothing, background, lighting, or original character should not carry over, say so explicitly.
- Describe sound separately. Specify dialogue, ambient noise, effects, and music instead of assuming the visual prompt controls every audio detail.
- Start with 768P for testing. Once the movement and composition work, use 2K for the final version.
Online, API, or Local MiniMax H3?
Online access is the fastest option for creators because it requires neither a local installation nor a compatible GPU. If H3 appears inside Hailuo AI, select it from the current model menu; Hailuo AI is the web interface, while H3 is the underlying model name shown in that workflow.
The official H3 and H3 Max API options are pay-as-you-go. Launch coverage reported a MiniMax H3 price of 0.8 yuan per second for 2K video, but you should verify the current API billing page before purchasing. A third-party website may also offer free trial credits, but that does not make the official model universally free.
MiniMax describes H3 as open and promotes local deployment and optimization. However, the official API documentation does not establish a universal 32 GB RAM requirement or confirm that four rather than eight sampling steps is correct. Those details depend on the checkpoint, quantization, resolution, sampler, and current ComfyUI implementation. For local use, verify the model card, license, checkpoint format, and supported nodes before downloading a safetensors file.
FAQ
Is MiniMax H3 free?
MiniMax H3 is not documented as universally free. Its official API is pay-as-you-go, and launch coverage reported a rate of 0.8 yuan per second for 2K output. Individual browser platforms may offer trial credits or subscription access, but their prices and limits are separate from official API pricing.
What video inputs does MiniMax H3 accept?
H3 accepts text prompts, first and last frames, and reference images, videos, or audio. A reference request can contain up to 9 images, 3 videos, and 3 audio clips, with no more than 12 files total. Reference clips must be 2–15 seconds each and total no more than 15 seconds.
Can MiniMax H3 generate image-to-video and 2K clips?
Yes. You can animate a supplied first frame, guide the result toward a supplied last frame, or provide both endpoints. H3 supports 768P and 2K output lasting 4–15 seconds. Review character details, motion, text, and transitions because 2K resolution alone does not guarantee perfect results.
Is MiniMax H3 the same as Hailuo AI?
No. H3 is the model; Hailuo AI is a web interface where H3 may be offered as a model option. Availability and menu labels can change, so check the selected model name before generating. H3 Max is also a separate model with lower maximum resolution and faster generation.
Can MiniMax H3 run locally with 32 GB of RAM?
Local deployment is a separate path from online use, and the official API guide does not provide a universal minimum. Claims that 32 GB of RAM is sufficient—or that four or eight steps is the correct setting—must be checked against the released model card and current ComfyUI workflow. Quantization and output resolution can materially change hardware needs.
Try the MiniMax H3 AI Video Generator
Run the MiniMax H3 video model online – from a text prompt, in about a minute.
Generate a video