MiniMax H3 Max Is Here — I Pushed It to the Limit
Yaroflasher
Create 5–15 second 480p/768p videos from text, first and last frames, or multimodal reference media. MiniMax H3 Max supports a dedicated reference-to-video workflow with images, video and audio.

MiniMax H3 Max is a video generation model in the H3 family for text-to-video, image-to-video with a first frame, a last frame, or both, and multimodal reference-to-video. The current EvoLink routes support 480p or 768p MP4 output from 5 to 15 seconds. Reference mode accepts images, videos and audio, making H3 Max useful when a shot needs stronger identity, motion or timing guidance than a text prompt alone.

Choose text, keyframes, or multimodal references based on how much control your shot needs.
Generate a 5–15 second clip from a prompt with 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16 framing.
Animate from a first frame, a last frame, or both when you need a defined visual start and finish.
Use up to 9 images, 3 videos and 3 audio clips, with no more than 12 reference assets in one request.
Use 480p for lower-cost exploration or 768p when you need a cleaner review-ready result.
Yaroflasher
ElevenLabs
AI Search
Excelerator
Joseph Martin
Benji’s AI Playground
MiniMax H3 Max — Video scene creation FAQ
MiniMax H3 Max is an H3-family AI video model for text-to-video, first/last-frame image-to-video and multimodal reference-to-video generation.
Reference mode supports up to 9 images, 3 videos and 3 audio clips, with 12 assets maximum in total. At least one reference image or video is required; audio alone is not accepted.
The current routes support 480p and 768p MP4 output with integer durations from 5 to 15 seconds.
H3 Max adds multimodal reference-to-video with image, video and audio inputs. H3 Max Turbo is the faster, lower-cost tier for text-to-video and first-frame image-to-video workflows.
The generator shows the current estimate before you submit. Cost depends on resolution and duration; reference mode can also include reference-video seconds and extra reference-image charges.
Start with text, keyframes, or multimodal references and generate directly in the browser.
10,000+ users