Direct the ingredients, not just the description

MiniMax’s July H3 demonstration combines a camera-motion reference, a character image and an audio reference. The instruction assigns a different job to each input. The resulting clip shows the café character performing while the camera changes its perspective. The illustration above is the supplied character reference; the accompanying social video shows the generated result.

That is a useful distinction for creators. A reference image can describe appearance more precisely than a paragraph, while a separate clip can communicate movement. The prompt explains how those ingredients should relate.

A better way to evaluate the demo

Start by asking which detail each reference is supposed to control. Then judge that detail in the output: does the character remain recognizable, does the camera perform the requested movement, and does the background remain coherent? A polished clip alone cannot answer whether every instruction was followed.

For your own experiment, keep the character reference fixed and change only the motion reference. Comparing those outputs would make the effect easier to inspect than changing every input at once. This is a suggested evaluation method, not a test we performed.

The scope of the example

MiniMax describes H3 as accepting text, images, video and audio context, with output up to 15 seconds at 2K and native stereo sound. These are the company’s stated capabilities. Our edited Reel focuses on visual references; its source audio is omitted, so it does not demonstrate audio quality. This is a July release being examined for its workflow, not a launch today.

Explore the original source ↗

Source published 2026-07-31. Coverage is based on the maker’s announcement and demonstration.