MiniMax H3 Prompt Structure
Turn an idea into an executable audiovisual timeline
MiniMax H3 generates video and audio together. Its official base guide uses three fields. ZHENSHENG first creates an editable draft, then Enhance Details converts supported choices into that public English structure.
Source review: . Examples are original editorial drafts, not measured generation results.
- 01Choose the H3 input mode and duration first
- 02Describe observable change in chronological order
- 03Separate on-screen sound from background music
Match duration, references, and the local task
The official H3 repository documents 4–15-second base generation. This builder offers 4, 6, 10, and 15-second presets. Match the local task duration to the prompt, especially the last-frame alignment time.
You can validate a silent single-shot action before adding sound. For precise object handling, start with a reference that clearly shows the hand and prop. Text cannot replace missing visual detail.
H3 documents five modes; ZHENSHENG currently supports three
The builder covers T2VA, I2VA, and FL2VA. L2VA and Ref2VA need extra input about the purpose of each asset, so this version does not generate empty reference labels.
- T2VA builds the complete audiovisual timeline from text.
- I2VA uses the first frame as the opening state and describes what happens next.
- FL2VA fixes the first and last frames and describes the path between them.
- L2VA supplies the last frame and requires a plausible lead-in.
- Ref2VA assigns images, video, and audio to subject, motion, camera, rhythm, or sound references.
Base modes use three fields
integrated_multimodal_description carries the visual scene, subject, action, camera, dialogue, and on-screen sound. overall_soundscape summarizes ambience, physical sound, and non-verbal vocalization. non_diegetic_music describes background music heard by the audience.
The official guide asks for structured descriptions in English. Dialogue, lyrics, and visible text may retain their original language. Use N/A for non_diegetic_music when there is no non-diegetic music. Use N/A for overall_soundscape only when the user explicitly requests complete silence.
I2VA maps Picture 1 to 0.00 seconds and Shot 1 before the three fields. FL2VA maps Picture 1 to 0.00 seconds and Picture 2 to the actual end of the video. The duration selected on this page must match the H3 task parameter.
Give the timeline a start, an observable change, and an endpoint
A continuous shot can use three stages: establish, execute, and settle. State the initial condition, causal order, object ownership, and final position. A first-last prompt should reduce differences in pose, location, composition, and lighting as it approaches the final frame.
The official camera vocabulary distinguishes movement type, optional amplitude, and speed. Write those choices as a complete sentence so the camera follows one spatial path.
Separate ambience, physical sound, and music
H3 produces native audio. Decide which sounds exist inside the scene, then decide whether the audience hears additional music. Café room tone, a cup contact sound, and a musical underscore have different jobs.
ZHENSHENG currently offers natural ambience, physical sounds only, soft music, rhythmic music, and silence. Dialogue needs speaker identity, language, and exact lines, so this version does not invent dialogue tags.
A draft and an H3-structured version of the same idea
The example below was written for ZHENSHENG. Use the draft to check the idea, then use the structured version when you are ready to test the local workflow.
Create a 10-second MiniMax H3 video. In a quiet café in the afternoon, a middle-aged man in a charcoal suit sits at a table. He lowers his gaze toward a white coffee cup, grips the handle with his right hand, lifts the cup steadily to chest height, and stops. Use a medium shot with a slow push-in and realistic cinematic texture. Keep café ambience and the cup contact sound. Use no background music.
integrated_multimodal_description: [Shot 1] Live-action, cinematic. The target video lasts exactly 10 seconds and remains one continuous shot. From 00:00.000 to 00:02.000, a middle-aged man in a tailored charcoal suit sits still at a wooden café table and lowers his gaze toward a white coffee cup. From 00:02.000 to 00:08.000, he leans forward slightly, reaches with his right hand, closes his fingers around the handle, and lifts the cup steadily to chest height while keeping it upright. From 00:08.000 to 00:10.000, he stops the movement and holds one stable final pose. The camera pushes in with small amplitude at slow speed from a medium shot. Preserve identity, clothing, cup geometry, hand-to-handle contact, light direction, and café layout. overall_soundscape: Soft café room tone, distant cup clinks, and the precise contact sound of the hand and cup; all physical sounds remain naturally synchronized. non_diegetic_music: N/A
Check causality, space, and sound before submitting
Check who owns each object, where the action travels, and which visible event motivates each sound. A cup must leave the table before its base loses contact; the return-to-table sound should coincide with the next visible contact.
Check the ending state as well. A person still holding a cup cannot simultaneously wave with that same empty hand. This is an editorial checklist, not a measured claim of improved generation success.
The open-source release does not include the hosted Context-IR implementation
MiniMax describes H3-Context-IR, H3-Base, and H3-Regenerate-2K as parts of the full system. The company says H3-Context-IR interprets and expands free-form text and multimodal input. Its implementation is not included in the open-source release; MiniMax provides the service through an official API.
ZHENSHENG formats a user's choices with the public prompt rules. It does not reproduce H3-Context-IR, and no official evidence shows equivalent output quality.
References
Model features can change. Check the current official pages before relying on a version-specific capability.
- MiniMax-AI. MiniMax H3 repository. Accessed 5 Sept. 2026.
- MiniMax-AI. base-mode prompt guide. Accessed 5 Sept. 2026.
- MiniMax-AI. full-reference prompt guide. Accessed 5 Sept. 2026.