Generating a video with dialogue, sound effects, and music is not simply a matter of listing every sound you want. These elements compete for timing and attention. If the prompt is unclear, music may cover the dialogue, effects may occur before the matching action, or a character may speak the wrong line. I get more reliable results by treating the soundtrack as three coordinated layers and arranging them along the visual timeline. This guide explains how to structure that kind of prompt for MiniMax H3, with examples for dialogue scenes, commercials, and cinematic action.
Understand the Three Audio Layers
Before writing the full prompt, I separate the soundtrack into dialogue, sound effects, and music. Each layer has a different job.
Dialogue
Dialogue includes the words, speaker, delivery, pace, and pauses. I always identify who is speaking and place the exact line in quotation marks.
Instead of writing:
The woman says that she thinks someone is outside.
I use:
Maya looks toward the window and quietly says, “I think someone is outside.” Her voice is tense and restrained, with a brief pause before “outside.”
The second version gives the model information about the speaker, wording, emotion, and timing. I avoid long lines in short clips because the character needs enough time to speak naturally.
Sound Effects
Sound effects should be connected to visible sources. Footsteps belong to someone walking, a click belongs to a switch, and a splash should happen when an object touches water.
I divide them into two groups:
- Continuous ambience: rain, traffic, wind, café noise, machinery
- Timed effects: a door closing, glass breaking, a phone vibrating, a bottle opening
“Add cinematic sound effects” is too broad. Describing the source and moment produces a soundtrack that feels more connected to the image.
Music
Music controls the emotional direction and pace of the scene. A useful music instruction covers the genre, instruments, mood, intensity, and progression.
For example:
A restrained electronic score begins with a low synth pulse. It grows slightly after the character opens the door but remains quieter than the dialogue.
This is more actionable than simply asking for “dramatic music.” It explains how the score should change and which sound deserves priority.
Plan the Audio Before Writing the Prompt
I start by deciding which layer carries the scene. In a conversation, the voices are usually the priority. In a product commercial, narration and product sounds may share attention. In an action sequence, effects and music can lead while dialogue remains brief.
I then create a simple timeline. It does not need to describe every frame. It only needs to identify the major visual and audio events:
- 0–3 seconds: opening action, ambience, and first line
- 3–7 seconds: a new action introduces a sound effect
- 7–10 seconds: music builds toward the final visual beat
This planning stage is also useful when working in Loova Creative Studio because I can develop the scene, references, and prompt around the same creative direction. Instead of treating the soundtrack as an extra instruction added at the end, I can plan it alongside the characters, setting, and camera movement.

Timestamps should be treated as timing guidance rather than guaranteed frame-accurate controls. Clear event order is often more important than filling the prompt with narrow time ranges.
How to Prompt Dialogue, Effects, and Music in MiniMax H3
MiniMax describes H3 as an omni-modal video system that jointly generates video and native stereo audio. Its official release information lists video durations from 4 to 15 seconds and stable dialogue support across 11 languages. That makes it suitable for short scenes in which visual action and several types of sound need to develop together.
Establish the Visual Scene First
Audio needs a visible context. I begin with the setting, characters, camera position, and main action before describing what anyone says or hears.
For example:
A woman waits alone on an almost empty train platform at night. She stands beneath a flickering overhead light while the camera slowly moves closer. A dark railway tunnel is visible behind her.
This establishes possible sound sources: the light, the open station, the woman, and the tunnel. I keep this part focused. Too many visual events make it harder to coordinate the soundtrack.
Assign Every Spoken Line
For each line, I specify:
- Speaker
- Exact words
- Vocal tone
- Volume
- Pace or pause
- Relevant facial reaction
A two-character café scene might include:
Daniel looks at the untouched coffee and asks, “Are you leaving tonight?” He speaks quietly and slowly. Lena waits one second before replying, “I haven’t decided.” Her voice is calm but uncertain.
The pause helps separate the speakers. I would not add narration or a third voice unless it serves a clear purpose. In a ten-second clip, two short lines are usually more manageable than a complete conversation.
Link Effects to Visible Actions
I place each effect next to the event that produces it. This reduces ambiguity:
Lena places the ceramic cup on the saucer, creating one soft clink.
This is clearer than listing “cup sound” in a separate audio paragraph. I use the same method for movement:
Daniel slides the chair back. Its wooden legs briefly scrape against the café floor.
For environmental ambience, I describe it once and indicate that it remains in the background:
Low café ambience continues throughout, with distant conversations and occasional dish sounds. No background voices are individually understandable.
That final instruction prevents the ambience from becoming competing dialogue.
Describe Music as a Changing Layer
Music rarely needs the same intensity throughout an entire scene. I describe where it enters and how it develops.
For a suspense scene:
Begin with no music. After the footsteps start, introduce a low ambient synth almost silently. Let it build gradually through the final three seconds without overpowering the footsteps or dialogue.
For a product commercial:
Use a clean, upbeat electronic track with light percussion. Keep it soft beneath the voiceover, then raise it slightly after the final spoken line.
These instructions connect the music to the scene’s structure. They also prevent an energetic soundtrack from competing with important speech.
State the Audio Hierarchy
Even when every layer is described clearly, I add a short mixing instruction. It tells the model what viewers should hear first.
A practical hierarchy might be:
- Dialogue remains clear and in the foreground
- Action-related effects are distinct but controlled
- Environmental ambience stays subtle
- Music lowers during speech and rises between lines
- No additional dialogue, vocals, or unrelated effects
I avoid using technical mixing numbers unless they are genuinely necessary. Simple relationships such as “quieter than the dialogue” are easy to understand and adapt.
Combine the Instructions into a Timeline
Here is a complete prompt for a suspenseful station scene:
Create a 10-second cinematic video set on an almost empty train platform at night. A woman stands beneath a flickering overhead light while the camera slowly moves toward her. Keep the scene grounded and realistic.
0–3 seconds: Quiet wind moves through the station, accompanied by a faint electrical buzz from the light. The woman looks toward the dark tunnel and softly says, “The last train should have arrived by now.” Her voice is nervous but controlled.
3–7 seconds: Slow footsteps begin echoing from inside the tunnel. The woman stops moving and turns toward the sound. Introduce a low ambient synth almost silently after the first footstep.
7–10 seconds: A metal station sign rattles sharply in the wind. A distant train horn sounds once. The footsteps grow closer as the music rises slightly. End before anyone enters the platform.
Keep the woman’s dialogue clear and foregrounded. The wind and electrical buzz should remain subtle. Make the footsteps easy to locate and synchronize every effect with its visible cause. Do not add narration, extra voices, singing, or unrelated sounds.
The strength of this prompt comes from its order. The visual action, dialogue, effects, and music all move toward the same final beat.
Practical Audio Prompt Examples
Product Commercial
For a beverage advertisement, I would let the physical product sounds lead:
A cold glass bottle stands beside a clear tumbler filled with ice. A hand opens the cap with a crisp metallic click, then pours the sparkling drink over the ice. Use detailed fizzing, pouring, and ice sounds. A warm voiceover says, “A brighter way to refresh.” Keep the light electronic music below the narration, then raise it slightly as the logo appears.
This prompt gives each sound a source and protects the voiceover from the music.
Two-Character Conversation
For dialogue, I keep the environment restrained:
Two friends sit across from each other in a quiet café. Soft room ambience and distant dish sounds remain in the background. Nora asks, “Did you bring the letter?” Sam places an envelope on the table with a soft paper sound and replies, “I never opened it.” Use a subtle piano note after his response. Do not add intelligible background voices.
The sound of the envelope marks the transition between the two lines without requiring loud music.
Cinematic Action

When testing MiniMax H3 with action, I avoid asking for dialogue during the noisiest moment:
A cyclist races through a narrow street in heavy rain. Begin with close tire spray, rainfall, and fast breathing. The cyclist looks behind and shouts, “They’re still following me!” After the line ends, bring in urgent percussion as a car skids around the corner. Synchronize the tire squeal with the turn and keep the rain present throughout.
The spoken line happens before the main impact of the music and vehicle effect, so each audio event has room to register.
Common Audio Prompting Problems
Dialogue Is Covered by Music
State that speech is the primary audio layer. Ask the music to lower during every line and rise only after the speaker finishes.
Effects Happen at the Wrong Time
Move each effect next to its triggering action. If timing still drifts, simplify the scene or reduce the number of effects.
Characters Speak the Wrong Lines
Label every speaker and place the line immediately after that character’s action. Avoid writing several unattributed quotes in one paragraph.
The Scene Contains Extra Sounds
End the prompt with a short exclusion list, such as:
No narration, singing, crowd dialogue, extra voices, or unrelated sound effects.
Negative instructions can reduce unwanted additions, but they may not work perfectly in every generation. Always review the final audio rather than assuming silence or exclusivity has been followed.
Reusable Prompt Structure
For a new scene, I use this order:
- Describe the setting, characters, and camera.
- Divide the action into a few time ranges.
- Assign every line to a named speaker.
- Connect each effect to its visible source.
- Describe when music begins and changes.
- State which audio layer has priority.
- List any voices or sounds that should not appear.
This structure is flexible enough for dialogue, advertising, comedy, suspense, and short cinematic sequences.
Conclusion
Good audio prompting is mostly about hierarchy and timing. I separate dialogue, effects, and music, then reconnect them through visible actions and a simple timeline. Short lines, clearly sourced effects, and music that changes with the scene give MiniMax H3 a much clearer production plan. If a result is imperfect, I simplify or revise the weak audio event instead of making the entire prompt more complicated.
Frequently Asked Questions
Can MiniMax H3 generate dialogue, sound effects, and music together?
Yes. MiniMax H3 generates video with native stereo audio. A prompt can describe spoken dialogue, environmental ambience, action effects, and music, although results should still be reviewed for timing and accuracy.
How should I format dialogue in a MiniMax H3 prompt?
Identify the speaker, put the exact line in quotation marks, and describe the tone and pace. Place the dialogue next to the action performed by that character.
Should I use timestamps?
Timestamps help organize short scenes with several events. Use them as sequence and timing guidance rather than assuming they will control every frame exactly.
How do I stop music from covering speech?
State that dialogue is the primary audio layer. Ask the music to remain soft during speech and increase only during pauses or after the final line.
How many speakers should I include?
For a short clip, one or two speakers are usually enough. Limiting the cast gives each line more time and reduces confusion over who should speak.




