Audio Plan
What it does
Builds the scene-by-scene sound design for your video — voice delivery, pauses, ambience, effects, music dynamics, and deliberate silence — as an editing timeline. Audio is where AI videos stop feeling like AI videos.
When to use it
- Visuals are generated, and you’re about to record or generate voiceover
- Your drafts feel flat even though the pictures are good (it’s the audio)
- Before the edit, so sound is designed rather than sprinkled
The skill
Act as a sound designer and voice director. Principles: the voice
carries the story, effects make the world believable, music carries
emotion, and silence is a tool — recommend it where it beats sound.
Write voice direction a non-actor can follow.
The script with scenes: [PASTE]
Voice tool I'll use: [ELEVENLABS / OTHER / RECORDING MYSELF]
Overall mood: [DESCRIBE]
For every scene:
1. The exact narration/dialogue, split per speaker — each
character's lines separated so I can generate voices
individually, never as one long file.
2. Voice direction: tone, pace, energy, and where the pauses go
(mark them in the text with [pause]).
3. Ambience bed: the constant background of this scene.
4. Spot effects: max 3 per scene, tied to visible actions.
5. Music: mood, intensity, and where it rises, falls, or drops out.
6. Any moment where silence does the work.
Then: the assembly timeline — scene by scene, the layer order and
rough levels (voice / effects / ambience / music), plus the 3
places in the whole edit where audio should peak.
Example output
[TO FILL AFTER TESTING]
Tweaks
- Generate voiceover scene by scene even when it’s one narrator — pacing control beats convenience
- Doing shorts? One ambience, one effect, one music decision. More gets muddy at phone volume
- Foley beats library music for believability: footsteps, cloth, doors first, soundtrack last