AI Filmmaking15 min read

How to Make an AI Music Video: The Method (+ Real Example)

The song is the script: a time map, shot lengths from the BPM, a script by section, the artist's reference sheet, a storyboard, two generation lanes and a cut on the waveform. Shown on a real six-shot demo, with a long section on keeping one artist consistent for three minutes.

Try Screenweaver
Two Neon Temple music video shots from ScreenWeaver, each grey 3D previz storyboard frame beside its final shot: the priestess under tall windows, and dancers in a ring seen from above

The short answer: treat the song as the screenplay. Lock the final mix, mark every section with a timecode, and turn the tempo into shot lengths (one bar lasts 240 divided by the BPM, in seconds). Then build the artist's reference sheet, storyboard each section, generate shots a little longer than their bar, and cut them on the waveform.

A good AI music video is decided on a timeline before anyone writes a prompt. Most failed clips generate first, then force the pictures onto a song nobody measured. This guide is the music video chapter of our how to make an AI film walkthrough: what a song changes, from a structure you did not write to a performer who carries it for three minutes.

What you need

  • The track, final mix, with the rights to use it. The music comes from the artist or a licensed source, not from the video tool.
  • Time: several working days for a full three-minute clip. A 15 to 30 second vertical teaser fits in a weekend.
  • Budget: writing the script is free in ScreenWeaver. Storyboards and video use credits, so the bill depends on shots kept and takes per shot. Plans are on the pricing section.
  • Hardware: a computer with a browser, plus an editor such as DaVinci Resolve, Premiere Pro or CapCut.
  • Level: you need to count bars. No drawing, no 3D skills.

What makes a good music video

Five rules come before any tool, because a clip that breaks them looks wrong whatever model made it.

1. The song is the script. A film script gives you the story and you choose the timing. A track gives you the timing and you find the story inside it.

2. Shot length comes from the tempo. In 4/4, one bar lasts 240 divided by the BPM, in seconds: 2 seconds at 120 BPM, 1.6 at 150. Verses hold a shot for two bars, a chorus cuts on every bar, a drum fill on half a bar.

3. Two tracks, intercut. The performance track is the artist delivering the song. The story track is the world around them. Plan them separately and fix the ratio early.

4. One motif per section. A light, a place, a costume or a colour that tells the viewer where they are in the song. The second chorus grows out of the first instead of repeating it.

5. Lip sync is a decision, not a hope. Video models do not reliably land a mouth on the syllables of a real vocal. Decide up front which shots show singing and how you will handle them.


The method, step by step

Step 1: Lock the track, the rights and the idea

Nothing starts until the master is final. A two-bar change in the bridge moves every timecode on the board, and generated shots do not stretch. Get the exact version and duration in writing.

Then settle the rights. If the song is not yours, you need permission from whoever owns the composition and the recording. On YouTube, rights holders can register a track in Content ID, and a matching upload can then be blocked, monetized for the owner, or tracked (YouTube Help, "How Content ID works", checked 8 October 2026).

Last, write the idea in one sentence: "A priestess wakes a stone temple with a drum, and the temple answers with light." If you cannot say it, the clip becomes a reel of pretty shots. Our treatment guide turns it into a pitch the artist can approve.

Step 2: Build the time map

This is the step most AI music videos skip. Load the track into your editor, drop a marker at every section change, and fill one row per section. Here is the start of a time map for an example song at 120 BPM, where one bar lasts 2 seconds:

SectionTimecodeBarsCut rhythmShotsMotifWhat the viewer must understand
Intro0:00 to 0:168one long take1smoke, candlesthe world before the artist
Verse 10:16 to 0:4816two bars per shot8window lightwho sings
Chorus 10:48 to 1:2016one bar per shot16neonthe signature image

Now do the arithmetic. A three-minute song at 120 BPM has 90 bars. At two bars per shot that is about 45 shots, closer to 70 if the choruses cut on every bar. That number is your shot list and the base of your budget, before you spend a credit.

Our article on story beats in a music video then gives each section one visual turn.

Step 3: Write the clip as a script

Write the clip as a screenplay where the sections are the scene headings. Here is the start of the Neon Temple demo in ScreenWeaver's free screenplay editor:

VERSE 1 (0:15) - INT. NEON TEMPLE - NIGHT

Under three tall windows she lifts her face. The spiked halo catches
the light. She sings to the lens.

HOOK (0:44)

Her palm lands on the drum on the downbeat. Dust jumps off the skin.

Write one action per shot, because a model asked for three things in two seconds does none of them well, and tag each shot as performance or story. When the script is done, the AI shot breakdown turns each section into a list of shots you can edit, reorder and time against your map.

Step 4: Build the reference sheets

Before the first storyboard panel, create every recurring element as a reference: Characters (the artist, dancers, story characters who come back), Places (each set, with its lighting states if they change) and Objects (a drum, a car, a necklace, anything the viewer must recognise). You can generate the sheets or upload your own images, including photos of the real artist if they agreed to it. The method is in the consistency section below.

Step 5: Fix one art direction

Pick one look and attach it to every generation: one or two style frames, a few light and colour rules, the aspect ratio.

Decide the ratios now too. Most artists need a 16:9 master and 9:16 cut-downs of the hook for Reels, TikTok and Shorts. A 16:9 performance shot rarely survives a vertical crop, so plan vertical shots as their own frames.

Step 6: Storyboard every section

Board the clip section by section, with a time in and a time out on every panel. In ScreenWeaver the storyboard is generated from the shot breakdown with the reference sheets attached, and it renders as a grey 3D previz.

Six shots of the Neon Temple music video, each grey 3D previz storyboard frame above its final shot: the nave in smoke, the priestess under the windows, hands on the drum, dancers in a ring from above, arms raised among candles, the last smile to camera

The six shots of Neon Temple, a demo clip the ScreenWeaver team made for our music video page: storyboard frame above, final shot below.

Then cut the boards into an animatic on the master. If chorus 2 reads like chorus 1, you find out in an afternoon, not after three hundred generations. For panel counts per section, see the music video storyboard guide. The AI storyboard generator page shows the boards in the app.

Step 7: Generate shot by shot, in two lanes

On the production canvas, two pipelines turn the board into video. Reference to video turns each shot's action into a directed prompt and attaches the characters, places and objects cited in the shot. Storyboard panels to video uses up to 14 boarded panels of a section to guide the order and composition of the generated sequence, alongside the references.

The default model is Seedance 2.5 (clips from 4 to 30 seconds, up to 1080p, up to 30 image references). Seedance 2, Veo 3.1, Kling 3.0, MiniMax H3 and Grok Imagine Video are also there. Choose per shot. Three rules apply:

Generate longer than the bar. A 1.6 second shot is not a generation length. Generate four seconds or more, describe a single action, and keep the stretch where the motion is clean.

Use the track as a reference, not as a promise. Seedance 2, Seedance 2.5 and MiniMax H3 accept an uploaded audio file as a reference input, so you can give the model the section it illustrates. Treat it as context for mood and energy. Do not count on it to land a cut on a beat or a mouth on a syllable. Timing is fixed in the edit.

Budget takes by lane. Story shots usually converge in a few takes. Performance shots, where a face, hands and an instrument all have to be right, take more. Keep the best take, not the first.

Step 8: Sound, when the song is already finished

ScreenWeaver does not generate music or voices, and a music video does not need it to: the track exists. Most video models can return ambient sound with a shot. Mute it, or keep a breath of it under an intro or a breakdown.

Step 9: Edit on the waveform, export both ratios

ScreenWeaver does not edit or export the final cut. Download your kept shots and assemble them in DaVinci Resolve, Premiere Pro or CapCut, on the master, never on a scratch mix. The real work is choosing which 1.6 seconds of a four-second generation land the hit. Make strobes in the edit, on the beat, from clean lit states, because generated flicker tends to smear.

Keep one frame rate end to end. If a label delivers through Vevo, its published HD spec asks for 1920x1080 in H.264 or ProRes, a native frame rate of at least 23.98, AAC stereo audio and no slates (Vevo, "Music Video Specs", checked 8 October 2026). Then cut the 9:16 versions from the frames you planned.


Keeping the artist the same for three minutes

A music video is the hardest consistency test in AI video. A short film might show its lead in fifteen shots. A clip shows the artist in thirty or forty, often in close-up, under changing light, and every chorus invites a comparison with the last one.

Why the artist drifts

A video model has no memory between generations. Write "a woman with long beaded braids and a metal headdress" in forty prompts and you get forty women who all match the sentence. Words describe a category, not a person, and more detail does not fix it. We go deeper in character drift in AI video.

The rule: attach the reference, do not re-describe it

Consistency comes from images that travel with every shot. In ScreenWeaver, the artist is a Character with a reference sheet. You cite that character in the shot's action line, the sheet follows the shot into the storyboard, and the same references are attached when the shot becomes video. The prompt only says what changes in this shot: action, framing, light.

The artist kit

Build it once, before the first panel:

  • A hero shot: front, neutral expression, even light. Every other image is checked against it.
  • A turnaround: front, three-quarter, profile, back. Clips use profiles and back shots far more than films do.
  • Outfit close-ups: fabric, cut, colours, and every accessory that matters, like a headdress, jewellery or an instrument.
  • Key expressions: singing with an open mouth, eyes closed on a held note, the final smile.
  • One sheet per outfit: if the artist changes look between sections, each look gets its own sheet, derived from the hero shot so the face stays the same.

The AI character consistency sheet tool helps you list what the sheet must show, and our guide to storyboard character consistency explains how to carry it panel to panel.

Performance shots and story shots are different problems

Performance shots show the artist to camera, singing, playing or dancing. The face is large and the viewer compares it to the last chorus. These shots need the full kit every time, tight framing and more takes. Keep them short: a held close-up is where drift and mouth problems show first.

Story shots show the world, other characters, hands, or the artist from a distance. In a wide, the viewer tracks the silhouette: outfit, hair, build. So the sheet still goes on. A wide shot without the reference is where the coat changes colour.

Outfits by section, decided on the board

A change of look is fine when it is planned. In Neon Temple, the boards give the priestess a spiked halo for the verse and a different crown for the chorus and the outro. That is a section motif written into the storyboard, not drift. What must never change by accident is the face, the braids and the skin tone.

The same priestess in three Neon Temple shots: under the windows in the verse, arms raised from behind in the chorus, smiling to camera in the outro

Verse, chorus, outro: three performance shots of the same character from the Neon Temple demo.

The drift test

Before generating a new section, put the first performance shot of the clip next to the latest one. Same face? Same braids? Same headdress for this section? Then compare chorus 1 and chorus 2 side by side, because that is what the audience does. Check on the storyboard first, where a fix costs one panel, then again on the video.

Real artist, generated performer, or no face

Decide this before any generation.

  • The real artist, filmed. Shoot the performance against a plain backdrop with playback, then generate the world around it. The mouth is real, so the lip sync is real. For slow motion, the video frames calculator does the playback speed arithmetic.
  • A generated performer. Build an original character, keep it consistent with the kit, and keep sung shots wide, in profile or from behind. For the few close-ups that need lip sync, plan a dedicated pass in a separate lip-sync tool, then check that the lips close on B, P and M.
  • No face. Hands, places, objects, distant dancers, text. No lip-sync pass, no likeness question.

A generated performer who looks like the real artist needs the artist's written consent. One who looks like any other real singer gets redesigned.


Try it free

Try Screenweaver for free on your script

It is free. Import your existing project, get a clearer view of your outline, and regain control of your story structure in minutes.

Start Free

One shot, taken apart

Shot 5 of Neon Temple, the first chorus, from page to screen.

Shot 5 of Neon Temple: on the left the grey 3D previz storyboard frame of the priestess seen from behind, arms raised among dancers and candles; on the right the final shot with white neon tubes lit
  • Time map: chorus 1, one bar per shot at 150 BPM, so 1.6 seconds on screen.
  • Script: "She raises both hands, a flame in each. On the next bar, every neon tube in the temple switches on."
  • Shot direction: "Wide from behind, arms raised, a flame in each hand, dancers facing her. The white tubes switch on halfway through the shot."
  • What the board fixed: the camera behind her, so the biggest image of the chorus needs no lip sync. Dancers facing her, so she is the centre without showing her face. Candles in the foreground, so the neon has something to change.
  • What the references fixed: the priestess's braids and build from behind, and the temple's two lighting states.
  • What the edit does: the light change lands on a downbeat. If a take had flickered, the fallback was two clean states, tubes off and tubes on, cut on the beat.

We did not log a take count for this demo, so we will not publish one.


What it costs and how long it takes

There is no honest single price, because the bill is shots kept times takes per shot.

Shots: from your time map. Around 45 for three minutes at two bars per shot, closer to 70 with a fast chorus.

Takes: the best public reference is still Paul Trillo's video for Washed Out's "The Hardest Part" (Sub Pop, 2024), made with Sora. According to fxguide, he generated roughly 700 clips, about 230 minutes of video, and used about 55 (fxguide, 8 May 2024, checked 8 October 2026). That is about 13 generations per clip kept, on a continuous-shot video and an early model. Treat it as a ceiling, and expect fewer takes when shots are boarded and references attached.

A paragraph from fxguide's May 2024 article on the Washed Out video, where Paul Trillo estimates he generated around 700 clips, about 230 minutes of video in total, and used about 55 of them, all rendered at 720p and upscaled to 2K

Credits: a generation's cost depends on model, length and resolution. Put your shot count and a take profile into the AI video cost calculator before you give the artist a number. Writing costs nothing in ScreenWeaver, and plans are on the pricing section.

Time: a day for the time map and script, one or two for sheets and boards, then generation in proportion to the shot count. For a long-form benchmark of the same pipeline, see the proof box below.


Common mistakes, and how to fix them

1. Generating before the time map. Every shot ends up the wrong length. Fix: map the track and count bars first.

2. Re-describing the artist in every prompt. Forty prompts, forty slightly different singers. Fix: one reference sheet, attached to every shot, wides included.

3. Counting on generated lip sync. Mouths drift off the words, and a held note makes it obvious. Fix: film the artist, keep generated sung shots wide or in profile, or plan a dedicated lip-sync pass.

4. Three actions in one shot. The model compromises and nothing lands on the beat. Fix: one action per shot, generated longer than the bar.

5. Stacking style adjectives. "Cinematic, moody, neon, gothic" gets averaged into a generic look. Fix: one art direction, attached as images.

6. Chorus 2 copies chorus 1. The clip stops growing halfway. Fix: same motif, new angles, one escalation, checked in the animatic.

7. Ignoring the rights. A track you do not control, a performer who looks like a real singer. Fix: written permission for the track, an original performer or the artist's consent.


Which tool to choose

Kind of toolWhat it does wellWhere it breaks on a music video
Standalone video generatorsStrong shots, newest models firstNo memory between clips, no time map, no storyboard
One-click music video appsA clip in minutes from a songGeneric visuals, little control over the artist or the cut
Production suites with script, storyboard and video in one projectTime map, shots, references and generations stay connectedYou still write, choose and edit

ScreenWeaver is in the third row: a screenplay editor, an AI shot breakdown, a storyboard and a production canvas, with the same Character, Place and Object sheets carried from the first word to the last shot. It does not compose music or cut the final edit. What it handles is where most AI clips fall apart: forty shots of one artist, in the right order, at the right length, with the same face.


Proof, not promises

Neon Temple is a demo the ScreenWeaver team made for our AI music video generator page: six shots cut one per bar at 150 BPM, 1.6 seconds each, in 16:9. The page shows every shot as its storyboard frame and its final shot, with its place in the song. It is a demo, not a commissioned video.

For the full pipeline over a long format, read the Lost Garden case study: two anime episodes, 38 minutes in total, made by one person from script to final shots in ScreenWeaver. Episode 2 runs 21 minutes and took 75 hours and $1,850.


Rights and YouTube labels before you publish

YouTube asks creators to disclose realistic content that makes a real person appear to say or do something they did not, alters a real event or place, or shows a realistic scene that did not happen. AI music that is the main focus of a video is on the list too. Clearly non-realistic content does not require it, and disclosure does not limit reach or eligibility to earn money (YouTube Help, "Disclosing use of altered or synthetic content", checked 8 October 2026).

YouTube's help page list of content creators need to disclose, including AI generated music, realistic AI footage of a real place, and making it appear as if a real person said or did something they did not

A photoreal generated performer singing the song qualifies. Our AI film disclosure rules cover other platforms.


FAQ

Can I make an AI music video for free?

Partly. Mapping the song and writing the clip script is free on ScreenWeaver's Screenwriter plan. Storyboards and video use credits. The planning, which decides the quality, costs nothing.

What is the best AI for music video lip sync?

No general video model does it reliably on a real vocal yet. Film the artist singing, keep generated sung shots wide or in profile, or run a dedicated lip-sync pass on the few close-ups that need it.

How do I keep the same character through the whole music video?

Use images, not descriptions. Build the artist's reference sheet before the first panel, attach it to every shot, and compare the first and latest performance shots before each new section.

How many shots does a three-minute music video need?

Count bars, not seconds. At 120 BPM, three minutes is 90 bars: about 45 shots at two bars per shot, closer to 70 if the choruses cut on every bar.

Can I make a vertical 9:16 music video?

Yes. ScreenWeaver works in 16:9 and 9:16, among other ratios. Plan the vertical shots as their own frames rather than cropping the master, and open the cut-down on the hook.

Can YouTube monetize an AI music video?

YouTube says disclosing AI use does not affect eligibility to earn money. What matters is controlling the rights to the track, since a rights holder can claim a matching upload through Content ID.


Start with the song

Your track already has a structure. Map it, write the clip and board every section in one project. Start writing for free on ScreenWeaver, then storyboard and generate once the time map holds.

Final Step

Build your next script with Screenweaver

Move from ideas to production-ready pages faster with timeline-native writing and AI-assisted story flow.

Try Screenweaver
ScreenWeaver Logo

About the Author

The ScreenWeaver Editorial Team is composed of veteran filmmakers, screenwriters, and technologists working to bridge the gap between imagination and production.

Film techniques

Film techniques in this article

Each one with a short clip, a screenplay excerpt and the shot list line.

All 263 techniques

Continue reading