Lesson 6 of 7

Stop Making Flat AI Shots: Direct Them Like a Film

Advanced 12 min videoFree

Same story, same script, same tool, same character. The difference between a flat first pass and a directed sequence is about fifteen decisions. This lesson makes them.

Chapters (21)

Lesson notes

Why first-pass AI storyboards feel flat

Every shot is a medium. Every character is centred. Every camera is at eye height, straight on. There are no reactions: we never cut back to her face after something happens. And there is no progression: shot one is the same size as shot seven, so nothing intensifies. It is not ugly, and that is the trap. It is competent and completely inert. It is coverage.

The tools to fix it are the four shot fields from Lesson 5. Every one of them is a directing decision.

Shot size is emotion

  • Wide makes a person small and the world large. That is isolation.
  • Medium is neutral, which is why six in a row feel like nothing.
  • Close-up removes the world and leaves only a face: intimacy, or its threat, depending on what came before.
  • Extreme close-up is pressure.
  • Insert is the film pointing at something and saying: this matters.

SIGNAL's weak version opened on a medium of Mara. The directed version opens on an extreme long establishing shot, because the first thing the audience should feel is not "here is a woman" but "here is a very large, very cold room with one person in it".

Composition and depth, in plain words

In that opening shot Mara is not in the middle. She is small on the left, and the right two thirds of the frame is window and empty room. Centred, the room is a backdrop. Off-centre, the room becomes pressure. You do not type "rule of thirds" into the action field; you type where she is.

Mara is small at the console on the left of frame.

Depth works the same way. Name what is in front and what is behind: the edge of a desk in the foreground, Mara in the middle, the sealed door at the back. Three layers make an image feel like a room instead of a poster.

Inserts and over-the-shoulder shots

In the weak version the needle moving on the dial happens in the background of a medium shot, and nobody sees it. An insert is the closest thing film has to a sentence in bold. Use it when a small thing changes everything, and no more than two or three times a minute, or it stops meaning anything.

Over the shoulder is not just a way to show two people talking. It puts the audience behind someone, looking at what they look at. Mara is listening to a speaker grille; with the camera behind her shoulder, the audience listens too. Add a slow dolly in, because the moment is contracting.

Camera movement: one per shot, and keep it small

Every camera move is one more thing that has to survive the jump from still frame to video. A slow dolly in is reliable. A crane that becomes a whip pan is a coin toss. Static is not a failure: six or seven of SIGNAL's ten shots are static, so the few that move mean something. If everything moves, nothing moves.

  • Reliable: static, slow dolly in, small pan, gentle tilt.
  • Ambitious: crane, whip pan, handheld. Wonderful in a still, a coin toss as video. Save them for the shot that needs them, and never on a shot shorter than 4 seconds.

Reactions, eyelines and screen direction

The reaction shot is the one the weak version does not have, and the most expensive to leave out. Audiences do not experience an event; they experience someone experiencing an event. Write it as direction to an actor: what the face is doing, and that she has not decided anything yet.

If she looks off to the right, the next shot shows what is on her right. That eyeline is the cheapest continuity you will ever get. Keep screen direction too: on the left of frame looking right, she stays on the left. Cross that line and she seems to turn around between cuts. Use Continues from to tell the system frames share a setup.

Read the shot sizes as a shape

Read the shot values of the whole sequence in order. The directed SIGNAL reads:

Extreme long, medium, insert, medium close-up, over the shoulder, close-up, close-up, POV, extreme close-up, extreme long.

The weak one read medium six times. The directed one opens at the widest size, works inward, hits its tightest point, an extreme close-up on her thumb over the transmit key at the moment of maximum hesitation, then pulls all the way out. Get closer as tension rises, pull back when it releases. The tightest shot goes on the moment of most pressure, not of most information.

  • Duration is a free directing tool: a held shot is a held breath. Give the hesitation an extra second; give the insert half a second less and it snaps.
  • Angle regenerates the same moment from another camera position. Camera height is a value judgment: low makes things larger than us, high makes them smaller.
  • Close the shape: SIGNAL ends on an extreme long like it opened. Same room, same distance, but she is smaller, the door is bigger, and the window holds a reflection that should not be there.

Direct for what AI video can actually do

Every choice must survive video generation later. A character walking the length of a room in one take is difficult; standing still while the camera moves an inch is easy. Two people interacting is hard; one person and an object is easy. Intricate hand work is where AI video still falls over most often.

In the script Mara crosses the room. In the sequence we never see it: we cut from her deciding to a close-up of her hand arriving, and the audience does the walk. That is not a workaround. It is editing, used for a hundred years for the same reason. Ask on every shot: what is the simplest image that carries this beat? It is usually the better image too. Then take the directed storyboard into Lesson 7.

Try this step in ScreenWeaver

Direct one sequence

Open your first-pass storyboard, read the shot values in order, and change the sizes so the sequence gets closer as the tension rises.

Open your storyboard

Go further

Free tools

Guides