AI Filmmaking28 min read

How I Made a Japanese-Style Anime With AI, From A to Z

Two episodes, 38 minutes of finished anime, one person. Episode 2 cost $1,852 and 65 hours. The full pipeline, in depth: the three documents every character needs, why the list of things a character never does is the most valuable page you will write, one line of dialogue traced from the character sheet to the finished frame, the four causes of unwanted voice shouting, and how to rebuild a 3D previz board from a finished film.

Try Screenweaver
Storyboard view of a Japanese-style anime made with AI: nine previz panels and the screenplay side by side
ScreenWeaver Logo
Frank Houbre

September 19, 2026

The short answer: two episodes, 38 minutes of finished anime, one person. Episode 2 cost $1,852 and 65 hours. The generator was never the bottleneck. Deciding was. Everything below is the actual pipeline, including the parts that failed and what replaced them.

Lost Garden is a dark fantasy series about empty suits of armour who wake up twenty centuries after the end of the world. Episode 1 runs 17 minutes, episode 2 runs 21. No studio, no crew, no outsourcing.

I am not going to tell you AI made this easy. It did not. What it removed was the wait. A project like this used to die at "I need to find the budget first", and then six months go by and you have moved on. It never died because I could not find someone who can draw. People who can draw are everywhere and they are better at it than any model.


Watch the result first

Everything below is process. Here is what it produced, so you can judge whether the method is worth reading.

Episode 1, The Awakening of the Lantern Knight, 17 minutes: youtu.be/eZ_JlaLDJ-8

Episode 2, The King Beneath the Vault, 21 minutes: youtu.be/z-YRrutXaFE

Both are at lostgarden.world with the press kit.

Six finished frames from Lost Garden: Bourdon's flame beard, the Source Tree, the castle on its cliff, Rose inside the medallion, the mice on their shelf, Serrure mid-leap with both blades

Those are finished frames, straight out of the episodes, no retouching. Keep them in mind while you read, because every screenshot after this point is the machinery behind them.


The one rule that shaped everything else

Cut every two to three seconds.

AI video falls apart when a shot runs long. Hands drift, faces slide, backgrounds breathe. Every model does it, and the failure is not linear: a shot is fine at two seconds, soft at four, and unusable at eight.

So I stopped letting shots run long. Not as a stylistic choice, as a survival tactic. Episode 2 has around 400 shots in 21 minutes.

Then something happened that I did not plan. The constraint turned into the style. Anime already cuts fast. Holding two seconds and cutting hard reads as deliberate direction rather than as damage control. The thing I did because I had to is the thing people compliment.

Write this into your process before you generate a single frame, because it changes the shot count, which changes the board, which changes the script.


Step 1: the script and the board have to be the same document

The single biggest time sink in this kind of production is not generation. It is the round trip.

You write in one tool. You board in another. You change a line, and now the board is wrong and you will not notice for two days. You change a shot, and the script no longer describes what happens.

ScreenWeaver puts the script vertically and the shot sequence horizontally, in the same window, in sync. Change a line and the shot moves with it. I never exported a screenplay to go board it somewhere else across two episodes.

The storyboard view on the bell foundry scene, with Bourdon and the mice as entities

And here is that same scene in the finished episode:

Bourdon in the finished episode: stone face, pale eyes, a beard of live flame lighting the bells behind him

That screenshot is scene 5 of episode 2. On the left, the entities in the scene. In the centre, the selected shot with its value, its camera movement, and the intention line. On the right, the actual screenplay. At the bottom, every shot in the scene with the playhead on the current one.

The important part is that none of those four panels can drift from the others, because there is only one document.

If you already have a screenplay, you do not retype it. Episode 2 came in as a Final Draft file and landed as 214 elements, seven characters and six places, already separated.

The AI Center, importing the episode 2 Final Draft file and listing the generations in progress

The characters and locations it extracts on import become the entities you are about to describe, which is the step most people do by hand and then skip halfway through.


Step 2: character sheets, and why one document is never enough

This is the step people skip, and it is the reason most AI films look like a folder of beautiful unrelated images.

It is also the step I got wrong on episode 1 and rebuilt properly for episode 2. Every character in Lost Garden now has three separate documents, and they do three jobs that cannot be merged.

  1. The model sheet. What the character looks like, from every angle.
  2. The pinned description. The paragraph the image model actually reads.
  3. The play sheet. How the character moves, speaks, reacts, and what they never do.

Skip the first and your character drifts. Skip the second and every shot is a negotiation. Skip the third, which almost everybody does, and your character looks perfect and behaves like a different person in every scene.

The model sheet

Model sheet for Serrure: front, profile and back view of a dark plate armour with a keyhole helm and a gold key on the breastplate

Three views, flat white background, flat cel colour, no dramatic lighting, no environment, no pose.

Every one of those choices is deliberate, and they are all the opposite of what you want to do.

Three views, not one. Front, profile, back. The back view is the one that saves you, because it is the view you never think about until you need an over the shoulder shot and discover you never decided what the back of the cape does.

White background, no lighting. A model sheet is a specification, not a picture. The moment you light it dramatically, the model starts copying the lighting into every shot that references it. You want the object, not the mood.

No action pose. A character standing still in a neutral stance reads as "this is what they are". A character mid-leap reads as "do this again", and you will get variations of that leap forever.

Model sheet for the Lantern Knight: white lantern helm with a carry ring, tan cape, ornate pauldrons, three views

Put the props on the sheet, because props drift first

Look at Bourdon's sheet. He is a stone giant who casts bells, and the sheet carries a pagoda roof structure on his back with three bells hanging from it, a hammer and a pair of tongs at his belt, and a small brass teapot.

Model sheet for Bourdon: stone giant with a flame beard, a bell frame on his back, hammer, tongs and a brass teapot, with a head detail inset

That teapot has nothing to do with his function. It is there because it tells you who he is in one object, and because props are the first thing to drift. A character's silhouette survives thirty shots. Their accessories mutate by shot four: the hammer becomes an axe, the tongs become pliers, the teapot disappears entirely.

Draw them on the sheet, name them in the description, and you get them back.

Note also the head detail inset at the top of the sheet. Bourdon appears in close-up several times, and a full body turnaround does not carry enough information about a face at that scale. If a character has a close-up in the episode, the sheet needs a head study.

Design the silhouette so it survives being small

Model sheet for Barrik: a barrel body with two eye holes, iron bands, short armoured legs, front and back

Barrik is a barrel with legs. At forty pixels tall he is still unmistakably Barrik.

That is not a joke about the design, it is the design. In a series that cuts every two seconds and uses a lot of extreme long shots, a character who is only recognisable in medium close-up is a character you cannot stage.

The test I use: shrink the model sheet to the height of your thumbnail strip. If you cannot tell who it is, the silhouette is not finished. Serrure is a tall thin vertical with a keyhole. Lanterne is a small block with a ring on top. Bourdon is a mountain with a roof on it. None of them can be confused at any size.

Barrik also carries the single hardest rule in the series, written at the top of his sheet in capitals: he has no head and no helm. Nothing sits on top of him. The barrel is his head and his body at once, and his eyes are two holes drilled in the shell. No dome, no hat, and above all never the lantern helm, which belongs to the Lantern Knight alone. Without that line written down, every generation eventually tries to give him a head, because barrels with faces are not in the training data and knights are.

The pinned description is not the model sheet in words

The second document is a single paragraph, and it is the one the image model reads. Here is Bourdon's, unedited:

Stone giant, the size of a gatehouse. Pale blind-looking eyes, a beard made of live flame, cracks that glow when he laughs. Casts bells for a kingdom that no longer answers. Laughs first, thinks out loud, decides slowly.

Two things about that paragraph matter more than they look.

It contains measurable facts. "The size of a gatehouse" is a constraint I can check in a frame. So is "cracks that glow when he laughs", which is a conditional: they glow when he laughs, not permanently. That is the sort of sentence that keeps a character alive instead of static.

The last sentence is not about appearance at all. "Laughs first, thinks out loud, decides slowly" is close to useless to an image model, and it is the most valuable line in the paragraph, because it is what keeps his dialogue consistent across twenty minutes of episode.

The play sheet: the document nobody writes

The third document is the one that made the difference, and it is written for me and for the voice performance, not for any model.

Mine open with a line that sets the boundary: play sheets, not design sheets. The exact design stays in the model sheet. This document answers a different question, which is how to write a line that sounds like this character.

Some of what is actually in Serrure's:

Size is a ratio, not a number. The canon fixes one thing: Serrure is half a head taller than the Lantern Knight. Never more. The sheet notes the metres as a proposal, explicitly flagged as awaiting approval, because a metre value is a guess and a ratio is a fact. Guesses and facts should not be written in the same tone in a document you will trust for two years.

A physical vocabulary. With Lanterne he is physical in specific, repeatable ways: two taps on the breastplate, a flick on the lantern helm, a hand on the shoulder, turning the helm by hand to show him where to look. That is a list of reusable gestures. It is what stops a character from having new body language every episode.

A speech rate. Warm, medium low, slightly rough. About two words per second, with real pauses. Then the line that actually directs the performance: a line that ends too early is right, a line that runs is wrong.

A writing tic. This is the single most useful thing in the document. Serrure reassures, and then undercuts his own reassurance three seconds later, half voice, without turning around. "I am never wrong... almost never." "It is only a demon, if it goes badly I will handle it... probably." Once that tic is written down, anyone can write a line that sounds like him, including me at two in the morning.

A rare signal, protected. When Serrure stops joking and his voice drops flat, something is genuinely wrong. The sheet says it plainly: this contrast is the best dramatic tool in the series, and it only works if it stays rare. That sentence has stopped me from using it three times.

The most valuable part of a play sheet is the list of things the character never does

Every one of my sheets ends with a "never" section, and it is the part I consult most.

Serrure never shouts to dominate. He never makes speeches. He never plays the king or the hero. He never stumbles, never loses his footing, never gets surprised by the scenery. Threatened socially, he answers with a stone smile and calm words rather than steel.

The Lantern Knight's list is stricter, because he is the harder character:

  • He never speaks. Not one word, ever, not even a whisper. He produces only hollow metallic sounds: soft creaks, resonances that read as breathing, clicks, timid little knocks, a questioning hiss, and a crash when something frightens him.
  • He never trips. He does not slip, does not fall, does not walk into the scenery. His comedy comes from curiosity, never from clumsiness. He walks and climbs steadily.
  • He never produces light himself, and his medallion is never visible.

That second rule is the one I have had to defend most often, because a small round character in a big dangerous world is constantly being offered a pratfall. The rule exists because a character who trips is a character you pity, and this one has to stay brave. The sheet says it in one line: he advances in front of danger, and he is never the first to run.

For the other one, here is what "never stumbles, never loses his footing" produces at the end of episode 1, when Serrure arrives to pull the Lantern Knight out of a fight he was losing:

Serrure in the finished episode, mid-leap in blue light with both blades drawn

All of his emotion has to pass through the body instead: helm tilts, posture, hand gestures, hesitations. And the carry ring on top of his helm oscillates when he moves quickly, which is what plays the trembling when he panics. That ring is a prop doing an actor's job, and it only works because it is on the model sheet.

The rest of the cast

Model sheets for the Vault King, the Unhooker, the Dark Knight and the Barded

Same treatment for everyone who appears more than once, including the antagonists and the creature the arena fight is against. A one scene character still gets a sheet if that scene has more than three shots, which in a two second cutting rhythm means most of them.

Write the description once, generate the sheet once, and reference it forever. That is why my two knights still look like themselves in shot 380.


Step 3: let the scene propose its own coverage

A blank shot list is the slowest page in filmmaking.

Generate shots from scene script reads the scene you already wrote and proposes the coverage. I keep maybe 60 percent and rewrite the rest, but it is far faster than a blank page, and it reliably catches the shot I always forget.

Generate shots from scene script on the arena scene, with the shooting plan written above the proposed shot list

The field above the shot list is the part that actually does the work. It is where you write your shooting plan in plain language before anything is proposed. Here is the one I used for the arena fight in episode 2:

Stay on the ledge with Serrure for the whole fight. We never take the fighter's point of view, only the point of view of the two who are watching him.

That single sentence decided nine shots. It is also the entire dramatic idea of the scene: the fight is not the subject, the two characters watching a stranger fight are. Without that line the breakdown gives you a competent action scene. With it, you get the scene the episode needed.

Write the intention, not the shots. The shots follow.

Here is the same view on a completely different scene, the arrival at the castle. Ten shots, different entities, different intention: scale first, faces later.

The storyboard view on the hall scene, ten previz panels, the Vault King as an entity

Step 4: name a compositional device in every prompt

"Cinematic" gives you nothing. It is the most common word in AI film prompts and it carries no information.

Name the device instead. Negative space. Chiaroscuro underlight. Deep two shot staging. Extreme long with the subject at one eighth frame height.

The shot panel for shot 3.6, with the shot value, the camera movement and the action line

Each shot in Lost Garden carries three fields, and all three go into the prompt:

  • Shot value. Extreme long, long, medium long, medium, medium close, close-up, extreme close-up, insert, over the shoulder, POV.
  • Camera movement. Static, slow push-in, slow dolly, whip pan, handheld.
  • The action line, which is the only free text, and which references entities by handle.

That last part is what keeps a 400 shot episode coherent. The action line for shot 3.1 reads:

@serrure and @lantern-knight flat on the ledge, looking down into the pit. Hold the wide until the blade answers.

The @ handles resolve to the reference sheets. The second sentence is a note to myself about timing that no model reads and that saved the shot in the edit.


Step 5: the reroll economy

Generation is cheap. Rerolls are cheap. Discovering in the edit that you never had a scene is not.

The expensive mistake in AI filmmaking is not a bad generation, it is generating beautifully for a week and finding out you have a mood board instead of a sequence. That is decided long before the model is ever opened.

Concretely, on episode 2:

  • Roughly 400 shots kept.
  • Around 2.4 generations per kept shot, so a bit under a thousand generations.
  • Total generation spend was the largest line of the $1,852.

The ratio is the number to watch. If you are rerolling six or eight times for one usable shot, the problem is almost never the model. It is that the prompt did not name a device, or the shot did not point at a reference sheet, or the shot should not exist.


Step 6: generate the picture silent, then lay the voices over it

This is the decision I would defend hardest, and it started as an accident.

My protagonists are empty suits of armour. They have no mouths. There is no lip sync to match, ever.

That means picture and dialogue are completely decoupled. I generate the image silent, cut the episode, and then perform the voices over a locked picture. No waiting on audio before I can see whether the scene works. No regenerating a shot because the mouth does not fit a line I rewrote.

If you are designing an AI series from scratch, think very hard about this before you design your characters. Helmets, masks, creatures, animals, anything without a visible mouth buys you an enormous amount of production freedom. It is the single best structural decision in Lost Garden and I did not make it on purpose.

What I learned about voice tags the hard way

Both episodes use ElevenLabs v3. Four things took me an embarrassing number of takes to work out.

Tag effect decays. A style tag does not hold for a paragraph. It fades within a sentence or two. If you want a character grave for twelve lines, the tag goes on every sentence, not once at the top.

Event tags replay. [laughs], [sighs], [chuckles] fire a sound each time they appear. Repeating them the way you repeat a style tag gives you a character laughing eleven times in a row.

Four different things cause unwanted shouting. This cost me the most time, so here they are in full.

  1. A sentence that starts like an imperative gets performed like one. "Open, in the dark, it's a candle on a table" came out as a shout. Rewritten as "To them, an open medallion in the dark is a candle on a table", it came out calm. Same information, different attack.
  2. A word placed immediately before a pause gets hit. Move the emphasised word into the middle of the sentence.
  3. Fragments of one to three words get shouted almost every time. Merge them into the preceding sentence with an ellipsis and a lowercase continuation.
  4. Consecutive short sentences escalate. "A fire. A meal. News from the road." produced three rising attacks. Merged into one sentence with ellipses, it produced the wistful line I wanted.

Pauses are the rate control, not the tag. <break time="2.5s" /> does more for a performance than any adjective. ElevenLabs caps a single break at three seconds, so chain them for longer holds. And the speed slider is the real tempo knob: 0.8 for my talkative knight, 0.7 for the King beneath the vault.

Every voice file for both episodes lives with its tags and its breaks, one file per character, so a re-record is a copy and paste rather than a rediscovery.


Try it free

Try Screenweaver for free on your script

It is free. Import your existing project, get a clearer view of your outline, and regain control of your story structure in minutes.

Start Free

One line, all the way through

Everything above is easier to believe as a single trace. Here is one line of dialogue from episode 2, from the character sheet to the finished frame.

1. The play sheet says what kind of line this is

Serrure's sheet says he thinks out loud, that he asks questions he already knows the answer to, and that a line which ends too early is right. It also says the moment he stops joking and drops flat is the rare signal.

This is that moment. He has just heard something in the pit. So the line has to be short, and it has to be a statement he is making to himself before he makes it to anyone else.

2. The screenplay

EXT. THE ARENA - CONTINUOUS

A voice comes up out of the pit. SERRURE stops dead.

                    SERRURE
              (startled, hushed)
        ...That's an armor voice.

                    SERRURE
             (shouting with joy)
        It's one of ours! Quick, Lanterne, come on!

Four words carry the whole episode's theme: they have not met another one of their kind in a very long time. The parenthetical is not decoration, it is the direction for the voice pass that comes four steps later.

3. The shooting plan, before any shot exists

Stay on the ledge with Serrure for the whole fight. We never take the fighter's point of view, only the point of view of the two who are watching him.

4. The shot

Shot 3.1. Long, static, low angle. Action line:

@serrure and @lantern-knight flat on the ledge, looking down into the pit. Hold the wide until the blade answers.

The @ handles resolve to the model sheets. "Hold the wide until the blade answers" is a note about timing that no model reads and that decided the cut point in the edit.

5. The voice, tagged

This is the actual text sent to ElevenLabs, unedited:

[calmly] ...You hear that? <break time="1.5s" /> [amazed] That's an armor
voice. <break time="2.0s" />

[shouting] [excited] It's one of ours! <break time="0.8s" /> [shouting]
[excited] Quick, Lanterne... come on! <break time="2.0s" />

Four things in five lines, and every one of them is a rule from further up this article.

The leading ellipsis on "...You hear that?" stops the attack. Without it the line starts loud and the whole beat is lost.

The tag repeats on the second sentence. [amazed] is not inherited from [calmly], and [calmly] would not have survived to the second sentence anyway. Tag effect decays.

"Quick, Lanterne... come on!" is one sentence, not two. As two short fragments it came out as two escalating shouts. Joined with an ellipsis, it reads as one impulse.

The two second break before and after the reveal is doing more work than any adjective. The silence is the performance.

6. The result, and the board rebuilt from it

Board on the left, final render on the right, four shots of the arena scene

Top left is shot 3.1: two knights flat on a ledge, looking down, held wide. Same framing in the board and in the film, because the board was made from the film.

That is the whole pipeline in one shot. A behaviour rule produced a line, the line produced a parenthetical, the parenthetical produced a voice direction, the shooting plan produced the framing, the framing produced the shot, and the shot produced the board.


Step 7: never let the video model write the music

Every single video prompt in Lost Garden contains the instruction no music.

If you let each clip generate its own audio bed, thirty-nine clips sound like thirty-nine different films. You cannot fix that in the mix. You cannot even fade between them.

Score the whole episode separately, as one piece of work, over a locked picture.

Six original tracks came out of that pass, written to picture after the cut was locked:

Working titleLengthWhere it sits
Inexorable Ascent5:00the long climb, the scale of the world
Song of the Empty Sky5:00the open passages, the world as it is now
Clockwork Requiem3:30the machines, the things that still run
The Giant and the Knight3:30the bell foundry, Bourdon
Ash Lantern Prayer3:18the quiet vigils
The Knight's Lullaby2:30the lullaby motif, reprised across both episodes

They were renamed for release, which is worth saying because it is a small thing that matters: a working title tells you where a cue sits in the episode, a release title has to work as a standalone track on a music platform. Those are two different jobs and they deserve two different names.

The order matters more than the tool. Cut first, score second. Every film has done it that way for a century, and AI tempts you to abandon it because the audio button is right there.


Step 8: the edit, and the one encoder setting that matters

The edit is DaVinci Resolve, the free version. That line of the budget is zero.

One technical note that cost me a full render to learn. If your source is a long master and ffmpeg sits at zero percent CPU doing nothing, it is not hung, it is trying to probe the file. Give it more room:

ffmpeg -probesize 20M -analyzeduration 10M -i master.mov ...

And when you mix down, -x264-params no-fast-pskip=1 is worth the few percent of bitrate. Without it, large flat dark areas, which is most of this series, develop blocking that looks like a frozen frame.


Step 9: subtitles, and a trick for catching model hallucinations

Both episodes are subtitled in English and French, burned in for the vertical cuts.

The transcription is Whisper large-v3-turbo with word level timestamps, which is what lets the subtitle land on the word rather than near it. The anti hallucination flags matter on a film with long musical passages:

--word-timestamps True --condition-on-previous-text False \
--no-speech-threshold 0.45 --logprob-threshold -0.8 \
--hallucination-silence-threshold 1.5

Even with those, episode 1 produced six phantom "I'm sorry" lines over instrumental sections.

Here is how to catch them for free. Re-transcribe the suspicious passage on its own, as an isolated clip. Real speech transcribes identically every time. A hallucination does not: the same eight seconds of music that gave me "I'm sorry" in the full pass gave me "I'm not going to die" in isolation. Disagreement between passes is the signature. That one check removed every phantom line and kept every real one, including a single whispered "Find me" at eighteen seconds that I would otherwise have deleted as noise.

Two more things, because they are not obvious:

Subtitles go on before rotation, never after. If you are making a vertical cut by rotating a landscape master, burn the subtitles into the landscape image first. Rotate afterwards and your text is sideways.

Measure your letterbox before you place the text. My master is 1920 by 1080 with the actual image at 1920 by 820, so there are 130 pixels of black top and bottom. Sitting the subtitles in that black bar means they never cover the picture. On black, drop the semi opaque box and use a thin outline instead. The box is invisible on black and it looks cheap where it shows.


Step 10: the part nobody films, distribution

Twenty one minutes of anime is not a TikTok. The full episode goes to YouTube. The vertical cut, rotated to 9:16 with burned subtitles, goes to TikTok and Instagram.

One practical finding that saved me an entire re-edit: on Instagram, a video post allows sixty minutes while a Reel allows three. The full episode goes up as a post. I had already cut a ninety-nine second excerpt as a Reel before I found that out, and it now sits in the drafts as a fallback I did not need.


The bonus step: boarding the film after the film

Here is the part that started as a favour to myself and turned into the most useful asset in the project.

Both episodes are finished. There was no storyboard left, because the board became the film. So for pitching episode 3, and for explaining the process, I rebuilt the board from the finished episodes.

Not by drawing. By taking one frame every ten seconds and converting each one into a raw 3D previsualisation blockout.

The raw material looks like this, one frame every ten seconds, named by timecode so any shot is one click away:

A contact sheet of episode 2, one frame every ten seconds through the bell foundry sequence

These two visuals work as a pair: the first shows A contact sheet of episode 2, one frame every ten seconds through the bell foundry sequence, and the second shifts to The bell foundry scene in previz: grey clay blockout, the fire flagged in flat orange - compare them briefly, then move on.

The bell foundry scene in previz: grey clay blockout, the fire flagged in flat orange

That is the forge scene. Untextured grey clay, low poly environment, featureless mannequins, and the fire flagged in a single flat orange because that is what previz does: everything grey, one colour for the thing the shot is about.

The arena fight uses cyan, and it only marks the blade, so your eye follows the weapon from shot to shot.

The arena fight in previz, nine panels, the blade flagged in cyan

And the arrival at the castle uses orange again, this time flagging every flame and every glowing eye, so you read where the light is before you read anything else.

The arrival at the castle in previz, ten panels, flames flagged in orange

The recipe, in full

Image to image, the finished frame as reference, aspect ratio forced to the source format. About one credit per frame.

Convert this frame into a raw 3D previsualization blockout render. Keep the
EXACT same camera angle, framing, composition and silhouettes as the reference.
Everything untextured matte grey clay, one single material, no line art, no
anime shading, no colour. Characters become smooth grey mannequins: keep the
silhouette and overall head shape readable, but no facial features, no ornament,
no fabric detail. Environment becomes simplified low-poly grey primitive
geometry, blocky, almost no detail. Soft neutral studio light, plain light grey
background, soft ambient occlusion. ONE EXCEPTION: <the key element> is a single
flat saturated <colour>, unshaded, like a previz colour tag. Everything else
stays pure grey clay.

Two phrases carry the entire result.

"Keep the EXACT same camera angle, framing, composition", without which the model recomposes the shot and your board no longer describes your film.

"No facial features, no eyes, no mouth", without which it redraws eyes and you get grey anime instead of previz.

One warning that cost me a frame. On a big close-up of a face, that second instruction is taken literally and the head becomes a smooth unrecognisable mass. On those shots, ask instead for the overall shape and the eye sockets to stay readable, with no texture and no detail.

Why this is worth doing on a finished film

Board on the left, final render on the right, four shots of the arena scene

Because a filter cannot do it. I tried first, and the honest result is worth describing: I wrote an image processing pass that kills the line art, kills the detail, and then removes the albedo itself by dividing the image by its own broad trend, so that a black breastplate and a white sleeve land on the same grey. It produces a convincing clay wash. It is also not previz, because no filter erases a drawn face or rebuilds geometry. The eyes of my tree spirit stayed anime eyes, blurred.

Regeneration rebuilds the form. That is the whole difference, and it is why the comparison above holds up: same framing, same silhouettes, different medium.

If you are pitching a season on the strength of one finished episode, this gives you a board to pitch with and a before and after that proves the shot was designed rather than found.


What it cost, and where the time went

$1,852 and 65 hours for episode 2. Sixty five hours is about eight working days for twenty one finished minutes.

The largest cost line is generation credits. Voices and music come next. The editing suite is free. The costs are real but they are not the interesting number.

The interesting number is the ratio between deciding and generating. Every hour I spent on the board saved something like five hours of generating shots I did not need. That is the whole economics of this kind of production, and it is why the pipeline above starts with the script and the reference sheets rather than with the generator.


What I would do differently

Write the entity descriptions before the first shot, not after the tenth. I learned consistency the expensive way on episode 1 and did it properly on episode 2.

Decide the cut length before the shot list. Two to three seconds per shot is not a post decision. It doubles your shot count and it belongs in the breakdown.

Design at least one main character without a mouth. Whatever your story is. The production freedom it buys is out of all proportion to the constraint.

Keep the board after you finish. I had to rebuild mine from the finished film. It worked, and it made a better pitch asset than the original board would have, but that was luck rather than planning.


Where to go next

If you want the pipeline in the abstract rather than this case study, start with the AI filmmaking workflow from script to screen. If you want to watch it happen, how to make an AI film is the walkthrough. If you want the boarding step on its own, how to generate a storyboard from a screenplay with AI covers it properly.

And if you take one thing from the whole of this: the tools are not the bottleneck. Deciding is.

Both episodes are free to watch. Episode 1, episode 2, and everything else at lostgarden.world. Episode 3 is being written in ScreenWeaver right now, and Barrik, the barrel, is the one carrying it.

Final Step

Build your next script with Screenweaver

Move from ideas to production-ready pages faster with timeline-native writing and AI-assisted story flow.

Try Screenweaver
ScreenWeaver Logo

About the Author

The ScreenWeaver Editorial Team is composed of veteran filmmakers, screenwriters, and technologists working to bridge the gap between imagination and production.

Continue reading