AI 视频片段规划器
把任意一场戏切成符合你模型时长上限的片段
粘贴一场戏、选好你模型的片段时长上限,这个规划器就会把它切成带编号、可直接粘贴的提示词,并且首尾相接成一段连续的序列。为那些模型停在 8 秒或 10 秒、而戏有 45 秒的 AI 电影人而造。
所有文生视频模型都只能一小段一小段地生成:Veo 3 大约到 8 秒,Sora 2 和 Kling 2.x 大约到 10 秒。更长的一场戏,只能以一串共用同一段一致性描述、并在每个切点交接状态的片段形式存在。规划器会估算每个节拍的银幕时长(动作按每秒约 3 个词,对白按每秒约 2.5 个说出口的词再加 20% 余量),然后把连续的节拍打包进不超过你上限的片段里。
全部在你的浏览器里运行。你的场次文本从不离开这里,也没有任何内容被发送到服务器。
模型预设只设定秒数,依据的是各模型在 2026 年年底的上限。
每个片段的秒数会逐字重复在每条片段提示词的最前面。它越具体,你的人物和场景在片段之间就越不会飘。
人物你的片段方案
把你的场次粘到上面,就能看到片段方案。
工作原理
规划器把你的场次切成节拍:每一句动作和每一次对白(一行「名字:」或一句带引号的话)都成为一个节拍。它按阅读节奏惯例估算每个节拍的银幕秒数,动作按每秒约 3 个词,对白按每秒约 2.5 个说出口的词再加 20% 余量用于呼吸和反应。然后它按顺序把连续的节拍打包进不超过你上限的片段里。每条片段提示词都带着同一段一致性表头、一行从上一个片段的 EXIT STATE 继承来的 ENTRY STATE、节拍文本和一个目标时长,这串片段于是能像同一场戏那样剪到一起。
怎么把一场戏切成 AI 视频片段
模型限制时长有两个原因:算力成本随片段长度陡增,而生成越久一致性就越飘(脸会变形、道具会瞬移、光会滑动)。所以没有任何设置能让 Veo 3 超过 8 秒,也没有办法让 Sora 2 在一次生成里更长。可靠地拉长 AI 视频的唯一办法是多片段工作:把戏切成节拍、分别生成每个片段,再剪到一起。
真正有效的串联方法既枯燥又可复制。在你的模型支持图像条件时,把一个片段的最后一帧当成下一个片段的第一帧,并且在每条提示词里重复完整的一致性描述。写「和刚才同一个角色」毫无用处,因为模型对上一次生成没有记忆;重复写「红色雨衣,黑色短发」才管用。正是这种重复保住了片段之间的一致性,而这也正是规划器该替你自动完成的事。
当你在动作上切开时,片段之间的接缝就消失了,这条剪辑惯例和连续性剪辑一样古老。让一个片段停在动作进行到一半(一扇正在推开的门、一个正在转过来的头),再让下一个片段把那个动作做完,观众的眼睛会自己跨过这个切点。每条提示词里的 EXIT STATE 和 ENTRY STATE 就是为了铺好这些交接,也正是它们让 Veo 3、Sora 2 或 Kling 的片段能串成一场连贯流动的戏,而不是一段幻灯片。
方案定下来之后,用我们的 免费 AI 视频提示词生成器 调好每个片段的摄影机和光线语言,让这一串提示词说同一套电影语法。
适合谁
- 做 AI 短片的人 把它当成场次规划器,在烧掉生成额度之前,先把剧本里的每一场戏变成带编号的镜头方案。
- 在 Veo 3 或 Sora 2 上单干的人 别再为了塞进 8 秒而删戏,开始用真的接得上的进出状态来规划片段串联。
- 代理商和社交团队 把一支 30 秒广告作为一串片段来做简报,整个团队都读得懂,一致性在第一次生成之前就已经锁好。
Scene to Clip Splitter: the complete guide
It segments a pasted scene into beats, estimates screen seconds per beat from word count, packs beats into clips under your max length, and writes one ready-to-paste prompt per clip with an identical continuity header.
For this workflow, the central problem is clear: text-to-video models cap generations at 8 to 12 seconds, so full scenes have to be planned as clip chains, and doing that by hand loses continuity and screen-time math. Left unresolved, this creates downstream friction and slower decisions. The practical target is a numbered clip plan with per-clip prompts, chained entry and exit states, and duration estimates that never exceed the model's cap.
Limitation to keep in mind: Word-count pacing is an estimate, not a stopwatch: performance, camera speed, and model interpretation shift real durations, so treat clip timings as planning targets and expect to trim in the edit.
Advanced workflow: Advanced users pair each clip prompt with last-frame conditioning: generate clip 1, export its final frame, feed it in as clip 2's start image, and keep the written continuity header identical so the model has both a visual and a textual anchor.
Step-by-Step Workflow
- Paste the scene text, dialogue included, and pick your model preset or a custom max clip length.
- Fill the continuity block with exact character looks, location, lighting, and style, since it repeats in every prompt.
- Review the clip cards: check durations, watch for over-cap warnings, and adjust beats that split awkwardly.
- Copy prompts clip by clip (or export the .txt), generate in order, and cut the results together on the action beats.
Use Cases By Profile
- AI short film maker: turn a script scene into a shot-by-shot generation plan before spending credits.
- Solo Veo 3 creator: plan an 8-second clip chain with handoffs that cut on action instead of jarring resets.
- Social team lead: brief a 30-second spot as a readable clip chain with continuity locked up front.
Common Mistakes To Avoid
- Writing "same character as before" instead of repeating the full description, which the model cannot remember.
- Packing clips to the exact cap with no handle, leaving nothing to trim when a generation runs short.
- Splitting mid-sentence on dialogue, which makes lip-sync and performance continuity nearly impossible.
Professional Best Practices
- Plan clips a second or two under the cap so every generation has trim room in the edit.
- End clips mid-movement and start the next completing it; cutting on action hides the seam between generations.
- Keep one master continuity block per project and paste it unchanged into every scene you split.
Treat this tool output as a decision support layer, not a replacement for authorship. Great scripts are remembered for specific choices, emotional precision, and clarity of dramatic movement. Tools help by removing noise so your energy can go where it matters: character, conflict, escalation, and payoff. If you review outcomes after each pass and keep an explicit log of accepted changes, your workflow becomes faster and more predictable from draft to draft. That consistency is exactly what professional collaborators value: fewer surprises, clearer rationale, and a script that evolves with intent.
Extended FAQ
Why do AI video models cap clips at 8 or 10 seconds?
Two reasons: compute cost grows steeply with duration, and temporal coherence drifts the longer a generation runs, so faces and props mutate. Short caps keep quality high, which is why longer scenes are built as clip chains instead.
What is the best AI video clip planner workflow for a full scene?
Split the scene into beats, pack beats into clips under your model's cap, repeat one continuity block in every prompt, chain entry and exit states, and generate in order. This tool automates the splitting, timing, and prompt assembly.
How does last frame to first frame chaining work?
You export the final frame of clip N and feed it to clip N+1 as a start image where the model supports it. Paired with a repeated text continuity block, it anchors character, wardrobe, and lighting across the cut.
Can I use this for Kling or Runway instead of Veo?
Yes. The presets only set the max seconds (10 for Kling 2.x and Runway Gen-4 as of late 2026). The prompts are plain text, so they paste into any model's prompt field.
How is clip duration estimated from my scene text?
Action is estimated at about 3 words per second of visualized screen time and dialogue at about 2.5 words per second spoken, plus 20 percent handles. Beats are clamped to a 1 second minimum, and clips never exceed your cap.
What should I do when one beat is longer than the cap?
The tool splits it at word boundaries and flags the affected clips. The better fix is a rewrite: break the long beat into two actions with a natural cut point, then re-run the split.
常见问题
单次生成做不到。做法是把戏切成不超过 8 秒的片段、在每条提示词里重复同一段一致性描述块、在支持时把每个片段的最后一帧当作下一个的第一帧喂进去,然后把结果剪到一起。这个规划器替你把那串片段搭好。
大约 24 个词的动作(按每秒约 3 个词),或者大约 16 个词的说出口的对白(按每秒约 2.5 个词再加 20% 余量用于呼吸和反应)。这就是为什么半页的一场戏需要好几个片段。
因为每次生成都从零开始,模型对你上一个片段没有记忆。解法是在每条提示词里逐字重复完整的人物和场景描述,也就是每张片段卡上那段一致性表头在做的事。
它们把片段串起来。每个片段的 EXIT STATE 描述最后一帧,下一个片段的 ENTRY STATE 把它重复一遍,模型于是从上一个片段停住的地方接着开始。如果你的模型接受起始图,就把这套做法和喂入上一个片段的最后一帧结合起来用。
你模型的上限就是天花板:Veo 3 是 8 秒,Sora 2、Kling 2.x 和 Runway Gen-4 在 2026 年底大约是 10 秒。很多电影人会有意规划更短的片段(4 到 6 秒),因为短生成飘得更少,也在剪辑里给出更多切点。











