---
title: "Script to Video AI: Convert a Script Into Video, Fast"
description: "Convert a script to AI video scene by scene: format beats, write per-scene prompts, pick Kling, Veo or Seedance per shot, add voice and captions, iterate."
canonical: "https://aicontentdrop.com/blog/script-to-video-guide"
source: "https://aicontentdrop.com/blog/script-to-video-guide"
---
Guide

March 28, 2026

19

min read

# Script to Video: Convert Any Script Into AI Video, Scene by Scene

Convert a script to AI video scene by scene: format beats, write per-scene prompts, pick Kling, Veo or Seedance per shot, add voice and captions, iterate.

guide

script-to-video

workflow

text-to-video

You have a script. It might be a 30-second product ad, a voiceover for a landing page, a founder story, or the first minute of a YouTube video. What you do not have is footage, a crew, or a week to get either. Script to video is the workflow that closes that gap: you convert the written script into a sequence of AI-generated scenes, each rendered by the video model that fits that scene, then assemble the pieces into one finished video. This guide is the full, practical version of that workflow on AI Content Drop, with the formatting rules that keep a script convertible, a worked example that turns a 30-second product script into five scene prompts, the model to use for each scene type with its credit cost, and the failure modes you will hit on your first attempt.

## What script to video means on this platform

On AI Content Drop, script to video is chat-driven rather than form-driven. You describe a scene in plain language, the platform recommends the video model that suits it, writes the generation prompt, shows you the model, credit cost and rough render time before you commit, then generates the clip. Credits are deducted only on successful generation; a render that fails or times out is not charged, and there is nothing to refund because nothing was taken. The platform runs the models directly, so the same credit balance covers Google Veo, Kling, Seedance, Hailuo, MiniMax and the rest of the catalogue, and a chat turn costs between 1 and 3 credits depending on what you ask it to do, separate from the video render itself.

There are two doors into the same engine. The [Chat-to-Ads Studio](https://aicontentdrop.com/chat) is the conversational one: paste a script, talk through the scenes, accept or edit the recommended model, and generate from the chat. Its Video Ads mode is tuned for exactly this, and its UGC Creator mode hands talking-head beats to the avatar pipeline. The [video generator](https://aicontentdrop.com/generate/video) is the direct one: a composer where you pick the model, aspect ratio and duration yourself, attach reference images, and watch each scene land in a generations feed. Most people draft in chat and finish in the generator, because the generator's Remix button reloads any finished scene with its exact prompt, model, aspect ratio and duration so you can change one line and re-run it.

The important mental shift is that script to video is not one prompt. A video model renders one shot at a time, from 3 to 30 seconds depending on the model, and a 30-second ad is four to six of those shots. The script is the plan for the sequence; each scene prompt is the instruction for one shot. Get the plan right and the prompts almost write themselves.

## Format the script so it converts cleanly

A script written for a human editor is a paragraph of voiceover with some stage directions. A script written for AI generation is a table of beats. Reformatting takes ten minutes and saves you a dozen wasted renders, so do it before you open the generator.

### Write in scene beats, not paragraphs

Split the script wherever the picture would change. Each beat needs one subject doing one action in one place. If a sentence of voiceover covers two locations, it is two beats. If a beat has no picture you can name in a single sentence, it is not ready; conceptual lines like "we believe in craft" need a visual stand-in, such as hands finishing a wooden edge, before a model can render them.

### Separate on-screen text from voiceover

Keep three columns: what is spoken, what appears as text on screen, and what the camera sees. This matters because video models render legible text unreliably, so any words the viewer must read should be added in your editor after generation, not requested in the prompt. The voiceover column drives your audio choice later (a native-audio model, an avatar with lip-sync, or a separate voice track), and the visual column becomes the prompt. Mixing the three is the single most common reason script-to-video output looks wrong.

### Give every beat a duration the model can hit

Duration is not free-form. Each model has a range, and a beat that falls outside it either gets clamped or split. The ranges that matter for a script workflow, taken from the model catalogue:

- Veo 3.1 and Veo-3 Fast
  
  render fixed 8-second clips. Write the beat for 8 seconds or plan to trim.
- Kling 3.0
  
  takes 3 to 15 seconds, so it covers the short hook and the long demo beat alike.
- Seedance 2.0 and Seedance 2.0 Fast
  
  take 4 to 15 seconds.
- Seedance 2.5
  
  holds a single take up to 30 seconds.
- Hailuo 2.3
  
  is 5 to 6 seconds;
  
  Kling 2.5 Turbo Pro
  
  is 5 to 10.

A good default is 4 to 8 seconds per beat. Shorter beats are easier to control and cheaper to regenerate; longer beats drift. If the voiceover for a beat runs longer than its picture, cut the picture into two shots rather than stretching one.

### Pick the aspect ratio per platform before you generate

Generate at the ratio you will publish, because cropping a 16:9 render down to 9:16 throws away two thirds of the frame and usually the subject with it. Use 9:16 for TikTok, Reels and Shorts, 16:9 for YouTube and landing pages, and 1:1 for feed placements. Kling, Veo and Hailuo support 16:9, 9:16 and 1:1; the Seedance 2.0 family adds 4:3, 3:4 and 21:9 if you need a cinematic letterbox. The generator's aspect control only shows the ratios the selected model supports, so if an option is missing, the model cannot render it. If you are producing one script for two platforms, write the beats once and render each beat twice at the two ratios; do not render once and crop.

For image-first ads, where a product photo is the anchor and the script describes how it moves, the formatting rules are slightly different. The [Kling image-ad script workflow](https://aicontentdrop.com/blog/how-to-write-kling-2-6-image-ad-script-workflow) walks through that version from blank page to published ad.

## Worked example: a 30-second product script in five beats

Here is a 30-second script for a glass cold-brew pitcher, already reformatted into beats. It is the kind of ad an ecommerce brand runs on Reels and TikTok, so the target ratio is 9:16.

| Beat | Seconds | Voiceover | On-screen text | Visual |
| --- | --- | --- | --- | --- |
| 1. Hook | 0 to 4 | "Your coffee maker has a 6 a.m. problem." | none | Dark kitchen, drip machine gurgling, a hand fumbles for the switch |
| 2. Problem | 4 to 10 | "Hot, rushed, bitter. Every single morning." | Bitter. Every. Morning. | Steam off a mug, a wince after the first sip |
| 3. Solution | 10 to 17 | "Cold brew overnight. Twelve hours, zero effort." | none | Grounds into the glass pitcher, water poured, fridge door closes |
| 4. Proof | 17 to 24 | "Smooth, low-acid, ready when you are." | Ready when you are | Morning light, pitcher out of the fridge, slow pour over ice |
| 5. CTA | 24 to 30 | "Brew tonight. Order today." | Order today. Link below. | Pitcher hero shot on a clean counter, logo held for the last second |

Notice what the table already decided. The on-screen text lives in its own column and will be added in the editor, so none of it goes into a prompt. Every beat has one subject and one action. Beats 3, 4 and 5 all show the same product, which means they should be generated from the same product photo, not from text alone, or the pitcher will change shape between shots.

## Breaking the script into per-scene prompts

A scene prompt is the visual column expanded into everything the model needs to render it: subject, action, setting, camera, lighting, style, audio, and what to avoid. Writing it as a structured block instead of a paragraph keeps you from forgetting a field, and the structure pastes straight into the chat, which turns it into the final generation prompt for whichever model you pick. Here is beat 1:

```
{
  "scene": 1,
  "beat": "hook",
  "duration_seconds": 4,
  "aspect_ratio": "9:16",
  "subject": "a countertop drip coffee maker with a blinking green light",
  "action": "gurgles and sputters while a sleepy hand reaches into frame and fumbles for the switch",
  "setting": "small apartment kitchen before dawn, only the machine's light and a blue window glow",
  "camera": "static close-up at counter height, shallow depth of field, slight handheld drift",
  "lighting": "low-key, single practical source from the machine, cool ambient fill from the window",
  "style": "photoreal, muted teal and amber, light film grain",
  "audio": "gurgling machine, faint kitchen hum, no music",
  "on_screen_text": "none",
  "avoid": ["text", "logos", "extra hands", "bright overhead light"]
}
```

And beat 4, which is the product beat and therefore anchored to a reference photo rather than described from scratch:

```
{
  "scene": 4,
  "beat": "proof",
  "duration_seconds": 7,
  "aspect_ratio": "9:16",
  "reference_image": "product-photo-pitcher-front.jpg",
  "subject": "the glass cold-brew pitcher from the reference image, three-quarters full of dark coffee",
  "action": "lifted out of a fridge, then a slow pour over a tall glass of ice, coffee swirling as it hits",
  "setting": "bright kitchen counter, morning window light from the left",
  "camera": "medium close-up, slow push-in, the pour in slow motion",
  "lighting": "soft daylight, warm highlights on the glass, no harsh reflections",
  "style": "photoreal, clean commercial look, matches scene 3 palette",
  "audio": "ice clink, liquid pour, quiet room tone",
  "on_screen_text": "none",
  "avoid": ["changing the pitcher shape", "text", "extra products", "hands covering the label"]
}
```

The remaining three beats follow the same shape. Beat 2 is a close-up of steam and a face reacting, a 6-second shot with no product in it, so it is the cheapest to get right. Beat 3 is the assembly sequence: grounds, water, fridge. Fit it in 7 seconds by choosing one continuous action, the pour and the door closing, rather than all three steps, and use the same reference photo as beat 4. Beat 5 is a static hero shot on a counter with a slow orbit; it needs 6 seconds, the same photo, and a completely clean frame so the logo and the call to action can be laid over it in the editor.

Three rules keep a set of scene prompts coherent. First, reuse the style and lighting lines verbatim across every beat; the model has no memory between generations, so continuity comes from you repeating yourself. Second, name the camera move in every prompt, because an unspecified camera gives you the same slow zoom on every clip. Third, keep the avoid list boring and consistent: text, logos, extra hands, extra products. If a hook needs to be stronger than what you have, the [25 UGC hook formulas](https://aicontentdrop.com/blog/winning-ai-ugc-hooks-for-video-ads) are written to drop into beat 1 without changing the rest of the script.

## Choosing the model for each scene type

The reason to route scene by scene rather than pick one model for the whole ad is that scene types reward different strengths. A talking head needs lip-sync. A product shot needs a reference image and shape fidelity. A cinematic beat needs motion quality and, ideally, native sound. A motion-control beat needs a reference video. The chat will suggest a model per scene; this is the reasoning behind the suggestion so you can override it. Credit costs are flat per generation and come from the platform's published rates.

| Scene type | Model | Credits | Why |
| --- | --- | --- | --- |
| Talking head, UGC-style | UGC Factory | 22 | Your avatar photo plus your script, AI voice and lip-sync in one pass |
| Talking head with dialogue in a scene | Seedance 2.0 | 56 | Native audio with lip-sync, multi-shot consistency, 1080p |
| Product b-roll from a photo | Kling 3.0 | 22 | Image-to-video with strong motion, 3 to 15 seconds, 1080p |
| Product b-roll, budget | Seedance 2.0 Fast | 22 | 720p, native audio, start and end frame control, up to 5 reference images |
| Cinematic beat with sound | Veo 3.1 | 26 | Synced audio from the prompt, 8-second clips, high realism |
| Cinematic long take | Seedance 2.5 | 30 for 5 s at 720p | Single takes to 30 seconds, billed per second, can extend a clip |
| Motion-controlled performance | Kling 3.0 Motion Control | 22 | Reference video drives the motion of your subject |
| Drafts and storyboards | Kling 2.5 Turbo Pro | 11 | Fastest lane; test the prompt before spending on the final |
| Cheap social b-roll | Hailuo 2.3 | 17 | 5 to 6 second clips with usable physics |

### Talking-head and UGC beats

If the beat is a person speaking to camera, do not render a silent clip and lay a voice over it; the mouth will not match. Use a lane that generates the voice and the lip-sync together. The [UGC Factory](https://aicontentdrop.com/generate/ugc) is the direct route: upload an avatar photo, pick a template, paste the beat's voiceover line as the script, and it produces the spoken, lip-synced clip for 22 credits. Seedance 2.0 is the alternative when the speaker needs to be inside a wider scene, because it renders picture, voice and lip-sync in the same pass at 56 credits, and Seedance 2.0 Fast keeps the native audio at 22 credits and 720p when the beat is not the hero shot.

### Product b-roll

Anything that shows the product should be image-to-video from a real product photo. Kling 3.0 at 22 credits is the workhorse here: attach the photo as a reference, describe the motion, and the pitcher keeps its shape across beats 3, 4 and 5 because every beat started from the same pixels. Seedance 2.0 Fast matches the price, drops to 720p, and adds start-and-end frame control, which is the tool to reach for when a beat has to begin on one product angle and end on another.

### Cinematic and brand beats

For atmosphere shots with no product and no dialogue, Veo 3.1 at 26 credits gives synced sound from the prompt and the most convincing realism per credit, at the cost of a fixed 8-second length; Veo-3 Fast at 14 credits is the same family for drafting, text-to-video only. Seedance 2.5 is the model for a long unbroken take: it holds a single shot up to 30 seconds, is billed at 6 credits per second at 720p or 3 at 480p, and can extend a clip you already have. That makes a whole 30-second script renderable as one take for 180 credits at 720p, which is tempting until the eighteenth second is wrong and the fix costs another 180. Scene-by-scene is cheaper to iterate. If you are writing for YouTube specifically, the [Seedance 2.0 Pro YouTube ad templates](https://aicontentdrop.com/blog/winner-youtube-video-ad-script-templates-seedance-2-0-pro) are already broken into beats at 16:9.

### Motion control

When the script calls for a specific performance, a hand gesture, a product turn, a dance step, describe it once on camera with your phone and hand the clip to Kling 3.0 Motion Control at 22 credits. The output follows the reference video's motion with your subject in it, which is far more reliable than describing choreography in words.

### A note on Sora 2

Sora 2 at 45 credits, Sora 2 Pro at 56 and Sora 2 Pro Storyboard at 84 still run on the platform today. OpenAI discontinued the Sora app and website on April 26, 2026, and has set September 24, 2026 as the date the Sora 2 models and the Videos API are removed, with no announced successor. That removal happens upstream of every platform, so do not build a script workflow on Sora now. Veo 3.1 is the like-for-like swap for audio-native clips and Seedance 2.0 replaces the Storyboard lane for multi-shot renders; the [Sora 2 pricing and shutdown guide](https://aicontentdrop.com/blog/sora-2-pricing-api-shutdown) has the dates and sources.

### What the example costs

Routing the cold-brew script through the table above gives beat 1 on Kling 3.0 (22), beat 2 on Hailuo 2.3 (17), and beats 3, 4 and 5 on Kling 3.0 from the product photo (22 each), for 105 credits of finals. A draft pass of all five beats on Kling 2.5 Turbo Pro first is 55 credits, so one full draft-then-final cycle is 160 credits. On the [current plans](https://aicontentdrop.com/pricing) that is one Starter month at 150 credits with a top-up, or about a third of Professional's 450, which leaves room for two more rounds of revisions. Swapping beat 3 for a UGC Factory talking head costs the same 22. If you route everything through Seedance 2.0 instead, the per-video math changes substantially; the [Seedance 2.0 pricing guide](https://aicontentdrop.com/blog/seedance-2-0-pricing-2026) works that through per plan. The free tier's 10 credits sit below every lane in the table, so treat it as a place to test the chat flow and image models, and expect to be on a paid plan for a real script.

## Voice, lip-sync and captions

Audio is where a script-to-video project either feels finished or feels like a slideshow, and the voiceover column of your script tells you which route to take for each beat.

### Spoken beats: avatar and lip-sync

The UGC Factory turns one photo into a speaking presenter. Its rail runs Template, Image, Action, Script, Voice, Ambience: choose from General, Selfie, Selling, Testimonial, Tutorial or Storytelling framing, upload the avatar photo, pick the on-camera action, paste the script, choose a voice type and emotion, and add an ambience bed. The script step shows a running estimate of spoken length and caps it at about 15 seconds of speech, so a talking-head beat longer than that becomes two clips with a cut between them, which is how real UGC is edited anyway. It requires a paid plan and costs 22 credits per clip.

### Ambient beats: native audio in the render

For beats with no speech, the cheapest audio is the audio the model generates with the picture. Veo 3.1 and the Seedance 2.0 family produce synced sound from the prompt, which is why the JSON blocks above carry an audio field; write the sound you want the same way you write the camera. MiniMax H3 is the budget option here: it renders stereo dialogue, foley and ambience in the same pass as the picture, at 2 credits per second at 768p, and any finished clip can be upgraded to 2K afterwards. If you would rather control the voice yourself, render the beats silent, record or generate the voiceover as one continuous track, and lay it under the cut; that is the right choice for a narrated explainer where no one is on camera.

### Captions

Do not ask a video model for captions. The [Marketing Studio](https://aicontentdrop.com/marketing-studio) produces UGC-style ads from a single brief with kinetic captions burned in, bold sans-serif with a black outline that reads on any background, and that is the fastest route when the whole script is one presenter. For a scene-by-scene build, add captions in your editor from the voiceover column. Write the spoken lines short enough that any two-second stretch fits on two lines of caption; the 30-second script above was written that way on purpose, with no sentence longer than nine words.

## Assembling and iterating

Generate the beats in script order, but review them as a set before rendering finals. The generations feed in the video generator shows every clip with its model, aspect ratio and duration, and opening a clip gives you Remix, download, and, on MiniMax H3 clips, a 2K upscale. Remix is the iteration loop: it reloads the composer with that scene's prompt, model, aspect and duration, you change the one line that was wrong, and you re-run only that beat. Never regenerate the whole ad because one scene missed; four good clips are already paid for and sitting in the feed.

The order of operations that wastes the fewest credits is: draft every beat once on the cheapest lane that fits, fix prompts until each draft is directionally right, then render finals on the models from the table. Drafting on Kling 2.5 Turbo Pro at 11 credits and finishing on Kling 3.0 at 22 means a wrong idea costs 11, not 22, and most prompt mistakes show up in the first second of a draft. If a render fails or times out, the feed says so and you were not charged; retry it as-is before rewriting anything, because a failed job is not feedback on your prompt.

Assembly happens in whatever editor you already use. Export each clip, place them in beat order, trim the fixed-length clips (an 8-second Veo shot for a 6-second CTA beat is trimmed, not regenerated), lay the on-screen text from the script table over the clean frames, add the voiceover track or keep the native audio, and cut to the voiceover's rhythm rather than to the clip boundaries. A tight edit hides more seams than a better model does.

## Common failure modes and fixes

- Garbled text in the frame.
  
  You asked for on-screen text in the prompt. Remove it, add "text" to the avoid list, and put the words back in the editor.
- The product changes shape between beats.
  
  The product beats were generated from text. Regenerate them as image-to-video from one product photo, and repeat the exact style and lighting lines in every prompt.
- The speaker's lips do not match.
  
  A silent clip had a voice laid over it. Re-render the beat on a lip-sync lane: UGC Factory for a presenter, Seedance 2.0 for a speaker inside a scene.
- The beat is longer than the model allows.
  
  Veo is 8 seconds flat, Hailuo 2.3 is 6 at most. Either split the beat into two shots or move it to Kling 3.0 or Seedance, which run to 15.
- Every clip has the same slow zoom.
  
  No camera move was specified. Name one per beat: static, push-in, orbit, handheld, tracking.
- The vertical version is cropped badly.
  
  It was rendered at 16:9 and cropped. Render it again at 9:16; the generator only offers the ratios the model supports, so pick a model that lists it.
- It all looks like AI video.
  
  Centered framing, flat lighting, no texture. The
  
  six fixes for generic script-to-video output
  
  are prompt-level changes that apply to every model in the table.
- The render failed.
  
  Nothing was charged. Retry once unchanged; if it fails again, shorten the prompt and remove any reference image over the model's cap.

## Frequently asked questions

### Can I paste a whole script into the chat and get a finished video?

You can paste the whole script, and the chat will help you break it into beats, recommend a model for each and generate them. The finished video is still assembled from those clips in an editor, because the on-screen text, the trims and the final voice mix belong there. Treat the chat as the director and the editor as the finishing room.

### How long can each scene be?

It depends on the model: 8 seconds fixed on Veo 3.1 and Veo-3 Fast, 3 to 15 on Kling 3.0, 4 to 15 on the Seedance 2.0 family, up to 30 in one take on Seedance 2.5, and 5 to 6 on Hailuo 2.3. Four to eight seconds per beat is the practical default.

### Do I need a different model for every scene?

No. Most ads use two: one lane for product or b-roll beats and one for the speaking beat. The table exists so you can pick deliberately, not so you use nine models.

### How do I keep the same product or person across scenes?

Generate every scene that shows them from the same reference image, and copy the style and lighting lines between prompts word for word. For a recurring presenter, the UGC Factory uses the same uploaded photo for every clip, which is the simplest continuity there is.

### Am I charged for a scene that fails?

No. Credits are deducted only on successful generation. A failed or timed-out render costs nothing and can be retried.

### Should I use Sora 2 for this?

Not for anything you will still be producing in October. The Sora 2 models and the Videos API are removed by OpenAI on September 24, 2026, with no replacement announced. Use Veo 3.1 or the Seedance 2.0 family instead.

## Where to start

Reformat your script into the five-column table, decide the aspect ratio, and open the [Chat-to-Ads Studio](https://aicontentdrop.com/chat) with beat 1. Accept the recommended model or override it from the table above, draft, remix, and move to beat 2. When the drafts read as a sequence, render finals in the [video generator](https://aicontentdrop.com/generate/video) and assemble. [Start with 10 free credits](https://aicontentdrop.com/register) to run the flow, then pick a plan sized to the number of scripts you ship a month.