Automate Video Production With AI Agents: A Working Setup
A working pipeline to automate video production with AI agents: script, shots, voice and assembly, with real per-stage costs and the human checkpoints that matter.
Automate Video Production With AI Agents: A Working Setup
A short product video used to mean a brief, a freelancer, two rounds of notes and a week. Today a coding agent can take a one-paragraph brief, write the script, generate every shot, record the voiceover and render an MP4 while you review something else. The pieces all exist. What most teams are missing is the wiring: which stage runs on which model, what each stage costs, and where a person still has to look before anything ships.
This is the setup we would build today. Seven steps, a cost sheet you can recompute for your own batch, two approval gates, and the platform disclosure steps that most automated pipelines forget until a video gets labelled or pulled.
The pipeline, stage by stage
Published write-ups of agent video workflows converge on the same shape. MindStudio describes five agents: script, storyboard, image or video generation, voice, and assembly. A second MindStudio write-up for content marketing orders it as script and shot list, voiceover with word-level timing, video generation, then assembly, subtitles and export. We use the same skeleton with one addition, an explicit review stage, because that is where the money is saved.
- Brief to script and shot list. An LLM turns the brief into a spoken script plus a numbered list of shots, each with a duration, a visual prompt and the line of voiceover it sits under.
- Storyboard stills. One cheap still image per shot, so you can judge composition before paying for motion.
- Gate one: human approval of script and stills.
- Shot generation. Each approved still or prompt becomes a video clip on the model that fits that shot.
- Voiceover. Text to speech from the approved script, split per shot or per paragraph.
- Assembly and captions. Clips, voice, captions and end card are stitched programmatically into the final file.
- Gate two: human review of the cut, then disclosure and publishing.
Stage 1: the brief and the shot list the agent must return
The single biggest quality lever in the whole pipeline is forcing the script stage to return a shot list in a fixed format, because every later stage reads from it. Here is the format we ask for. Paste it into your agent instructions as is.
Copy-paste: shot list contract
- shot_id: s01, s02, and so on, in order.
- duration_s: a whole number of seconds for the clip.
- vo_line: the exact words spoken over this shot, or empty for a silent beat.
- visual: one sentence describing subject, action, setting and camera move. No brand names the model cannot render.
- needs_audio: true only if the shot needs native sound (dialogue, ambient sound tied to the action).
- model_hint: fast, premium or audio, chosen by the rules in the routing section below.
- on_screen_text: the caption or overlay, kept short enough to read on a phone.
Then add three hard rules to the brief: total duration within the platform target, one idea per shot, and a final shot reserved for the call to action.
This stage is the cheapest in the pipeline. MindStudio puts the script step for a five-minute explainer at $0.05 to $0.20.
Stages 2 and 4: routing each shot to the right model
The mistake we see most often is picking one video model and sending every shot to it. Shots differ. A slow product rotation, a talking scene with ambient sound and a quick establishing shot have different requirements, and the prices spread wide. At the time of writing, per-second video prices on Aitachyon range from Hailuo 02 at $0.085/s to Seedance 2.5 at $0.44/s at 720p, with Kling v3 at $0.16/s silent or $0.32/s with native audio, Veo 3.1 Fast from $0.19/s and Wan 2.7 at $0.19/s. The full grid lives on the models page with every price per call.
Decision rules for the model_hint field
- Storyboard stills: use the cheapest image model that holds your composition. Seedream 4.0 at $0.057 per image or FLUX.2 [pro] at $0.057 to $0.086 are the defaults at the time of writing. Reserve Nano Banana ($0.13) or gpt-image-2 ($0.10 to $0.40) for stills that go straight into the final cut. Our Nano Banana vs FLUX.2 Pro comparison covers where each one holds up on product shots.
- Fast: filler, establishing shots and B-roll under the voice. Hailuo 02 is the price floor here.
- Audio: only shots where needs_audio is true. Kling v3 with native audio costs twice its silent rate, so the agent should never request audio for a shot that will sit under a voiceover anyway.
- Premium: the hero shot and the first two seconds, where the viewer decides whether to keep watching. This is where Seedance 2.5 or Veo 3.1 Fast earns its price. We compared the top two in Kling v3 vs Seedance 2.5.
One pricing note worth being honest about. Buying direct from a provider can be cheaper per unit. Google lists Veo 3.1 Fast at $0.10 per second at 720p, $0.12 at 1080p and $0.30 at 4K, with Veo 3.1 Lite at $0.05 per second at 720p. If your pipeline only ever uses one model, going direct makes sense. The case for a single account is routing: one key, one balance and one billing trail across a dozen models, which is what the rules above need.
Stage 5: voiceover, timing and what it costs
Voice is the stage where automation saves the most calendar time and the least money, because it is already cheap. ElevenLabs' own API pricing lists text to speech at $0.08 per 1,000 characters for v3 and Multilingual v2, and $0.04 for Flash and Turbo, with a discounted v4 rate of $0.022 per 1,000 characters shown until October 12. Through Aitachyon, ElevenLabs voiceover is $0.043 to $0.086 per 450 characters at the time of writing. The cost detail is in our ElevenLabs voiceover cost breakdown.
Two timing decisions to make up front
- Voice drives the edit, or the edit drives the voice. For explainers, generate the voice first and cut shots to its timing. The content marketing workflow above specifically calls for voiceover with word-level timing before video generation, so the agent can set each clip's duration from the real audio.
- One file or one file per shot. Per-shot files make re-edits cheap: change one line and you regenerate a few hundred characters instead of the whole read.
Stage 6: assembly without a timeline editor
Assembly is where most "automated" pipelines quietly fall back to a person dragging clips in an editor. Two approaches keep it in code.
FFmpeg for straight cuts
If the video is a sequence of clips under a voice track with burned-in captions, FFmpeg handles it. MindStudio notes that assembly can be done programmatically with FFmpeg.
Remotion for anything with motion design
For animated text, lower thirds, price callouts or templated end cards, Remotion is the stronger fit: the video is a React component and the agent writes it. Remotion ships official agent skills, installed with npx skills add remotion-dev/skills, covering markup, rendering, captions and more, and designed for coding assistants such as Claude Code and Cursor. The /remotion-best-practices skill is the catch-all that includes the others, so install that one first.
Whichever you choose, captions belong in this stage, generated from the script text rather than transcribed back from audio. Our note on why captions are no longer optional covers placement for vertical formats.
The cost sheet: what a real batch costs
Here is the arithmetic, using Aitachyon prices at the time of writing. Swap in your own durations and the formula holds.
Formula
Batch cost = (stills x image price) + (seconds of video x per-second price, per model bucket) + (voice characters / 450 x voice price) + reroll buffer. Assembly on FFmpeg or Remotion runs on your own machine and adds compute, not per-call fees.
Worked examples
- 20 hook variants, 5 seconds each, on Hailuo 02: 20 x 5 s x $0.085 = $8.50.
- 50 product stills: on Seedream 4.0, 50 x $0.057 = $2.85. On FLUX.2 [pro], $2.85 to $4.30. On Nano Banana, 50 x $0.13 = $6.50.
- A week of shorts (7 videos, 6 shots of 5 seconds each, 210 seconds of footage): all on Hailuo 02, 210 x $0.085 = $17.85. All on Veo 3.1 Fast from $0.19/s, $39.90. All on Kling v3 with audio, $67.20. Add 7 storyboards of 6 Seedream stills (42 x $0.057 = $2.39) and two 450-character voice blocks per video (14 x $0.043 to $0.086 = $0.60 to $1.20).
- The routed version of that week: shot one of each video on Veo 3.1 Fast (7 x 5 s x $0.19 = $6.65) and the other 175 seconds on Hailuo 02 (175 x $0.085 = $14.88). Video total $21.53, against $39.90 for everything on Veo.
Budget a reroll buffer. If you expect to regenerate half your shots once, multiply the video line by 1.5. That buffer is the argument for the storyboard gate: rerolls on a still cost cents, rerolls on video cost dollars.
For a sense of the range, MindStudio estimates a five-minute explainer at $2 to $7 in total, and $10 to $20 with premium models such as Veo and ElevenLabs premium voices. The same vendor contrasts a traditional 60-second brand spot at $2,000 to $20,000 with $5 to $50 generated. Treat both as vendor estimates. Our model-by-model API pricing guide has the rest of the grid.
Where humans stay in the loop
Every serious write-up keeps a person in the loop. MindStudio recommends a human review step for anything public-facing, and its marketing workflow specifies manual output review, automated flags for clips below duration or confidence thresholds, and a human approval gate before final assembly. We put the gates at the two points where a wrong call is most expensive.
Gate one checklist: before any video is generated
- The script says one thing, and the first line earns the second.
- Every claim in the voiceover is one you could defend to a platform reviewer.
- Each still shows the right product, with no invented logos or garbled packaging text.
- Only shots that need native audio have needs_audio set to true.
- Premium routing is limited to the hook and the hero shot.
- The projected cost from the formula above is under the batch budget.
Gate two checklist: before publishing
- Watch the full cut with sound, then once muted to check the captions carry it.
- Check hands, faces and text on screen frame by frame in the hook.
- Decide the disclosure answer (next section) and record it with the file.
- Match the end card to the landing page the video will send people to.
Automate the flags so the reviewer only looks where it matters: a clip shorter than its duration_s, a failed render, a voice file longer than its shot.
Disclosure belongs inside the pipeline
Platforms now expect you to say when realistic footage was generated, and an automated pipeline should record that answer per video.
- YouTube requires disclosure for realistic content that makes a real person appear to say or do something they didn't, alters footage of a real event or place, or generates a realistic scene that didn't occur. Clearly unrealistic content and production help such as scriptwriting, captions or color correction do not need it. The answer is set in YouTube Studio at upload. Creators who consistently fail to disclose can get a label applied for them, content removed or suspension from the YouTube Partner Program, and YouTube says proper disclosure does not limit reach or monetization eligibility.
- TikTok requires creators to label realistic AI-generated content and was the first video platform to implement C2PA Content Credentials, reading them from May 9, 2024 to auto-label content made on other platforms. TechCrunch reported that unlabeled AI content may be taken down, and that C2PA is a coalition led by Adobe with Google, Meta, OpenAI, Microsoft and Intel among its members.
The practical rule: add a disclose field to each video record, default it to true for any realistic generated footage, and have the publishing step refuse to run while it is empty. For ad accounts, the policy traps in getting video ads approved on Meta and TikTok apply on top.
Running it from Claude Code or Cursor
The whole pipeline can run inside one agent session. The agent reads the brief, writes the shot list to a file, calls the image and video models, polls until each render is done, writes the FFmpeg command or the Remotion composition and renders. What makes it hold together is that each generated file has a stable reference the agent can come back to, and that every call reports what it cost. On Aitachyon, each output gets a ref such as img_, scn_ or vo_ that code can fetch with GET /api/generations/{ref}, and every job is itemised with its model and price, so the agent can write an actual cost next to each shot instead of an estimate. Connecting Claude Code to the hosted MCP server is one command: claude mcp add --transport http aitachyon https://aitachyon.com/api/mcp --header "Authorization: Bearer ait_...". The same key works against the plain HTTP API if you would rather run the pipeline as a script.
A brief you can hand the agent
Make a 30-second vertical video for [product]. Return the shot list in the contract format first and stop for my approval. After approval, generate one still per shot on the cheapest image model, then stop again. After my second approval, generate video per the routing rules, generate the voiceover per shot, assemble with Remotion with captions from the script, and write a cost table listing each ref, model and price. Set disclose to true.
The two explicit stops are the human gates. Without them, the agent will happily spend the whole budget on the first version of the script. For the setup steps, see generating AI video from Claude Code with an MCP server.
FAQ
Can AI fully automate video production end to end?
Technically yes: script, shots, voice and assembly can all run from one agent. In practice every published workflow we found keeps a human review before anything goes public, and the platform disclosure rules make a person accountable for the final file anyway. Let the agent do the production and keep the final judgment with a person.
How much does it cost to automate a short video with AI?
It depends on seconds of footage and which model renders them. At the time of writing on Aitachyon, 30 seconds on Hailuo 02 is $2.55, on Veo 3.1 Fast from $5.70 and on Kling v3 with audio $9.60, plus cents for stills and voice. MindStudio's estimate for a five-minute explainer is $2 to $7, or $10 to $20 with premium models.
How long does an automated AI video take to produce?
MindStudio reports 5 to 15 minutes from prompt to export for a 60-second social clip and 20 to 40 minutes for a 2 to 3 minute explainer, and 30 to 60 minutes for a 3 to 5 minute video built from generated clips.
Should I use FFmpeg or Remotion to assemble AI video?
FFmpeg for straight cuts under a voice track with burned captions. Remotion when you need animated text, templates or anything you would otherwise build in a motion design tool, since the agent can write the React composition and render it to MP4.
Do I have to label AI-generated videos on YouTube and TikTok?
For realistic generated footage, yes on both. YouTube asks at upload in YouTube Studio, and TikTok requires a label and can detect C2PA Content Credentials on its own. Fantastical or clearly unrealistic content, and AI used only for scripts or captions, does not need a YouTube disclosure.
Sources
- Google AI for Developers: Gemini Developer API pricing (Veo 3.1), September 2026
- ElevenLabs: API Pricing, September 2026
- YouTube Help: Disclosing use of altered or synthetic content
- TikTok Newsroom: Partnering with our industry to advance AI transparency and literacy, May 2024
- TechCrunch: TikTok will automatically label AI-generated content created on other platforms, May 2024
- Remotion: Agent Skills
- MindStudio: AI Agent Workflow, YouTube Video From One Prompt
- MindStudio: AI Video Generation for Content Marketing, Multi-Agent Workflow
If you want to run this pipeline without five provider accounts, Aitachyon puts the image, video and voice models above behind one prepaid balance and one key, usable from the web studio, the HTTP API or a hosted MCP server that Claude Code connects to in one command. Failed renders are refunded automatically, each key has its own spend alert, and the API and MCP quick start takes a few minutes.
Related articles
AI Credits Pricing vs Pay-Per-Call: Which Actually Costs Less
How AI credits pricing really works, where expiry, re-rolls and failed renders hide cost, and a worked example comparing credits with per-second dollar pricing.
StrategiesAd Landing Page Congruence: The Conversion Leak Nobody Audits
A strong ad still converts cold when the page breaks its promise. Why ad landing page congruence leaks money and a checklist to align it end to end.
StrategiesThe Video Ad Metrics That Actually Predict Winners
Which video ad metrics forecast a winner at low spend: hook rate, hold rate, cost per outbound click. With cited 2026 benchmarks and a 2x2.
StrategiesRetargeting Ad Creative: What to Show the People Who Bounced
Retargeting ad creative needs a different message than prospecting. The objection, proof, and urgency shifts that recover lost carts and cold leads.
StrategiesBrand Consistency in Ads: A Brand-Kit System for 50 Variants
A brand kit for ad creative—locked colors, type, voice, and do-not rules—that holds brand consistency in ads across 50 variants without per-asset review.
StrategiesThe SaaS demo ad: a format most founders get wrong
Most SaaS demo ads are 60 seconds of UI tour over elevator music. The structure that converts: face hooks, screens convince, face closes. Script included.
Free tools to try
Free AI product photo generator
Generate clean, studio-style product photos for your store and listings in seconds. Crisp lighting and tidy backgrounds, free to try with no account, keep your first shot on signup.
Try it freeFree toolFree AI image generator
Describe what you want and get a high-quality AI image in seconds. A free AI image generator, no account needed to preview, keep your first image when you sign up.
Try it freeFree toolFree background remover
Remove the background from any image in seconds and get a clean, transparent cutout. A free background remover, no account needed to preview, keep your first cutout when you sign up.
Try it freeStop describing your brand. Paste your URL.
Aitachyon reads your whole brand from your website, then creates videos, images, carousels, posts and banners, on-brand, every format, every feed.