Skip to content
StrategiesSeptember 30, 2026· 10 min read

Automate Video Production With AI Agents: A Working Setup

A working pipeline to automate video production with AI agents: script, shots, voice and assembly, with real per-stage costs and the human checkpoints that matter.

automationai agentsvideo productionmcp
Strategies

Automate Video Production With AI Agents: A Working Setup

A short product video used to mean a brief, a freelancer, two rounds of notes and a week. Today a coding agent can take a one-paragraph brief, write the script, generate every shot, record the voiceover and render an MP4 while you review something else. The pieces all exist. What most teams are missing is the wiring: which stage runs on which model, what each stage costs, and where a person still has to look before anything ships.

This is the setup we would build today. Seven steps, a cost sheet you can recompute for your own batch, two approval gates, and the platform disclosure steps that most automated pipelines forget until a video gets labelled or pulled.

The pipeline, stage by stage

Published write-ups of agent video workflows converge on the same shape. MindStudio describes five agents: script, storyboard, image or video generation, voice, and assembly. A second MindStudio write-up for content marketing orders it as script and shot list, voiceover with word-level timing, video generation, then assembly, subtitles and export. We use the same skeleton with one addition, an explicit review stage, because that is where the money is saved.

  1. Brief to script and shot list. An LLM turns the brief into a spoken script plus a numbered list of shots, each with a duration, a visual prompt and the line of voiceover it sits under.
  2. Storyboard stills. One cheap still image per shot, so you can judge composition before paying for motion.
  3. Gate one: human approval of script and stills.
  4. Shot generation. Each approved still or prompt becomes a video clip on the model that fits that shot.
  5. Voiceover. Text to speech from the approved script, split per shot or per paragraph.
  6. Assembly and captions. Clips, voice, captions and end card are stitched programmatically into the final file.
  7. Gate two: human review of the cut, then disclosure and publishing.

Stage 1: the brief and the shot list the agent must return

The single biggest quality lever in the whole pipeline is forcing the script stage to return a shot list in a fixed format, because every later stage reads from it. Here is the format we ask for. Paste it into your agent instructions as is.

Copy-paste: shot list contract

  1. shot_id: s01, s02, and so on, in order.
  2. duration_s: a whole number of seconds for the clip.
  3. vo_line: the exact words spoken over this shot, or empty for a silent beat.
  4. visual: one sentence describing subject, action, setting and camera move. No brand names the model cannot render.
  5. needs_audio: true only if the shot needs native sound (dialogue, ambient sound tied to the action).
  6. model_hint: fast, premium or audio, chosen by the rules in the routing section below.
  7. on_screen_text: the caption or overlay, kept short enough to read on a phone.

Then add three hard rules to the brief: total duration within the platform target, one idea per shot, and a final shot reserved for the call to action.

This stage is the cheapest in the pipeline. MindStudio puts the script step for a five-minute explainer at $0.05 to $0.20.

Stages 2 and 4: routing each shot to the right model

The mistake we see most often is picking one video model and sending every shot to it. Shots differ. A slow product rotation, a talking scene with ambient sound and a quick establishing shot have different requirements, and the prices spread wide. At the time of writing, per-second video prices on Aitachyon range from Hailuo 02 at $0.085/s to Seedance 2.5 at $0.44/s at 720p, with Kling v3 at $0.16/s silent or $0.32/s with native audio, Veo 3.1 Fast from $0.19/s and Wan 2.7 at $0.19/s. The full grid lives on the models page with every price per call.

Decision rules for the model_hint field

  • Storyboard stills: use the cheapest image model that holds your composition. Seedream 4.0 at $0.057 per image or FLUX.2 [pro] at $0.057 to $0.086 are the defaults at the time of writing. Reserve Nano Banana ($0.13) or gpt-image-2 ($0.10 to $0.40) for stills that go straight into the final cut. Our Nano Banana vs FLUX.2 Pro comparison covers where each one holds up on product shots.
  • Fast: filler, establishing shots and B-roll under the voice. Hailuo 02 is the price floor here.
  • Audio: only shots where needs_audio is true. Kling v3 with native audio costs twice its silent rate, so the agent should never request audio for a shot that will sit under a voiceover anyway.
  • Premium: the hero shot and the first two seconds, where the viewer decides whether to keep watching. This is where Seedance 2.5 or Veo 3.1 Fast earns its price. We compared the top two in Kling v3 vs Seedance 2.5.

One pricing note worth being honest about. Buying direct from a provider can be cheaper per unit. Google lists Veo 3.1 Fast at $0.10 per second at 720p, $0.12 at 1080p and $0.30 at 4K, with Veo 3.1 Lite at $0.05 per second at 720p. If your pipeline only ever uses one model, going direct makes sense. The case for a single account is routing: one key, one balance and one billing trail across a dozen models, which is what the rules above need.

Stage 5: voiceover, timing and what it costs

Voice is the stage where automation saves the most calendar time and the least money, because it is already cheap. ElevenLabs' own API pricing lists text to speech at $0.08 per 1,000 characters for v3 and Multilingual v2, and $0.04 for Flash and Turbo, with a discounted v4 rate of $0.022 per 1,000 characters shown until October 12. Through Aitachyon, ElevenLabs voiceover is $0.043 to $0.086 per 450 characters at the time of writing. The cost detail is in our ElevenLabs voiceover cost breakdown.

Two timing decisions to make up front

  • Voice drives the edit, or the edit drives the voice. For explainers, generate the voice first and cut shots to its timing. The content marketing workflow above specifically calls for voiceover with word-level timing before video generation, so the agent can set each clip's duration from the real audio.
  • One file or one file per shot. Per-shot files make re-edits cheap: change one line and you regenerate a few hundred characters instead of the whole read.

Stage 6: assembly without a timeline editor

Assembly is where most "automated" pipelines quietly fall back to a person dragging clips in an editor. Two approaches keep it in code.

FFmpeg for straight cuts

If the video is a sequence of clips under a voice track with burned-in captions, FFmpeg handles it. MindStudio notes that assembly can be done programmatically with FFmpeg.

Remotion for anything with motion design

For animated text, lower thirds, price callouts or templated end cards, Remotion is the stronger fit: the video is a React component and the agent writes it. Remotion ships official agent skills, installed with npx skills add remotion-dev/skills, covering markup, rendering, captions and more, and designed for coding assistants such as Claude Code and Cursor. The /remotion-best-practices skill is the catch-all that includes the others, so install that one first.

Whichever you choose, captions belong in this stage, generated from the script text rather than transcribed back from audio. Our note on why captions are no longer optional covers placement for vertical formats.

The cost sheet: what a real batch costs

Here is the arithmetic, using Aitachyon prices at the time of writing. Swap in your own durations and the formula holds.

Formula

Batch cost = (stills x image price) + (seconds of video x per-second price, per model bucket) + (voice characters / 450 x voice price) + reroll buffer. Assembly on FFmpeg or Remotion runs on your own machine and adds compute, not per-call fees.

Worked examples

  • 20 hook variants, 5 seconds each, on Hailuo 02: 20 x 5 s x $0.085 = $8.50.
  • 50 product stills: on Seedream 4.0, 50 x $0.057 = $2.85. On FLUX.2 [pro], $2.85 to $4.30. On Nano Banana, 50 x $0.13 = $6.50.
  • A week of shorts (7 videos, 6 shots of 5 seconds each, 210 seconds of footage): all on Hailuo 02, 210 x $0.085 = $17.85. All on Veo 3.1 Fast from $0.19/s, $39.90. All on Kling v3 with audio, $67.20. Add 7 storyboards of 6 Seedream stills (42 x $0.057 = $2.39) and two 450-character voice blocks per video (14 x $0.043 to $0.086 = $0.60 to $1.20).
  • The routed version of that week: shot one of each video on Veo 3.1 Fast (7 x 5 s x $0.19 = $6.65) and the other 175 seconds on Hailuo 02 (175 x $0.085 = $14.88). Video total $21.53, against $39.90 for everything on Veo.

Budget a reroll buffer. If you expect to regenerate half your shots once, multiply the video line by 1.5. That buffer is the argument for the storyboard gate: rerolls on a still cost cents, rerolls on video cost dollars.

For a sense of the range, MindStudio estimates a five-minute explainer at $2 to $7 in total, and $10 to $20 with premium models such as Veo and ElevenLabs premium voices. The same vendor contrasts a traditional 60-second brand spot at $2,000 to $20,000 with $5 to $50 generated. Treat both as vendor estimates. Our model-by-model API pricing guide has the rest of the grid.

Where humans stay in the loop

Every serious write-up keeps a person in the loop. MindStudio recommends a human review step for anything public-facing, and its marketing workflow specifies manual output review, automated flags for clips below duration or confidence thresholds, and a human approval gate before final assembly. We put the gates at the two points where a wrong call is most expensive.

Gate one checklist: before any video is generated

  1. The script says one thing, and the first line earns the second.
  2. Every claim in the voiceover is one you could defend to a platform reviewer.
  3. Each still shows the right product, with no invented logos or garbled packaging text.
  4. Only shots that need native audio have needs_audio set to true.
  5. Premium routing is limited to the hook and the hero shot.
  6. The projected cost from the formula above is under the batch budget.

Gate two checklist: before publishing

  1. Watch the full cut with sound, then once muted to check the captions carry it.
  2. Check hands, faces and text on screen frame by frame in the hook.
  3. Decide the disclosure answer (next section) and record it with the file.
  4. Match the end card to the landing page the video will send people to.

Automate the flags so the reviewer only looks where it matters: a clip shorter than its duration_s, a failed render, a voice file longer than its shot.

Disclosure belongs inside the pipeline

Platforms now expect you to say when realistic footage was generated, and an automated pipeline should record that answer per video.

The practical rule: add a disclose field to each video record, default it to true for any realistic generated footage, and have the publishing step refuse to run while it is empty. For ad accounts, the policy traps in getting video ads approved on Meta and TikTok apply on top.

Running it from Claude Code or Cursor

The whole pipeline can run inside one agent session. The agent reads the brief, writes the shot list to a file, calls the image and video models, polls until each render is done, writes the FFmpeg command or the Remotion composition and renders. What makes it hold together is that each generated file has a stable reference the agent can come back to, and that every call reports what it cost. On Aitachyon, each output gets a ref such as img_, scn_ or vo_ that code can fetch with GET /api/generations/{ref}, and every job is itemised with its model and price, so the agent can write an actual cost next to each shot instead of an estimate. Connecting Claude Code to the hosted MCP server is one command: claude mcp add --transport http aitachyon https://aitachyon.com/api/mcp --header "Authorization: Bearer ait_...". The same key works against the plain HTTP API if you would rather run the pipeline as a script.

A brief you can hand the agent

Make a 30-second vertical video for [product]. Return the shot list in the contract format first and stop for my approval. After approval, generate one still per shot on the cheapest image model, then stop again. After my second approval, generate video per the routing rules, generate the voiceover per shot, assemble with Remotion with captions from the script, and write a cost table listing each ref, model and price. Set disclose to true.

The two explicit stops are the human gates. Without them, the agent will happily spend the whole budget on the first version of the script. For the setup steps, see generating AI video from Claude Code with an MCP server.

FAQ

Can AI fully automate video production end to end?

Technically yes: script, shots, voice and assembly can all run from one agent. In practice every published workflow we found keeps a human review before anything goes public, and the platform disclosure rules make a person accountable for the final file anyway. Let the agent do the production and keep the final judgment with a person.

How much does it cost to automate a short video with AI?

It depends on seconds of footage and which model renders them. At the time of writing on Aitachyon, 30 seconds on Hailuo 02 is $2.55, on Veo 3.1 Fast from $5.70 and on Kling v3 with audio $9.60, plus cents for stills and voice. MindStudio's estimate for a five-minute explainer is $2 to $7, or $10 to $20 with premium models.

How long does an automated AI video take to produce?

MindStudio reports 5 to 15 minutes from prompt to export for a 60-second social clip and 20 to 40 minutes for a 2 to 3 minute explainer, and 30 to 60 minutes for a 3 to 5 minute video built from generated clips.

Should I use FFmpeg or Remotion to assemble AI video?

FFmpeg for straight cuts under a voice track with burned captions. Remotion when you need animated text, templates or anything you would otherwise build in a motion design tool, since the agent can write the React composition and render it to MP4.

Do I have to label AI-generated videos on YouTube and TikTok?

For realistic generated footage, yes on both. YouTube asks at upload in YouTube Studio, and TikTok requires a label and can detect C2PA Content Credentials on its own. Fantastical or clearly unrealistic content, and AI used only for scripts or captions, does not need a YouTube disclosure.

Sources

  1. Google AI for Developers: Gemini Developer API pricing (Veo 3.1), September 2026
  2. ElevenLabs: API Pricing, September 2026
  3. YouTube Help: Disclosing use of altered or synthetic content
  4. TikTok Newsroom: Partnering with our industry to advance AI transparency and literacy, May 2024
  5. TechCrunch: TikTok will automatically label AI-generated content created on other platforms, May 2024
  6. Remotion: Agent Skills
  7. MindStudio: AI Agent Workflow, YouTube Video From One Prompt
  8. MindStudio: AI Video Generation for Content Marketing, Multi-Agent Workflow

If you want to run this pipeline without five provider accounts, Aitachyon puts the image, video and voice models above behind one prepaid balance and one key, usable from the web studio, the HTTP API or a hosted MCP server that Claude Code connects to in one command. Failed renders are refunded automatically, each key has its own spend alert, and the API and MCP quick start takes a few minutes.

Related articles

Free tools to try

Stop describing your brand. Paste your URL.

Aitachyon reads your whole brand from your website, then creates videos, images, carousels, posts and banners, on-brand, every format, every feed.