Skip to content
StrategiesSeptember 30, 2026· 9 min read

Bulk AI Video Generator: Producing Hundreds of Clips Safely

How to run hundreds of AI video generations from a script or agent: concurrency limits, backoff, idempotent retries, refunds, per-key spend and saving outputs.

bulkapiautomationcost control
Strategies

Bulk AI Video Generator: Producing Hundreds of Clips Safely

The first time you run three hundred video generations from a script, something breaks around job forty. A provider returns a 429, your loop retries immediately, the retries collide with each other, and by the time the run finishes you have 280 files, a few duplicates you paid for twice, and a dozen jobs whose output links have already expired. The models are fine; the failure comes from treating a batch job like a for-loop.

Below is the plumbing that makes a bulk AI video generator safe to run unattended: pricing a batch up front, queuing within concurrency limits, retrying without paying twice, reconciling refunds, and naming outputs so you can find them next month.

Three limits that decide how fast a batch can go

Generation APIs throttle on three separate axes, which Leonardo.Ai's documentation defines them cleanly:

  • Concurrency is the number of generation jobs processed in parallel.
  • Queue (pending) limit is how many requests may wait when every concurrent slot is busy.
  • API rate limit is how many API requests you can make in a given time window, regardless of what they do.

Concurrency does not cap total requests: extra work waits in the queue if it has room. Concurrency sets your throughput; the rate limit sets how hard you can poll and submit.

What the defaults look like

Defaults are low, and they are usually set per organization rather than per key:

  • fal: new accounts start at 2 concurrent requests, scaling with paid invoices over the last four weeks up to 40 self-serve. Requests sitting IN_QUEUE do not count toward the limit; only IN_PROGRESS ones do.
  • Vidu: up to 5 concurrent tasks per organization, and the limit applies per organization, not per API key.
  • LTX: according to its rate-limit page, a default of up to 2 simultaneous generations, with over-limit requests answered by a 429 and a Retry-After header.
  • Replicate: 600 requests per minute for creating predictions and 3,000 per minute for other endpoints, with short bursts tolerated before throttling.

At a concurrency of 2, a run of 300 clips is a long evening however fast your code is. Treat it as a background job with a manifest and a finish line.

Price the batch before you press run

With pay-per-call pricing, batch cost is a product of numbers you already know. The formula for your runbook:

Batch cost = clips x seconds per clip x price per second + images x price per image + (characters / 450) x price per 450 characters of voice.

Using Aitachyon's public prices at the time of writing (the full list is on the models page with each price per call), here is what realistic batches cost end to end, assuming 5-second clips:

  • 20 hooks on Hailuo 02 at $0.085/s: 20 x 5 x $0.085 = $8.50.
  • The same 20 hooks on Kling v3: $16.00 silent at $0.16/s, or $32.00 with native audio at $0.32/s.
  • The same 20 hooks on Seedance 2.5: $20.00 at 480p ($0.20/s) or $44.00 at 720p ($0.44/s).
  • 50 product shots on Seedream 4.0 at $0.057 per image: $2.85. On FLUX.2 [pro] at $0.057 to $0.086: $2.85 to $4.30. On Nano Banana at $0.13: $6.50.
  • A week of shorts, seven videos each made of three 5-second Hailuo 02 clips plus about 900 characters of ElevenLabs voiceover ($0.043 to $0.086 per 450 characters): per short, 15 x $0.085 = $1.275 for video and $0.086 to $0.172 for voice, so about $9.53 to $10.13 for the week.
  • 300 clips: 1,500 seconds of video. That is $127.50 on Hailuo 02 and $480.00 on Kling v3 with audio.

The same batch varies by roughly four to one across models, so model choice matters more at volume. A common pattern: generate every variant on a cheaper model, pick winners, and re-render only those on the model with the look or native audio you want. The model-by-model API pricing breakdown covers the trade-offs.

Two lines people forget. Retries: with bad retry logic, a 300-clip run bills well over 300 clips. And the keep ratio: if you keep one clip in five, a usable clip costs five times the per-call price. The upstream question is how many ads you should actually run.

The job manifest: name every job before you submit it

The most useful habit in bulk generation is writing the whole batch to a manifest file before any request goes out. It records what should exist, what was submitted, what finished and what it cost. When the script crashes at job 140, you resume from it instead of guessing.

A manifest row that survives a crash

  1. job_key: a deterministic ID built from campaign, shot, model and variant, for example spring-launch_hook-07_kling3_v2.
  2. model and params: model slug, duration, resolution, audio on or off, seed if the model takes one.
  3. prompt: the exact text sent, so you can reproduce a winner.
  4. estimated_cost: from the formula above.
  5. status: planned, submitted, running, done, failed.
  6. provider_id: the request or job ID the API returns on submit.
  7. output_ref and local_path: where the file lives once you have it.
  8. actual_cost and attempts: filled in on completion.

Sum estimated_cost and refuse to start above budget. That check catches the typo that turns 30 clips into 300.

Idempotency: the rule that stops double billing

The dangerous moment is a network timeout on submit: you do not know whether the job was created, and a blind retry may bill it twice. Stripe documents the standard defence for paid POSTs: the server stores the status code and body of the first request per idempotency key, including 500 errors, and returns that same result to repeats. Keys can be up to 255 characters, may be pruned after at least 24 hours, and reusing a key with different parameters returns an error.

Many generation APIs do not document idempotency; fal's reliability page does not explicitly address idempotency guarantees. When the provider offers no key, your manifest is the idempotency layer: before resubmitting a job whose submit timed out, look for it on the provider side and resubmit only once you have confirmed it does not exist.

Queue the work, cap what is in flight

Video APIs are asynchronous: submit, receive an ID, come back for the result. fal's queue returns a request_id to store for later status checks and moves each request through IN_QUEUE, IN_PROGRESS and COMPLETED. OpenAI's video guide used the same shape, with queued, in_progress, completed and failed statuses.

The pattern that avoids mass rejections, per LTX's documentation, is a local queue that limits active jobs to the allowed concurrency. In practice:

  1. Load all planned rows from the manifest into a local queue.
  2. Start a worker pool sized to your concurrency limit. If the provider counts only running jobs, as fal does, you can submit ahead into its queue; fal states that queued requests are never dropped due to concurrency limits unless a start_timeout expires.
  3. Each worker submits one job, records the provider ID in the manifest immediately, then waits for completion.
  4. On completion, the worker downloads the output, writes the row, and takes the next job.
  5. Put the jobs you need first at the front, and mark bulk variants low priority where supported; fal accepts priority set to low for this.

Webhooks or polling

Prefer webhooks when you can receive them. Replicate sends HTTP POSTs when predictions are created, updated and finished, and fal accepts a webhook_url on submit and posts the request ID, status and output on completion. Make the handler idempotent: the same completion can arrive more than once, so look up the job by ID, and do nothing if the row is already done.

If you poll, poll slowly. OpenAI suggested every 10 to 20 seconds, with exponential backoff if needed. Polling 300 jobs every second burns your general rate limit (Replicate's is 3,000 requests per minute for non-creation endpoints) and gets you throttled on the calls that matter.

Retries that do not multiply the bill

Retrying after a fixed delay synchronizes your workers: they wake up together and collide again. AWS measured this in its study of exponential backoff and jitter. With 100 contending clients, full jitter cut total calls by more than half compared with un-jittered exponential backoff, and un-jittered backoff took so much longer that it was left off the completion-time graph.

The full jitter formula:

sleep = random(0, min(cap, base x 2^attempt))

With base = 1 second and cap = 60 seconds, attempt 3 sleeps a random time between 0 and 8 seconds. fal's own guidance for raw HTTP clients follows the same shape, backing off 1s, 2s, 4s, 8s, and its SDK retries up to 10 times.

A retry decision rule

Not every error deserves a retry:

  1. 429 with a Retry-After header: wait exactly that long, plus a little jitter. Replicate's throttle message tells you when the limit resets, around 30 seconds.
  2. 429 without a header, or fal's concurrent_requests_limit (fal also sets X-Fal-needs-retry: 1): full-jitter backoff, and reduce your worker count by one if it keeps happening.
  3. 503, 504 or a connection error: retry with backoff. These are the errors fal's queue retries automatically, alongside 429.
  4. Timeout on submit: do not resubmit until you have checked whether the job exists (see idempotency above).
  5. 4xx validation or content errors: do not retry. Mark the row failed with the message, fix the prompt or parameters, and requeue it as a new variant.
  6. Job completed with status failed: retry once with the same parameters. If it fails again, treat it as a prompt problem and move on.

Cap attempts per job (three is a reasonable default) and total retries per run. When a tenth of the jobs are retrying, something is wrong upstream: pause the run.

Failures, refunds and reconciliation

At volume, some renders fail, and the question is who pays. fal states that failed requests returning 5xx incur no charges. On Aitachyon, a render that fails is refunded automatically, to the cent, and every job is itemised with the model it ran on and what it cost. Whatever your provider's policy, reconcile after every run instead of trusting it.

Post-run reconciliation checklist

  • Count rows by status. Planned minus done minus failed should be zero; anything left is stuck and needs a status check.
  • Sum actual_cost and compare it with estimated_cost. A gap above a few percent usually means duplicate submissions or a model or duration that differed from the plan.
  • Match every charged job to a manifest row. A charge with no row is a duplicate from a retry.
  • Check that every failed job shows its refund or a zero charge, and every done row has a local file of the expected duration.
  • Record the selection: which outputs you kept. That ratio feeds the next batch budget.

With a script this takes minutes, and it carries straight over to producing 50 UGC variants a week, where it keeps weekly spend flat.

Get the files out before they disappear

Provider-hosted outputs are temporary. Replicate deletes output files of API-created predictions after an hour and tells you to save a copy; web-interface predictions keep their files, so a workflow that worked by hand can lose files once scripted. OpenAI's video download URLs were valid for at most 1 hour after generation.

Run overnight, download in the morning, and you get dead links. Download in the completion handler, in the step that marks the row done.

A naming convention you can search

Rename each file on download to its job key, for example spring-launch_hook-07_kling3_v2.mp4, and keep the provider's reference in the manifest. Aitachyon gives every generated file a stable ref (img_, scn_, vo_ and so on) that code can fetch at any time with GET /api/generations/{ref}, so the ref is the durable handle and the local file is your working copy. With both in the manifest, a question like "which prompt made the hook that won in March" is a grep.

Provider dependency is a batch risk too

OpenAI's guide now notes that the Sora 2 models and Videos API were discontinued on September 24, 2026. Pipelines wired to that endpoint had to be rewritten. With the model kept as a manifest parameter, a retirement costs a config change.

Cost control per key, and running batches from an agent

A buggy script with an unlimited balance is the expensive combination. Replicate applies progressively stricter limits as credit decreases and recommends keeping the balance above $20 with auto-reload, while accounts without a payment method are held to 1 request per second. Your own guardrails should sit on top:

Spend guardrails for unattended runs

  1. One API key per pipeline. The nightly shorts job, the ad-variant script and your agent session each get their own key, so spend is attributable. On Aitachyon each key has its own spend, an alert when it spends unusually fast, and a one-click revoke.
  2. A prepaid ceiling. A prepaid balance is a hard cap by construction. Aitachyon's balance never expires, so topping up for a large run does not create use-it-or-lose-it pressure. The comparison of credits and pay-per-call goes through where each model leaks money.
  3. A budget check in code. The manifest total, checked before the run and again every 50 jobs against actual spend.
  4. A kill switch. If the per-key alert fires, revoke the key. The workers fail fast on auth errors, which is what you want.

The same batch, driven by an agent

Increasingly a coding agent writes and runs the script. With a hosted MCP server connected to Claude Code or Cursor, you describe the batch ("twenty 5-second hooks from these prompts on Hailuo 02, save them to ./hooks named by hook number") and the agent submits, waits, downloads and reports each file with its cost. Connecting Claude Code to Aitachyon is one command:

claude mcp add --transport http aitachyon https://aitachyon.com/api/mcp --header "Authorization: Bearer ait_..."

Agent runs need the same guardrails: a dedicated key, a budget stated in the prompt, and the manifest written to disk as it goes. The walkthrough on generating video from Claude Code over MCP shows the setup, and a working agent setup for video production covers chaining scripts, clips and voice into finished pieces.

Pre-flight checklist for a bulk run

Copy this into the top of your batch script or your agent prompt:

  1. Manifest written, one row per output, deterministic job keys.
  2. Estimated total below budget.
  3. A dedicated API key with spend alerts.
  4. Workers capped at the concurrency limit.
  5. Full-jitter backoff, Retry-After honored, three attempts max.
  6. Timed-out submits checked before any resubmit.
  7. Webhook handler or slow polling, idempotent either way.
  8. Download and rename on completion.
  9. A test batch of five run end to end.
  10. Reconciliation script ready.

The test batch is the step people skip and the cheapest: on Hailuo 02 at the time of writing, five 5-second clips cost $2.13, and they catch the wrong aspect ratio before it costs you three hundred times over.

FAQ

Do I pay for failed AI video generations?

It depends on the provider. fal does not charge for requests that fail with a 5xx error, and Aitachyon refunds failed renders automatically to the cent. Either way, reconcile your itemised charges against your job list after each run, because a duplicate submission from a bad retry is a successful job that you did pay for.

How long do generated video files stay available?

Often not long. Replicate deletes API output files after an hour, and OpenAI's video download URLs lasted at most an hour. Download each file as soon as its job completes and store it under your own name. On Aitachyon, each output also has a stable ref you can fetch later with GET /api/generations/{ref}.

What does it cost to generate 100 short AI videos?

For 100 five-second clips at Aitachyon's prices at the time of writing, that is $42.50 on Hailuo 02, $80 on Kling v3 silent, and $100 on Seedance 2.5 at 480p. Divide by your keep ratio to get the cost per usable clip.

Sources

  1. Replicate: Rate limits
  2. Replicate: Output files
  3. Replicate: Webhooks
  4. fal: Concurrency limits
  5. fal: Async inference and queue
  6. fal: Reliability
  7. Leonardo.Ai: Guide to concurrency, queue and rate limit
  8. Vidu: Usage and limits
  9. Stripe: Idempotent requests
  10. AWS Architecture Blog: Exponential backoff and jitter

If you want to run this kind of batch without wiring up a separate account per model, Aitachyon puts Seedance, Kling, Veo, Hailuo, the image models and ElevenLabs voice behind one key, usable from an HTTP API or the hosted MCP server, with every job itemised, failed renders refunded, and spend tracked per key. The API and MCP quick start has the setup.

Related articles

Free tools to try

Stop describing your brand. Paste your URL.

Aitachyon reads your whole brand from your website, then creates videos, images, carousels, posts and banners, on-brand, every format, every feed.