Veo 3.1 Fast API: Price, Limits and How to Call It
Veo 3.1 Fast API pricing per second, clip lengths, resolutions, latency and quotas from Google's docs, plus how to call it from code or from Claude Code.
Veo 3.1 Fast API: Price, Limits and How to Call It
An 8-second Veo 3.1 Fast clip at 720p costs $0.80 on Google's own API. The same clip at 4K costs $2.40. Both figures come from the Gemini Developer API pricing page, and they shape most of the decisions that follow: which resolution you draft at, how many variants you can afford to discard, and whether you call Google directly or go through a gateway.
Below: what a second of Veo 3.1 Fast costs in each place you can buy it, the hard limits in Google's documentation, and the call sequence from a script or from an agent such as Claude Code. Every price is at the time of writing, September 2026.
What Veo 3.1 Fast is and where it sits in the family
Google announced Veo 3.1 and Veo 3.1 Fast on October 15, 2025, in paid preview across the Gemini API, AI Studio and Vertex AI, according to the Google Developers Blog announcement. The same release added reference image guidance with up to three images, scene extension for longer videos, and generation from a first and last frame. Google priced Veo 3.1 the same as Veo 3 at launch.
The Gemini API video overview describes Veo 3.1 as a model that generates video with native audio, and lists video extension, frame-specific generation and image-based direction as supported inputs. There are three tiers, each with its own model code in the Veo 3.1 developer docs:
- veo-3.1-generate-preview: Standard, the most expensive tier.
- veo-3.1-fast-generate-preview: Fast, the subject of this guide.
- veo-3.1-lite-generate-preview: Lite, the cheapest, with no 4K output.
Vercel's gateway listing sums up the Fast tier's purpose well: it is optimized for workflows where volume and iteration speed matter as much as final quality.
A decision rule for the three tiers
- Use Lite for throwaway drafts where you only need to judge composition and motion, and never need 4K.
- Use Fast as the default for anything that could ship: hooks, B-roll, product shots, shorts.
- Use Standard only for a shot you have already validated with Fast and want to re-render, because it costs four times as much per second at 720p.
Veo 3.1 Fast pricing per second and per clip
Google bills Veo per second of output. The Gemini API pricing page lists these rates at the time of writing, with audio included:
- Veo 3.1 Fast: $0.10/s at 720p, $0.12/s at 1080p, $0.30/s at 4K.
- Veo 3.1 Standard: $0.40/s at 720p and 1080p, $0.60/s at 4K.
- Veo 3.1 Lite: $0.05/s at 720p, $0.08/s at 1080p, 4K unsupported.
The same page states: "You will only be charged if your video is successfully generated." Filtered or failed jobs do not appear on the bill.
Converted to clips (Veo renders 4, 6 or 8 seconds, and 1080p and 4K only at 8), Google's list price gives:
- Fast, 720p: $0.40 for 4s, $0.60 for 6s, $0.80 for 8s.
- Fast, 1080p (8s only): $0.96.
- Fast, 4K (8s only): $2.40.
- Standard, 720p, 8s: $3.20. Standard, 4K, 8s: $4.80.
- Lite, 720p, 8s: $0.40. Lite, 1080p, 8s: $0.64.
BenchLM's pricing tracker notes that at 720p, Veo 3.1 Fast sits at or below Sora 2 standard at $0.10/s, which puts it at the low end of the market for a model that produces audio in the same pass.
Gateways price the same model differently
Google is not the only place to buy it, and the gateways do not all copy Google's grid:
- fal.ai splits audio out: $0.10/s without audio and $0.15/s with audio at 720p or 1080p, and $0.30/s or $0.35/s at 4K. Its generate_audio parameter defaults to true, so a default 8-second call costs $1.20, and turning audio off brings it back to $0.80.
- Leonardo.Ai bills in its own API credits: 546 credits for 4s, 819 for 6s and 1,092 for 8s, which scales linearly with duration.
- Vercel AI Gateway exposes it under the model ID google/veo-3.1-fast-generate-001.
- Aitachyon lists Veo 3.1 Fast from $0.19/s at the time of writing, billed from a prepaid balance. That is above Google's direct list price.
For a side-by-side of every major video model's per-second rate, see the model-by-model video API pricing breakdown.
What a realistic batch costs
Batch 1: 20 hook variants. Twenty 8-second clips at 720p is 160 seconds of output.
- Google direct: 160 x $0.10 = $16.00.
- fal with audio: 160 x $0.15 = $24.00.
- Aitachyon, from: 160 x $0.19 = $30.40.
Then re-render the three winners at 1080p on Google: 3 x $0.96 = $2.88. Drafting all twenty at 1080p would cost $19.20, a small difference. At 4K it is large: twenty 4K drafts cost $48.00, three times the 720p batch.
Batch 2: a week of shorts. Three shorts a day, each stitched from three 8-second clips, is 63 clips or 504 seconds a week.
- Google direct at 720p: $50.40. At 1080p: $60.48. At 4K: $151.20.
- Aitachyon, from $0.19/s: $95.76.
Those totals assume every clip is usable first time. If you keep one clip in two, double them. Hook volume is where this adds up fastest, and a library of tested hook formulas reduces the number of blind renders you pay for.
The hard limits in Google's documentation
These constraints come from the Veo 3.1 page of the Gemini API docs unless stated otherwise:
- Duration: durationSeconds accepts 4, 6 or 8. Eight seconds is required for higher resolutions and for reference images.
- Resolution: 720p (default), 1080p or 4k. 1080p and 4k only work at 8 seconds.
- Aspect ratio: 16:9 (default) or 9:16. Nothing else.
- Latency: a minimum of 11 seconds and up to 6 minutes during peak hours.
- Retention: generated videos are stored on Google's servers for 2 days, then removed.
- Watermark: all videos receive SynthID watermarking.
- Regional rules: in the EU, UK, Switzerland and MENA, allow_adult is the only permitted personGeneration value.
- Rate limits: the Gemini rate limits page does not list Veo in its published tables. Per-project limits are visible in AI Studio at aistudio.google.com/rate-limit.
- Frame rate: Vercel's listing gives 24 fps for both text-to-video and image-to-video.
- Exact frame sizes: Leonardo documents only 1280x720, 720x1280, 1920x1080 and 1080x1920 on its integration.
What those limits mean in a real pipeline
- Eight seconds is the unit of work. Anything longer is assembled from several clips or built with scene extension. Write scripts as a sequence of 8-second beats from the start.
- There is no cheap short 1080p clip. A 4-second 1080p render does not exist, so a 1080p shot always costs at least $0.96 on Google's grid. If you only need 4 seconds, draft at 720p and trim.
- Vertical is native. 9:16 is a first-class option, so you do not have to crop a landscape render for Shorts, Reels or TikTok. Framing still matters, and the rules for producing 9:16 video apply to generated footage as much as to filmed footage.
- Design for asynchronous work. A range of 11 seconds to 6 minutes rules out a synchronous request that holds an HTTP connection open. Queue jobs, poll, and let the rest of the pipeline continue.
- Download immediately. After 2 days the file is gone. A pipeline that stores Google's URL instead of the file will break two days later.
- Plan for the watermark. SynthID marks every output as AI-generated, so write your disclosures accordingly.
- Check person generation by market. If your team or your users are in the EU, UK, Switzerland or MENA, test a prompt with people in it from that region before you promise a feature built on it.
How to call Veo 3.1 Fast from code
The Gemini API flow is a long-running operation. The official docs describe the same four steps for every SDK:
- Submit. In Python, call client.models.generate_videos with the model veo-3.1-fast-generate-preview and your prompt. Over REST, send a POST to models/veo-3.1-fast-generate-preview:predictLongRunning. Set durationSeconds, resolution and aspectRatio explicitly rather than relying on defaults, so your cost per call is predictable.
- Receive an operation. The response is an operation handle, not a video.
- Poll. Re-fetch the operation until it reports done.
- Download. Fetch the generated file and write it to your own storage within the 2-day window.
Polling without wasting requests
The documented latency floor is 11 seconds, so make the first check at around 10 to 15 seconds, then poll at a fixed interval with a hard timeout above the 6-minute peak. Log the operation name with the prompt, so a timed-out job can be resumed instead of re-submitted and billed twice.
The same call through a gateway
- fal.ai (Python): fal_client.subscribe('fal-ai/veo3.1/fast', arguments={...}) handles the queue for you. The documented parameters are aspect_ratio, duration (4s, 6s or 8s, default 8s), resolution, generate_audio (default true), negative_prompt, seed, auto_fix and safety_tolerance from 1 to 6.
- Vercel AI SDK: experimental_generateVideo({ model, prompt }) with the model set to google/veo-3.1-fast-generate-001, per Vercel's model page.
One gateway detail with a real cost: if you plan to lay a separate voiceover over the footage, turning off generate_audio on fal saves $0.05 per second, or $0.40 per 8-second clip. On a 20-clip batch that is $8.00. The voice track then becomes its own line item, and the voiceover cost breakdown shows how to price it.
Calling it from an agent instead of a script
Running generation from an agent removes the polling loop from your codebase. In Claude Code or Cursor, an MCP server exposes video generation as a tool; the agent writes the prompts, calls the tool, waits for the job and saves the file, and you review the output and the cost. Generating AI video from Claude Code over MCP walks through that setup end to end.
With a hosted MCP server the connection is one command. For Aitachyon it is:
claude mcp add --transport http aitachyon https://aitachyon.com/api/mcp --header "Authorization: Bearer ait_..."
The API and MCP quick start covers the same key used over plain HTTP. Every generated file comes back with a stable ref (scn_ for a scene) that code can fetch later with GET /api/generations/{ref}, which avoids the 2-day expiry problem of storing a provider's temporary URL.
A brief you can paste into the agent
Agents spend money as fast as you let them. Put the constraints in the brief, not in your head:
- Task: generate one Veo 3.1 Fast clip per hook listed in hooks.md.
- Format: 8 seconds, 9:16, 720p. Do not change resolution or duration without asking.
- Output: save each file as renders/[hook-id].mp4 and append model, duration, resolution, prompt and cost to renders/costs.csv.
- Budget: stop and report if the running total passes $20.
- Escalation: ask before any 1080p or 4K render, and before retrying a failed job more than once.
- Review: when done, list the three clips whose first two seconds best match the hook text.
Twelve hooks under that brief cost $9.60 at Google's 720p list price or from $18.24 at $0.19/s, so a $20 cap holds either way. For a fuller pipeline with scripting, voice and assembly in one run, see a working setup for automating video production with agents.
Writing prompts that fit an 8-second clip
Eight seconds holds one idea. Most failed renders come from prompts that describe a whole scene arc. A structure that fits the limits:
- Subject: who or what, with one or two concrete visual details.
- Single action: one movement that can finish in the clip.
- Camera: framing and one camera move.
- Light and setting: time of day, location, surface.
- Sound: since audio is generated in the same pass, say what should be heard, or say silence.
Before: "A woman discovers our skincare product, tries it, loves it, and shows the results to her friends at a party."
After: "Close-up, vertical frame. A woman in her thirties presses a pump bottle of face serum into her palm in a bright bathroom at morning light. Slow push-in on her hands. Sound: the soft click of the pump and a quiet room tone, no music."
The first prompt asks for four scenes in eight seconds; the second asks for one shot. When you need continuity across shots, use the controls Google shipped with 3.1: up to three reference images for a consistent product or character, and first and last frame generation to pin where a clip starts and ends. Reference images require the 8-second duration, so budget them at the 8-second price. On fal, fixing the seed lets you change one word of a prompt and compare like with like.
When to pick Veo 3.1 Fast over other video models
The choice is usually a constraint problem. Using Aitachyon's own grid at the time of writing, the per-second rates for the other video models are: Hailuo 02 at $0.085/s, Kling v3 at $0.16/s silent and $0.32/s with native audio, Wan 2.7 at $0.19/s, and Seedance 2.5 from $0.20/s at 480p and $0.44/s at 720p. Veo 3.1 Fast is listed from $0.19/s on the same grid.
- Lowest cost per drafted second: Hailuo 02 on that grid, or Veo 3.1 Lite at $0.05/s if you buy from Google directly.
- Audio in the same render: Veo 3.1 Fast includes audio in Google's price; Kling v3 charges double for its audio mode.
- 4K output: Veo 3.1 Fast or Standard. Lite does not render 4K.
- Clips longer than 8 seconds in one call: Veo cannot, so either chain Veo clips with scene extension or test another model.
For a head-to-head of the two most-requested alternatives, read the Kling v3 vs Seedance 2.5 comparison. Most teams do best keeping two models in rotation and routing each shot to the cheapest one that meets its constraint.
Pre-flight checklist before a Veo 3.1 Fast batch
- Confirm your project's Veo quota in AI Studio.
- Pin the model code in config and set durationSeconds, resolution and aspectRatio on every call.
- Draft at 720p. Promote only the winners to 1080p or 4K.
- Decide whether you need generated audio; on fal it changes the price by half again.
- Multiply the batch cost by your expected reject rate and set a hard budget cap.
- Check personGeneration rules for the region you operate in.
- Copy every file to your own storage on completion, before the 2-day deletion.
- Log prompt, seed, model, resolution, duration and cost per clip so winners can be reproduced.
FAQ
How much does the Veo 3.1 Fast API cost per second?
On Google's Gemini API, at the time of writing, $0.10 per second at 720p, $0.12 at 1080p and $0.30 at 4K, audio included, so an 8-second 720p clip is $0.80. Gateways price it differently.
What is the model name for Veo 3.1 Fast in the Gemini API?
veo-3.1-fast-generate-preview. Vercel AI Gateway uses google/veo-3.1-fast-generate-001 and fal uses fal-ai/veo3.1/fast.
How long can a Veo 3.1 Fast video be?
Each call returns 4, 6 or 8 seconds, and 1080p, 4K and reference images require 8. Longer videos are built by chaining clips or with scene extension.
What are the Veo 3.1 Fast rate limits?
Google does not list Veo in its published rate-limit tables. Your project's limits are shown in AI Studio at aistudio.google.com/rate-limit, and they are the number to plan concurrency around.
Does Veo 3.1 Fast generate audio, and can I remove the watermark?
Veo 3.1 generates native audio, and Google's per-second price includes it. Every output carries a SynthID watermark, and the docs do not document an option to disable it.
Sources
- Google AI for Developers: Gemini Developer API pricing
- Google AI for Developers: Generate videos with Veo 3.1 in Gemini API
- Google AI for Developers: Video generation with Veo
- Google AI for Developers: Gemini API rate limits
- Google Developers Blog: Introducing Veo 3.1 and new creative capabilities in the Gemini API (October 15, 2025)
- fal.ai: Veo 3.1 Fast
- Vercel: Veo 3.1 Fast Generate on AI Gateway
- Leonardo.Ai: Generate with Veo3.1, Veo3.1 Fast using text prompts
- BenchLM.ai: Veo 3.1 API pricing
If Veo 3.1 Fast is one of several models in your pipeline, Aitachyon puts it behind the same key as Kling, Seedance, Hailuo, the image models and ElevenLabs, callable from the studio, the HTTP API or Claude Code over MCP. Each job is itemised with its model and cost, failed renders are refunded to the cent, and the prepaid balance does not expire. The per-second rate is higher than Google's direct price, so if you only ever call Veo, going to Google directly is the cheaper option.
Related articles
Seedance 2.5 API Pricing: What a Clip Really Costs
Seedance 2.5 API pricing per second at 480p, 720p and 1080p, worked costs for 5s, 10s and 30s clips and batches, and where to get access with no subscription.
GuidesAI Video Generation API Pricing in 2026, Model by Model
How AI video APIs bill (per second, per clip, credits), what drives the price, a model-by-model rate list and worked batch costs with retries included.
GuidesElevenLabs API Pricing: What a Voiceover Really Costs (2026)
ElevenLabs API pricing per 1,000 characters by model, what a 30s or 60s voiceover costs, and how to batch-generate voiceovers from code without 429 errors.
GuidesReal Estate Video Ads: The Media Buyer's Playbook for Booked Viewings
A data-driven guide to real estate video ads—per-listing cost math, platform-by-platform CPLs, the geo-first targeting compliance rules, and CTA funnel logic.
GuidesFitness Studio Video Ads: The Gym Owner's 2026 Playbook
How to run fitness studio video ads that fill classes—compliant transformation framing, paid trials, local targeting, and weekly creative refresh.
GuidesBlack Friday Video Ads: A Two-Week Production Plan
Black Friday video ads are short offer creatives built and tested before Cyber Five CPMs spike. Here is the day-by-day plan to ship them on time.
Free tools to try
Free AI image generator
Describe what you want and get a high-quality AI image in seconds. A free AI image generator, no account needed to preview, keep your first image when you sign up.
Try it freeFree toolFree AI product photo generator
Generate clean, studio-style product photos for your store and listings in seconds. Crisp lighting and tidy backgrounds, free to try with no account, keep your first shot on signup.
Try it freeFree toolFree background remover
Remove the background from any image in seconds and get a clean, transparent cutout. A free background remover, no account needed to preview, keep your first cutout when you sign up.
Try it freeStop describing your brand. Paste your URL.
Aitachyon reads your whole brand from your website, then creates videos, images, carousels, posts and banners, on-brand, every format, every feed.