Skip to content
GuidesSeptember 30, 2026· 8 min read

Veo 3.1 Fast API: Price, Limits and How to Call It

Veo 3.1 Fast API pricing per second, clip lengths, resolutions, latency and quotas from Google's docs, plus how to call it from code or from Claude Code.

veogooglevideo api
Guides

Veo 3.1 Fast API: Price, Limits and How to Call It

An 8-second Veo 3.1 Fast clip at 720p costs $0.80 on Google's own API. The same clip at 4K costs $2.40. Both figures come from the Gemini Developer API pricing page, and they shape most of the decisions that follow: which resolution you draft at, how many variants you can afford to discard, and whether you call Google directly or go through a gateway.

Below: what a second of Veo 3.1 Fast costs in each place you can buy it, the hard limits in Google's documentation, and the call sequence from a script or from an agent such as Claude Code. Every price is at the time of writing, September 2026.

What Veo 3.1 Fast is and where it sits in the family

Google announced Veo 3.1 and Veo 3.1 Fast on October 15, 2025, in paid preview across the Gemini API, AI Studio and Vertex AI, according to the Google Developers Blog announcement. The same release added reference image guidance with up to three images, scene extension for longer videos, and generation from a first and last frame. Google priced Veo 3.1 the same as Veo 3 at launch.

The Gemini API video overview describes Veo 3.1 as a model that generates video with native audio, and lists video extension, frame-specific generation and image-based direction as supported inputs. There are three tiers, each with its own model code in the Veo 3.1 developer docs:

  • veo-3.1-generate-preview: Standard, the most expensive tier.
  • veo-3.1-fast-generate-preview: Fast, the subject of this guide.
  • veo-3.1-lite-generate-preview: Lite, the cheapest, with no 4K output.

Vercel's gateway listing sums up the Fast tier's purpose well: it is optimized for workflows where volume and iteration speed matter as much as final quality.

A decision rule for the three tiers

  1. Use Lite for throwaway drafts where you only need to judge composition and motion, and never need 4K.
  2. Use Fast as the default for anything that could ship: hooks, B-roll, product shots, shorts.
  3. Use Standard only for a shot you have already validated with Fast and want to re-render, because it costs four times as much per second at 720p.

Veo 3.1 Fast pricing per second and per clip

Google bills Veo per second of output. The Gemini API pricing page lists these rates at the time of writing, with audio included:

  • Veo 3.1 Fast: $0.10/s at 720p, $0.12/s at 1080p, $0.30/s at 4K.
  • Veo 3.1 Standard: $0.40/s at 720p and 1080p, $0.60/s at 4K.
  • Veo 3.1 Lite: $0.05/s at 720p, $0.08/s at 1080p, 4K unsupported.

The same page states: "You will only be charged if your video is successfully generated." Filtered or failed jobs do not appear on the bill.

Converted to clips (Veo renders 4, 6 or 8 seconds, and 1080p and 4K only at 8), Google's list price gives:

  • Fast, 720p: $0.40 for 4s, $0.60 for 6s, $0.80 for 8s.
  • Fast, 1080p (8s only): $0.96.
  • Fast, 4K (8s only): $2.40.
  • Standard, 720p, 8s: $3.20. Standard, 4K, 8s: $4.80.
  • Lite, 720p, 8s: $0.40. Lite, 1080p, 8s: $0.64.

BenchLM's pricing tracker notes that at 720p, Veo 3.1 Fast sits at or below Sora 2 standard at $0.10/s, which puts it at the low end of the market for a model that produces audio in the same pass.

Gateways price the same model differently

Google is not the only place to buy it, and the gateways do not all copy Google's grid:

For a side-by-side of every major video model's per-second rate, see the model-by-model video API pricing breakdown.

What a realistic batch costs

Batch 1: 20 hook variants. Twenty 8-second clips at 720p is 160 seconds of output.

  • Google direct: 160 x $0.10 = $16.00.
  • fal with audio: 160 x $0.15 = $24.00.
  • Aitachyon, from: 160 x $0.19 = $30.40.

Then re-render the three winners at 1080p on Google: 3 x $0.96 = $2.88. Drafting all twenty at 1080p would cost $19.20, a small difference. At 4K it is large: twenty 4K drafts cost $48.00, three times the 720p batch.

Batch 2: a week of shorts. Three shorts a day, each stitched from three 8-second clips, is 63 clips or 504 seconds a week.

  • Google direct at 720p: $50.40. At 1080p: $60.48. At 4K: $151.20.
  • Aitachyon, from $0.19/s: $95.76.

Those totals assume every clip is usable first time. If you keep one clip in two, double them. Hook volume is where this adds up fastest, and a library of tested hook formulas reduces the number of blind renders you pay for.

The hard limits in Google's documentation

These constraints come from the Veo 3.1 page of the Gemini API docs unless stated otherwise:

  • Duration: durationSeconds accepts 4, 6 or 8. Eight seconds is required for higher resolutions and for reference images.
  • Resolution: 720p (default), 1080p or 4k. 1080p and 4k only work at 8 seconds.
  • Aspect ratio: 16:9 (default) or 9:16. Nothing else.
  • Latency: a minimum of 11 seconds and up to 6 minutes during peak hours.
  • Retention: generated videos are stored on Google's servers for 2 days, then removed.
  • Watermark: all videos receive SynthID watermarking.
  • Regional rules: in the EU, UK, Switzerland and MENA, allow_adult is the only permitted personGeneration value.
  • Rate limits: the Gemini rate limits page does not list Veo in its published tables. Per-project limits are visible in AI Studio at aistudio.google.com/rate-limit.
  • Frame rate: Vercel's listing gives 24 fps for both text-to-video and image-to-video.
  • Exact frame sizes: Leonardo documents only 1280x720, 720x1280, 1920x1080 and 1080x1920 on its integration.

What those limits mean in a real pipeline

  1. Eight seconds is the unit of work. Anything longer is assembled from several clips or built with scene extension. Write scripts as a sequence of 8-second beats from the start.
  2. There is no cheap short 1080p clip. A 4-second 1080p render does not exist, so a 1080p shot always costs at least $0.96 on Google's grid. If you only need 4 seconds, draft at 720p and trim.
  3. Vertical is native. 9:16 is a first-class option, so you do not have to crop a landscape render for Shorts, Reels or TikTok. Framing still matters, and the rules for producing 9:16 video apply to generated footage as much as to filmed footage.
  4. Design for asynchronous work. A range of 11 seconds to 6 minutes rules out a synchronous request that holds an HTTP connection open. Queue jobs, poll, and let the rest of the pipeline continue.
  5. Download immediately. After 2 days the file is gone. A pipeline that stores Google's URL instead of the file will break two days later.
  6. Plan for the watermark. SynthID marks every output as AI-generated, so write your disclosures accordingly.
  7. Check person generation by market. If your team or your users are in the EU, UK, Switzerland or MENA, test a prompt with people in it from that region before you promise a feature built on it.

How to call Veo 3.1 Fast from code

The Gemini API flow is a long-running operation. The official docs describe the same four steps for every SDK:

  1. Submit. In Python, call client.models.generate_videos with the model veo-3.1-fast-generate-preview and your prompt. Over REST, send a POST to models/veo-3.1-fast-generate-preview:predictLongRunning. Set durationSeconds, resolution and aspectRatio explicitly rather than relying on defaults, so your cost per call is predictable.
  2. Receive an operation. The response is an operation handle, not a video.
  3. Poll. Re-fetch the operation until it reports done.
  4. Download. Fetch the generated file and write it to your own storage within the 2-day window.

Polling without wasting requests

The documented latency floor is 11 seconds, so make the first check at around 10 to 15 seconds, then poll at a fixed interval with a hard timeout above the 6-minute peak. Log the operation name with the prompt, so a timed-out job can be resumed instead of re-submitted and billed twice.

The same call through a gateway

  • fal.ai (Python): fal_client.subscribe('fal-ai/veo3.1/fast', arguments={...}) handles the queue for you. The documented parameters are aspect_ratio, duration (4s, 6s or 8s, default 8s), resolution, generate_audio (default true), negative_prompt, seed, auto_fix and safety_tolerance from 1 to 6.
  • Vercel AI SDK: experimental_generateVideo({ model, prompt }) with the model set to google/veo-3.1-fast-generate-001, per Vercel's model page.

One gateway detail with a real cost: if you plan to lay a separate voiceover over the footage, turning off generate_audio on fal saves $0.05 per second, or $0.40 per 8-second clip. On a 20-clip batch that is $8.00. The voice track then becomes its own line item, and the voiceover cost breakdown shows how to price it.

Calling it from an agent instead of a script

Running generation from an agent removes the polling loop from your codebase. In Claude Code or Cursor, an MCP server exposes video generation as a tool; the agent writes the prompts, calls the tool, waits for the job and saves the file, and you review the output and the cost. Generating AI video from Claude Code over MCP walks through that setup end to end.

With a hosted MCP server the connection is one command. For Aitachyon it is:

claude mcp add --transport http aitachyon https://aitachyon.com/api/mcp --header "Authorization: Bearer ait_..."

The API and MCP quick start covers the same key used over plain HTTP. Every generated file comes back with a stable ref (scn_ for a scene) that code can fetch later with GET /api/generations/{ref}, which avoids the 2-day expiry problem of storing a provider's temporary URL.

A brief you can paste into the agent

Agents spend money as fast as you let them. Put the constraints in the brief, not in your head:

  • Task: generate one Veo 3.1 Fast clip per hook listed in hooks.md.
  • Format: 8 seconds, 9:16, 720p. Do not change resolution or duration without asking.
  • Output: save each file as renders/[hook-id].mp4 and append model, duration, resolution, prompt and cost to renders/costs.csv.
  • Budget: stop and report if the running total passes $20.
  • Escalation: ask before any 1080p or 4K render, and before retrying a failed job more than once.
  • Review: when done, list the three clips whose first two seconds best match the hook text.

Twelve hooks under that brief cost $9.60 at Google's 720p list price or from $18.24 at $0.19/s, so a $20 cap holds either way. For a fuller pipeline with scripting, voice and assembly in one run, see a working setup for automating video production with agents.

Writing prompts that fit an 8-second clip

Eight seconds holds one idea. Most failed renders come from prompts that describe a whole scene arc. A structure that fits the limits:

  1. Subject: who or what, with one or two concrete visual details.
  2. Single action: one movement that can finish in the clip.
  3. Camera: framing and one camera move.
  4. Light and setting: time of day, location, surface.
  5. Sound: since audio is generated in the same pass, say what should be heard, or say silence.

Before: "A woman discovers our skincare product, tries it, loves it, and shows the results to her friends at a party."

After: "Close-up, vertical frame. A woman in her thirties presses a pump bottle of face serum into her palm in a bright bathroom at morning light. Slow push-in on her hands. Sound: the soft click of the pump and a quiet room tone, no music."

The first prompt asks for four scenes in eight seconds; the second asks for one shot. When you need continuity across shots, use the controls Google shipped with 3.1: up to three reference images for a consistent product or character, and first and last frame generation to pin where a clip starts and ends. Reference images require the 8-second duration, so budget them at the 8-second price. On fal, fixing the seed lets you change one word of a prompt and compare like with like.

When to pick Veo 3.1 Fast over other video models

The choice is usually a constraint problem. Using Aitachyon's own grid at the time of writing, the per-second rates for the other video models are: Hailuo 02 at $0.085/s, Kling v3 at $0.16/s silent and $0.32/s with native audio, Wan 2.7 at $0.19/s, and Seedance 2.5 from $0.20/s at 480p and $0.44/s at 720p. Veo 3.1 Fast is listed from $0.19/s on the same grid.

  • Lowest cost per drafted second: Hailuo 02 on that grid, or Veo 3.1 Lite at $0.05/s if you buy from Google directly.
  • Audio in the same render: Veo 3.1 Fast includes audio in Google's price; Kling v3 charges double for its audio mode.
  • 4K output: Veo 3.1 Fast or Standard. Lite does not render 4K.
  • Clips longer than 8 seconds in one call: Veo cannot, so either chain Veo clips with scene extension or test another model.

For a head-to-head of the two most-requested alternatives, read the Kling v3 vs Seedance 2.5 comparison. Most teams do best keeping two models in rotation and routing each shot to the cheapest one that meets its constraint.

Pre-flight checklist before a Veo 3.1 Fast batch

  1. Confirm your project's Veo quota in AI Studio.
  2. Pin the model code in config and set durationSeconds, resolution and aspectRatio on every call.
  3. Draft at 720p. Promote only the winners to 1080p or 4K.
  4. Decide whether you need generated audio; on fal it changes the price by half again.
  5. Multiply the batch cost by your expected reject rate and set a hard budget cap.
  6. Check personGeneration rules for the region you operate in.
  7. Copy every file to your own storage on completion, before the 2-day deletion.
  8. Log prompt, seed, model, resolution, duration and cost per clip so winners can be reproduced.

FAQ

How much does the Veo 3.1 Fast API cost per second?

On Google's Gemini API, at the time of writing, $0.10 per second at 720p, $0.12 at 1080p and $0.30 at 4K, audio included, so an 8-second 720p clip is $0.80. Gateways price it differently.

What is the model name for Veo 3.1 Fast in the Gemini API?

veo-3.1-fast-generate-preview. Vercel AI Gateway uses google/veo-3.1-fast-generate-001 and fal uses fal-ai/veo3.1/fast.

How long can a Veo 3.1 Fast video be?

Each call returns 4, 6 or 8 seconds, and 1080p, 4K and reference images require 8. Longer videos are built by chaining clips or with scene extension.

What are the Veo 3.1 Fast rate limits?

Google does not list Veo in its published rate-limit tables. Your project's limits are shown in AI Studio at aistudio.google.com/rate-limit, and they are the number to plan concurrency around.

Does Veo 3.1 Fast generate audio, and can I remove the watermark?

Veo 3.1 generates native audio, and Google's per-second price includes it. Every output carries a SynthID watermark, and the docs do not document an option to disable it.

Sources

  1. Google AI for Developers: Gemini Developer API pricing
  2. Google AI for Developers: Generate videos with Veo 3.1 in Gemini API
  3. Google AI for Developers: Video generation with Veo
  4. Google AI for Developers: Gemini API rate limits
  5. Google Developers Blog: Introducing Veo 3.1 and new creative capabilities in the Gemini API (October 15, 2025)
  6. fal.ai: Veo 3.1 Fast
  7. Vercel: Veo 3.1 Fast Generate on AI Gateway
  8. Leonardo.Ai: Generate with Veo3.1, Veo3.1 Fast using text prompts
  9. BenchLM.ai: Veo 3.1 API pricing

If Veo 3.1 Fast is one of several models in your pipeline, Aitachyon puts it behind the same key as Kling, Seedance, Hailuo, the image models and ElevenLabs, callable from the studio, the HTTP API or Claude Code over MCP. Each job is itemised with its model and cost, failed renders are refunded to the cent, and the prepaid balance does not expire. The per-second rate is higher than Google's direct price, so if you only ever call Veo, going to Google directly is the cheaper option.

Related articles

Free tools to try

Stop describing your brand. Paste your URL.

Aitachyon reads your whole brand from your website, then creates videos, images, carousels, posts and banners, on-brand, every format, every feed.