AI Avatar Video API: Talking Heads From a Photo and a Voice
How photo-plus-audio avatar APIs work, what they cost per second across providers, where they break, and how to call one from code or an AI agent.
AI Avatar Video API: Talking Heads From a Photo and a Voice
A 30-second talking-head clip from one portrait and one voice track costs somewhere between $0.75 and $6.40 through an API, depending on which provider and which resolution you pick. That spread is wider than most people expect, and it comes from a single variable: the per-second rate. The model, the inputs and the request shape are close to identical across vendors.
This guide covers how these models turn a still photo into a speaking face, what each provider charges per second of output, where the results fall apart, and how to wire the whole thing into a script or an agent so that a brief goes in and finished MP4s come out with their cost attached.
How a photo-plus-audio avatar model works
Every provider in this category follows the same contract. You send a portrait and a speech track, the model animates the face to match the audio, and you get back a video whose length equals the audio. Replicate's Kling Avatar V2 page states it plainly: the duration automatically matches the audio length. You never pick a duration. You pick a voice track, and the voice track sets the bill.
APIframe's avatar API guide breaks the process into three stages:
- Identity extraction. The model reads the face from the portrait and locks it, so the person in frame one is the person in the last frame.
- Audio mapping. The audio is mapped frame by frame to mouth shapes and facial expressions. This is where lip sync quality is won or lost.
- Asynchronous rendering. The video is rendered as a background job. You submit, you wait, you download.
Most models also accept an optional text prompt to steer expression or motion. On Kie.ai's Kling AI Avatar 2.0 page the prompt field takes 0 to 5,000 characters and is optional for basic use. In practice, the portrait and the audio do almost all the work.
The subject does not have to be a real person. fal's Kling Avatar v2 Pro listing says it handles realistic humans, animals, cartoons and stylized characters, and VEED describes its Fabric model as accepting photos, illustrations, mascots, 3D renders and anime. That matters for brands that want a mascot or a generated presenter instead of a filmed founder. If you are still deciding whether a talking head is the right format at all, our breakdown of when avatar ads work and when they don't covers the creative side.
What an avatar video API costs per second
Almost every provider bills per second of output video. APIframe puts the typical range at $0.02 to $0.07 per second at standard resolutions. Published rates at the time of writing, from provider pages and third-party guides:
- P-Video-Avatar: $0.025/s at 720p and $0.045/s at 1080p (inference.sh, 2026).
- Kling Avatar v2 Standard on WaveSpeed: $0.056 per second of audio, so 5 seconds costs $0.28.
- VEED Fabric 1.0: $0.08/s at 480p and $0.15/s at 720p, served through fal.ai.
- Kling Avatar v2 Pro on fal: $0.115 per second of output, about $6.90 per minute.
- Pixverse Avatar: $0.053333/s at 360p, $0.106667/s at 540p, $0.16/s at 720p and $0.213333/s at 1080p (Empirio Labs catalog).
Two caveats. These figures come from vendor pages and guides and have not been checked against real invoices, so run a handful of test renders before you commit a budget. And the same underlying model can cost very different amounts depending on who serves it: Kling Avatar v2 appears at $0.056/s (Standard, WaveSpeed) and $0.115/s (Pro, fal), while Kie.ai uses credit-based billing with no per-second rate on the page and Replicate's model page showed no price when we checked.
The cost formula
Because output length follows the audio, the math for one clip is short:
clip cost = audio seconds x avatar rate + voice cost + portrait cost
The portrait is a one-time cost if you reuse the same face. The voice is cheap next to the avatar. On Aitachyon, ElevenLabs voiceover costs $0.043 to $0.086 per 450 characters at the time of writing, and a generated portrait on Nano Banana is $0.13 per image. The avatar seconds dominate every line of the budget.
Three batches, worked out
Using the published rates above:
- 20 avatar hooks of 10 seconds each (200 seconds of video). On Kling Avatar Standard at $0.056/s: $11.20. On Kling Avatar Pro at $0.115/s: $23.00. Add 20 short voice lines under 450 characters each ($0.86 to $1.72) and one reused portrait ($0.13), and the batch lands between roughly $12.19 and $24.85.
- A week of shorts, seven 45-second talking heads (315 seconds). Standard: $17.64. Pro: $36.23. P-Video-Avatar at 720p: $7.88. Assume each script runs to two 450-character blocks of voice, which adds 14 blocks at $0.60 to $1.20.
- One 60-second explainer at 1080p. P-Video-Avatar: $2.70. Kling Pro: $6.90. Pixverse: $12.80.
The same week of content can cost $8 or $37 before a single creative decision is made. If you are pricing video models more broadly, our model-by-model video API pricing breakdown uses the same per-second method.
Where avatar models break
Duration
The hard caps are generous. WaveSpeed bills Kling Avatar Standard up to a maximum of 300 seconds per job, Kie.ai states that audio cannot exceed 5 minutes, and VEED lists Fabric at 5 minutes per generation.
The quality ceiling arrives much earlier. APIframe warns that some models show identity drift and warping beyond 30 seconds. The P-Video-Avatar guide recommends keeping clips under three minutes and splitting longer content. A five-minute cap tells you what the API will accept. It says little about what will still look like the same person at minute four.
Lip sync and expression
Lip sync is only as good as the audio mapping stage, and that stage is only as good as the audio. WaveSpeed asks for a clean voice track, recorded or TTS, with long silences trimmed. Silence costs money on a per-second model and gives the face nothing to do, which is where uncanny idle frames tend to appear.
Gestures and body
Most of these models animate a face and head from a single image. VEED describes Fabric as producing natural head gestures and expressive body language, but that is a vendor description. Plan for a head-and-shoulders shot and treat hand gestures as a bonus.
Framing and resolution
The P-Video-Avatar guide notes that the aspect ratio follows the input image. If you want a 9:16 short, generate or crop a 9:16 portrait first. Resolution depends on tier: Kie.ai lists up to 720p on Standard and up to 1080p at 48fps on Pro.
Rejected jobs
APIframe lists the common failure causes as poor source images, unsupported audio formats and content policy rejections. All three are preventable before you spend a cent, which is what the checklist below is for.
The input checklist: run it before every job
Input quality drives output quality more than model choice does. Copy this and run it as a validation step in your script:
- Portrait angle. Front-facing or a slight three-quarter view, as WaveSpeed recommends. No profile shots, no face partly out of frame.
- Portrait lighting. Even, well-lit face. Hard shadows across the mouth give the model less to map.
- Portrait aspect ratio. Match the final placement (9:16 for shorts, 1:1 or 16:9 elsewhere), since output framing follows the image.
- Image file. JPEG or PNG under 10MB for Kie.ai; PNG, JPEG or WebP under 20MB and under 10,000 px on the long edge for Pixverse. fal accepts JPG, JPEG, PNG, WebP, GIF and AVIF.
- Audio format. MP3, OGG, WAV, M4A or AAC is safe across fal and Pixverse. Kie.ai accepts MPEG, WAV, AAC, MP4 and OGG up to 100MB.
- Audio cleanliness. One speaker, no music bed, no room echo. Add music after the render.
- Silence trimmed. Cut leading and trailing silence and long pauses. Every second is billed.
- Length. Keep each segment under 30 seconds unless you have tested your model past that point. Split longer scripts at sentence boundaries and render the segments separately.
- Policy. No real person's likeness without consent, and nothing the provider's content policy would reject. A rejected job wastes wall-clock time even when it costs nothing.
If you generate the portrait instead of photographing one, choose the image model for skin texture and consistency. Our Nano Banana vs FLUX.2 Pro comparison covers that trade-off for product and people shots.
Choosing a provider: a decision rule
Pick with three questions, in this order:
- How long is the finished segment? Under 30 seconds, every option is on the table. Longer than that, split it, or test the specific model for drift before you run a batch.
- What resolution does the placement need? Feed placements viewed on a phone rarely need 1080p. If you do need it, the cheapest published 1080p rate in our research is P-Video-Avatar at $0.045/s, and the most expensive is Pixverse at $0.213333/s.
- How much wall-clock time can you wait? WaveSpeed lists a typical generation time of about 177 seconds for Kling Avatar Standard. inference.sh claims P-Video-Avatar processes at about 1.83 seconds per second of output, against 26 s/s for HeyGen and 34 s/s for Veed Fabric. That comparison is a vendor claim about its own product, so measure it yourself.
A reasonable default: prototype on the cheapest tier that meets your resolution, render five clips with your real portrait and voice, compare them side by side, and only then move up a tier if the cheaper output fails on a specific, nameable problem (mouth artefacts, drift, stiff head).
API or seat-based tool
Seat-based avatar tools price by month instead of by second. According to VEED's 2026 roundup: HeyGen starts at $29/month on credits with lip sync across 40+ languages, D-ID Lite starts at $5.90/month with an API tier from $18/month that includes 16 minutes of regular video, Synthesia starts at $18/month for 120 minutes per year, and Tavus uses custom enterprise pricing with a digital twin built from about 2 minutes of footage. If you use all 16 D-ID minutes, $18 works out to about $1.13 per minute. The trade-off is that unused minutes go to waste and overage has its own rules. Per-second APIs cost more per unit at the top of the range and nothing in a month you do not render.
Calling an avatar API from code
The request flow is the same across providers. APIframe describes it as a POST that returns a job ID, then polling or a webhook, then downloading the video. Concretely, for Kling Avatar Standard on WaveSpeed:
- Submit. POST to https://api.wavespeed.ai/api/v3/kwaivgi/kling-v2-ai-avatar-standard with a Bearer token, an image URL and an audio URL. The response carries a prediction ID to poll.
- Poll. WaveSpeed's docs start polling every 2 seconds. Back off after the first minute, since typical jobs take a few minutes.
- Download. On fal the result is JSON with a video.url field pointing to an MP4 with synced audio. Copy the file to your own storage right away; provider URLs are not a place to keep assets.
- Record the cost. Multiply audio seconds by the rate and store it next to the file. You will want this number when a batch of 200 clips comes back.
Fabric is served over REST via fal.ai with Python and JavaScript clients, so the same four steps apply with a different endpoint name.
Upstream steps: portrait and voice
The avatar call needs two URLs, and in an automated pipeline both are usually generated. On Aitachyon, one API key (Authorization: Bearer ait_...) covers the image and voice models, every generated file gets a stable ref (img_ for images, vo_ for voiceovers) that you can fetch with GET /api/generations/{ref}, and every job is itemised with the model it ran on and what it cost. The API and MCP quick start has the request shapes. For voice pricing and pacing in detail, see our ElevenLabs API cost breakdown.
Running it from an agent
For developers who already work in Claude Code or Cursor, the fastest loop is to let the agent handle the upstream assets and the bookkeeping while a small script handles the avatar provider. Connecting Aitachyon's hosted MCP server to Claude Code is one command:
claude mcp add --transport http aitachyon https://aitachyon.com/api/mcp --header "Authorization: Bearer ait_..."
A working session then looks like this:
- You give the agent a brief: "Twenty 10-second hooks for the pricing page, same presenter, 9:16."
- The agent writes twenty scripts, trims each to one 450-character voice block, and generates one 9:16 portrait and twenty voice tracks through the MCP server. Each comes back with a ref and a cost.
- The agent runs your avatar script against the provider of your choice, passing the portrait and audio URLs, polling, and saving the MP4s.
- It writes a manifest: script, voice ref, portrait ref, avatar file, seconds, cost per line, and the batch total.
The setup is covered step by step in our guide to generating video from Claude Code with an MCP server. Two guardrails matter once an agent is spending money in a loop: cap the batch size in the brief, and give the agent its own key so its spend is isolated. On Aitachyon each key has its own spend, an alert when it spends unusually fast, and a one-click revoke. For concurrency limits, retries and file handling at larger volume, our guide to bulk AI video generation goes further.
The hybrid pattern: short avatar, then B-roll
The single most effective cost and quality lever is to keep the talking head short. A 5 to 10 second avatar hook, followed by B-roll under the same voiceover, stays well inside the 30-second zone where drift is less likely, and moves most of the runtime to cheaper footage.
Worked example, 30 seconds total at the time of writing:
- Full avatar, 30 s on Kling Avatar Standard at $0.056/s: $1.68. On Kling Pro at $0.115/s: $3.45.
- Hybrid: 8 s of avatar on Kling Standard ($0.45) plus 22 s of B-roll on Hailuo 02 at $0.085/s ($1.87). Total about $2.32.
On the cheapest avatar tier, the hybrid costs more than a full avatar. Against the Pro tier it costs less, and in both cases the viewer sees the face only for the hook, where it earns attention, while the product is on screen for the rest. B-roll that still reads as real footage takes some care, which our guide to AI B-roll that doesn't look fake covers shot by shot.
FAQ
How much does an AI avatar video API cost per minute?
From published rates at the time of writing: about $1.50 per minute on P-Video-Avatar at 720p, $3.36 on Kling Avatar Standard via WaveSpeed, $6.90 on Kling Avatar Pro via fal, $9.00 on VEED Fabric at 720p, and $12.80 on Pixverse at 1080p. Add voice and portrait costs, which are small by comparison.
What is the maximum length of an AI avatar video?
Kling Avatar and VEED Fabric accept up to 5 minutes per job. Quality is the tighter limit: some models drift or warp beyond 30 seconds, and P-Video-Avatar's guide recommends staying under three minutes. Split long scripts into segments and render each one separately.
Can I make a talking avatar from a single photo?
Yes. A single front-facing or slight three-quarter portrait plus an audio file is all these APIs need. Illustrations, mascots, animals and stylized characters also work on Kling Avatar v2 and VEED Fabric.
Do I need to record my own voice?
No. Every provider in this guide accepts a TTS track, and P-Video-Avatar and Pixverse can take a text script directly. Clean audio with silences trimmed gives the best lip sync regardless of source.
Is an avatar API cheaper than HeyGen or Synthesia?
It depends on volume. Seat plans charge monthly whether or not you render. Per-second APIs charge only for what you produce, so they win for irregular or bursty volume and for anything automated from code.
Sources
- WaveSpeedAI: Kling V2 AI Avatar Standard API
- fal.ai: Kling AI Avatar v2 Pro
- Kie.ai: Kling AI Avatar 2.0 API
- Replicate: Kling Avatar V2
- VEED: Best Avatar APIs (2026), Build Talking Videos at Scale
- APIframe: AI Avatar API Guide
- inference.sh: P-Video-Avatar, the Fastest AI Talking Head Generator
- Empirio Labs: Pixverse Avatar API
Aitachyon covers the inputs around the avatar call: portraits on Nano Banana or FLUX.2 [pro], voice on ElevenLabs, and B-roll on Seedance, Kling, Veo, Wan or Hailuo, all from one prepaid balance that never expires, billed per call, with failed renders refunded automatically and every file itemised with its cost. You can drive it from the web studio, the HTTP API or the MCP server in Claude Code, and the full price grid is public at /api/pricing.
Related articles
Seedance 2.5 API Pricing: What a Clip Really Costs
Seedance 2.5 API pricing per second at 480p, 720p and 1080p, worked costs for 5s, 10s and 30s clips and batches, and where to get access with no subscription.
GuidesAI Video Generation API Pricing in 2026, Model by Model
How AI video APIs bill (per second, per clip, credits), what drives the price, a model-by-model rate list and worked batch costs with retries included.
GuidesVeo 3.1 Fast API: Price, Limits and How to Call It
Veo 3.1 Fast API pricing per second, clip lengths, resolutions, latency and quotas from Google's docs, plus how to call it from code or from Claude Code.
GuidesElevenLabs API Pricing: What a Voiceover Really Costs (2026)
ElevenLabs API pricing per 1,000 characters by model, what a 30s or 60s voiceover costs, and how to batch-generate voiceovers from code without 429 errors.
GuidesHow to Run a Faceless YouTube Channel With AI in 2026
The real pipeline for a faceless YouTube channel with AI: script, voice, visuals, edit, what each stage costs per video, and how YouTube's AI rules affect monetization.
GuidesFaceless Channel Cost: What Each AI Video Costs to Make
A line-by-line faceless channel cost model for 30, 60 and 600-second AI videos, with live model prices, the arithmetic shown, and editor and voice actor rates
Free tools to try
Free AI image generator
Describe what you want and get a high-quality AI image in seconds. A free AI image generator, no account needed to preview, keep your first image when you sign up.
Try it freeFree toolFree AI avatar generator
Create a unique AI avatar for your profile in seconds, from realistic portraits to stylized and anime looks. Free to try with no account, keep your first avatar when you sign up.
Try it freeFree toolFree AI product photo generator
Generate clean, studio-style product photos for your store and listings in seconds. Crisp lighting and tidy backgrounds, free to try with no account, keep your first shot on signup.
Try it freeStop describing your brand. Paste your URL.
Aitachyon reads your whole brand from your website, then creates videos, images, carousels, posts and banners, on-brand, every format, every feed.