Skip to content
GuidesSeptember 30, 2026· 9 min read

AI Avatar Video API: Talking Heads From a Photo and a Voice

How photo-plus-audio avatar APIs work, what they cost per second across providers, where they break, and how to call one from code or an AI agent.

avatarsapilip synctalking head
Guides

AI Avatar Video API: Talking Heads From a Photo and a Voice

A 30-second talking-head clip from one portrait and one voice track costs somewhere between $0.75 and $6.40 through an API, depending on which provider and which resolution you pick. That spread is wider than most people expect, and it comes from a single variable: the per-second rate. The model, the inputs and the request shape are close to identical across vendors.

This guide covers how these models turn a still photo into a speaking face, what each provider charges per second of output, where the results fall apart, and how to wire the whole thing into a script or an agent so that a brief goes in and finished MP4s come out with their cost attached.

How a photo-plus-audio avatar model works

Every provider in this category follows the same contract. You send a portrait and a speech track, the model animates the face to match the audio, and you get back a video whose length equals the audio. Replicate's Kling Avatar V2 page states it plainly: the duration automatically matches the audio length. You never pick a duration. You pick a voice track, and the voice track sets the bill.

APIframe's avatar API guide breaks the process into three stages:

  1. Identity extraction. The model reads the face from the portrait and locks it, so the person in frame one is the person in the last frame.
  2. Audio mapping. The audio is mapped frame by frame to mouth shapes and facial expressions. This is where lip sync quality is won or lost.
  3. Asynchronous rendering. The video is rendered as a background job. You submit, you wait, you download.

Most models also accept an optional text prompt to steer expression or motion. On Kie.ai's Kling AI Avatar 2.0 page the prompt field takes 0 to 5,000 characters and is optional for basic use. In practice, the portrait and the audio do almost all the work.

The subject does not have to be a real person. fal's Kling Avatar v2 Pro listing says it handles realistic humans, animals, cartoons and stylized characters, and VEED describes its Fabric model as accepting photos, illustrations, mascots, 3D renders and anime. That matters for brands that want a mascot or a generated presenter instead of a filmed founder. If you are still deciding whether a talking head is the right format at all, our breakdown of when avatar ads work and when they don't covers the creative side.

What an avatar video API costs per second

Almost every provider bills per second of output video. APIframe puts the typical range at $0.02 to $0.07 per second at standard resolutions. Published rates at the time of writing, from provider pages and third-party guides:

Two caveats. These figures come from vendor pages and guides and have not been checked against real invoices, so run a handful of test renders before you commit a budget. And the same underlying model can cost very different amounts depending on who serves it: Kling Avatar v2 appears at $0.056/s (Standard, WaveSpeed) and $0.115/s (Pro, fal), while Kie.ai uses credit-based billing with no per-second rate on the page and Replicate's model page showed no price when we checked.

The cost formula

Because output length follows the audio, the math for one clip is short:

clip cost = audio seconds x avatar rate + voice cost + portrait cost

The portrait is a one-time cost if you reuse the same face. The voice is cheap next to the avatar. On Aitachyon, ElevenLabs voiceover costs $0.043 to $0.086 per 450 characters at the time of writing, and a generated portrait on Nano Banana is $0.13 per image. The avatar seconds dominate every line of the budget.

Three batches, worked out

Using the published rates above:

  1. 20 avatar hooks of 10 seconds each (200 seconds of video). On Kling Avatar Standard at $0.056/s: $11.20. On Kling Avatar Pro at $0.115/s: $23.00. Add 20 short voice lines under 450 characters each ($0.86 to $1.72) and one reused portrait ($0.13), and the batch lands between roughly $12.19 and $24.85.
  2. A week of shorts, seven 45-second talking heads (315 seconds). Standard: $17.64. Pro: $36.23. P-Video-Avatar at 720p: $7.88. Assume each script runs to two 450-character blocks of voice, which adds 14 blocks at $0.60 to $1.20.
  3. One 60-second explainer at 1080p. P-Video-Avatar: $2.70. Kling Pro: $6.90. Pixverse: $12.80.

The same week of content can cost $8 or $37 before a single creative decision is made. If you are pricing video models more broadly, our model-by-model video API pricing breakdown uses the same per-second method.

Where avatar models break

Duration

The hard caps are generous. WaveSpeed bills Kling Avatar Standard up to a maximum of 300 seconds per job, Kie.ai states that audio cannot exceed 5 minutes, and VEED lists Fabric at 5 minutes per generation.

The quality ceiling arrives much earlier. APIframe warns that some models show identity drift and warping beyond 30 seconds. The P-Video-Avatar guide recommends keeping clips under three minutes and splitting longer content. A five-minute cap tells you what the API will accept. It says little about what will still look like the same person at minute four.

Lip sync and expression

Lip sync is only as good as the audio mapping stage, and that stage is only as good as the audio. WaveSpeed asks for a clean voice track, recorded or TTS, with long silences trimmed. Silence costs money on a per-second model and gives the face nothing to do, which is where uncanny idle frames tend to appear.

Gestures and body

Most of these models animate a face and head from a single image. VEED describes Fabric as producing natural head gestures and expressive body language, but that is a vendor description. Plan for a head-and-shoulders shot and treat hand gestures as a bonus.

Framing and resolution

The P-Video-Avatar guide notes that the aspect ratio follows the input image. If you want a 9:16 short, generate or crop a 9:16 portrait first. Resolution depends on tier: Kie.ai lists up to 720p on Standard and up to 1080p at 48fps on Pro.

Rejected jobs

APIframe lists the common failure causes as poor source images, unsupported audio formats and content policy rejections. All three are preventable before you spend a cent, which is what the checklist below is for.

The input checklist: run it before every job

Input quality drives output quality more than model choice does. Copy this and run it as a validation step in your script:

  1. Portrait angle. Front-facing or a slight three-quarter view, as WaveSpeed recommends. No profile shots, no face partly out of frame.
  2. Portrait lighting. Even, well-lit face. Hard shadows across the mouth give the model less to map.
  3. Portrait aspect ratio. Match the final placement (9:16 for shorts, 1:1 or 16:9 elsewhere), since output framing follows the image.
  4. Image file. JPEG or PNG under 10MB for Kie.ai; PNG, JPEG or WebP under 20MB and under 10,000 px on the long edge for Pixverse. fal accepts JPG, JPEG, PNG, WebP, GIF and AVIF.
  5. Audio format. MP3, OGG, WAV, M4A or AAC is safe across fal and Pixverse. Kie.ai accepts MPEG, WAV, AAC, MP4 and OGG up to 100MB.
  6. Audio cleanliness. One speaker, no music bed, no room echo. Add music after the render.
  7. Silence trimmed. Cut leading and trailing silence and long pauses. Every second is billed.
  8. Length. Keep each segment under 30 seconds unless you have tested your model past that point. Split longer scripts at sentence boundaries and render the segments separately.
  9. Policy. No real person's likeness without consent, and nothing the provider's content policy would reject. A rejected job wastes wall-clock time even when it costs nothing.

If you generate the portrait instead of photographing one, choose the image model for skin texture and consistency. Our Nano Banana vs FLUX.2 Pro comparison covers that trade-off for product and people shots.

Choosing a provider: a decision rule

Pick with three questions, in this order:

  1. How long is the finished segment? Under 30 seconds, every option is on the table. Longer than that, split it, or test the specific model for drift before you run a batch.
  2. What resolution does the placement need? Feed placements viewed on a phone rarely need 1080p. If you do need it, the cheapest published 1080p rate in our research is P-Video-Avatar at $0.045/s, and the most expensive is Pixverse at $0.213333/s.
  3. How much wall-clock time can you wait? WaveSpeed lists a typical generation time of about 177 seconds for Kling Avatar Standard. inference.sh claims P-Video-Avatar processes at about 1.83 seconds per second of output, against 26 s/s for HeyGen and 34 s/s for Veed Fabric. That comparison is a vendor claim about its own product, so measure it yourself.

A reasonable default: prototype on the cheapest tier that meets your resolution, render five clips with your real portrait and voice, compare them side by side, and only then move up a tier if the cheaper output fails on a specific, nameable problem (mouth artefacts, drift, stiff head).

API or seat-based tool

Seat-based avatar tools price by month instead of by second. According to VEED's 2026 roundup: HeyGen starts at $29/month on credits with lip sync across 40+ languages, D-ID Lite starts at $5.90/month with an API tier from $18/month that includes 16 minutes of regular video, Synthesia starts at $18/month for 120 minutes per year, and Tavus uses custom enterprise pricing with a digital twin built from about 2 minutes of footage. If you use all 16 D-ID minutes, $18 works out to about $1.13 per minute. The trade-off is that unused minutes go to waste and overage has its own rules. Per-second APIs cost more per unit at the top of the range and nothing in a month you do not render.

Calling an avatar API from code

The request flow is the same across providers. APIframe describes it as a POST that returns a job ID, then polling or a webhook, then downloading the video. Concretely, for Kling Avatar Standard on WaveSpeed:

  1. Submit. POST to https://api.wavespeed.ai/api/v3/kwaivgi/kling-v2-ai-avatar-standard with a Bearer token, an image URL and an audio URL. The response carries a prediction ID to poll.
  2. Poll. WaveSpeed's docs start polling every 2 seconds. Back off after the first minute, since typical jobs take a few minutes.
  3. Download. On fal the result is JSON with a video.url field pointing to an MP4 with synced audio. Copy the file to your own storage right away; provider URLs are not a place to keep assets.
  4. Record the cost. Multiply audio seconds by the rate and store it next to the file. You will want this number when a batch of 200 clips comes back.

Fabric is served over REST via fal.ai with Python and JavaScript clients, so the same four steps apply with a different endpoint name.

Upstream steps: portrait and voice

The avatar call needs two URLs, and in an automated pipeline both are usually generated. On Aitachyon, one API key (Authorization: Bearer ait_...) covers the image and voice models, every generated file gets a stable ref (img_ for images, vo_ for voiceovers) that you can fetch with GET /api/generations/{ref}, and every job is itemised with the model it ran on and what it cost. The API and MCP quick start has the request shapes. For voice pricing and pacing in detail, see our ElevenLabs API cost breakdown.

Running it from an agent

For developers who already work in Claude Code or Cursor, the fastest loop is to let the agent handle the upstream assets and the bookkeeping while a small script handles the avatar provider. Connecting Aitachyon's hosted MCP server to Claude Code is one command:

claude mcp add --transport http aitachyon https://aitachyon.com/api/mcp --header "Authorization: Bearer ait_..."

A working session then looks like this:

  1. You give the agent a brief: "Twenty 10-second hooks for the pricing page, same presenter, 9:16."
  2. The agent writes twenty scripts, trims each to one 450-character voice block, and generates one 9:16 portrait and twenty voice tracks through the MCP server. Each comes back with a ref and a cost.
  3. The agent runs your avatar script against the provider of your choice, passing the portrait and audio URLs, polling, and saving the MP4s.
  4. It writes a manifest: script, voice ref, portrait ref, avatar file, seconds, cost per line, and the batch total.

The setup is covered step by step in our guide to generating video from Claude Code with an MCP server. Two guardrails matter once an agent is spending money in a loop: cap the batch size in the brief, and give the agent its own key so its spend is isolated. On Aitachyon each key has its own spend, an alert when it spends unusually fast, and a one-click revoke. For concurrency limits, retries and file handling at larger volume, our guide to bulk AI video generation goes further.

The hybrid pattern: short avatar, then B-roll

The single most effective cost and quality lever is to keep the talking head short. A 5 to 10 second avatar hook, followed by B-roll under the same voiceover, stays well inside the 30-second zone where drift is less likely, and moves most of the runtime to cheaper footage.

Worked example, 30 seconds total at the time of writing:

  • Full avatar, 30 s on Kling Avatar Standard at $0.056/s: $1.68. On Kling Pro at $0.115/s: $3.45.
  • Hybrid: 8 s of avatar on Kling Standard ($0.45) plus 22 s of B-roll on Hailuo 02 at $0.085/s ($1.87). Total about $2.32.

On the cheapest avatar tier, the hybrid costs more than a full avatar. Against the Pro tier it costs less, and in both cases the viewer sees the face only for the hook, where it earns attention, while the product is on screen for the rest. B-roll that still reads as real footage takes some care, which our guide to AI B-roll that doesn't look fake covers shot by shot.

FAQ

How much does an AI avatar video API cost per minute?

From published rates at the time of writing: about $1.50 per minute on P-Video-Avatar at 720p, $3.36 on Kling Avatar Standard via WaveSpeed, $6.90 on Kling Avatar Pro via fal, $9.00 on VEED Fabric at 720p, and $12.80 on Pixverse at 1080p. Add voice and portrait costs, which are small by comparison.

What is the maximum length of an AI avatar video?

Kling Avatar and VEED Fabric accept up to 5 minutes per job. Quality is the tighter limit: some models drift or warp beyond 30 seconds, and P-Video-Avatar's guide recommends staying under three minutes. Split long scripts into segments and render each one separately.

Can I make a talking avatar from a single photo?

Yes. A single front-facing or slight three-quarter portrait plus an audio file is all these APIs need. Illustrations, mascots, animals and stylized characters also work on Kling Avatar v2 and VEED Fabric.

Do I need to record my own voice?

No. Every provider in this guide accepts a TTS track, and P-Video-Avatar and Pixverse can take a text script directly. Clean audio with silences trimmed gives the best lip sync regardless of source.

Is an avatar API cheaper than HeyGen or Synthesia?

It depends on volume. Seat plans charge monthly whether or not you render. Per-second APIs charge only for what you produce, so they win for irregular or bursty volume and for anything automated from code.

Sources

  1. WaveSpeedAI: Kling V2 AI Avatar Standard API
  2. fal.ai: Kling AI Avatar v2 Pro
  3. Kie.ai: Kling AI Avatar 2.0 API
  4. Replicate: Kling Avatar V2
  5. VEED: Best Avatar APIs (2026), Build Talking Videos at Scale
  6. APIframe: AI Avatar API Guide
  7. inference.sh: P-Video-Avatar, the Fastest AI Talking Head Generator
  8. Empirio Labs: Pixverse Avatar API

Aitachyon covers the inputs around the avatar call: portraits on Nano Banana or FLUX.2 [pro], voice on ElevenLabs, and B-roll on Seedance, Kling, Veo, Wan or Hailuo, all from one prepaid balance that never expires, billed per call, with failed renders refunded automatically and every file itemised with its cost. You can drive it from the web studio, the HTTP API or the MCP server in Claude Code, and the full price grid is public at /api/pricing.

Related articles

Free tools to try

Stop describing your brand. Paste your URL.

Aitachyon reads your whole brand from your website, then creates videos, images, carousels, posts and banners, on-brand, every format, every feed.