Claude Code MCP Video Generation: A Step-by-Step Setup
Connect a hosted MCP server to Claude Code, ask for images, clips and voice in plain words, and get files back with their cost. Setup, prompts, arithmetic.
Claude Code MCP Video Generation: A Step-by-Step Setup
You are in a terminal, halfway through a landing page, and you need three 5-second product clips, a hero image and a 20-second voiceover. The usual route is four browser tabs, three logins, a downloads folder full of files named output(7).mp4, and no clear idea what the afternoon cost. The other route is one sentence typed into Claude Code, with the files landing in your project, each labelled with the model it ran on and what it cost.
The second route runs on the Model Context Protocol. Claude Code speaks MCP natively, so any generation service that exposes a hosted MCP server becomes a set of tools the agent can call. What follows covers the setup, the permission settings worth changing on day one, how to phrase requests so the agent picks the right model and length, and the arithmetic of what real batches cost.
What happens when Claude Code calls a video tool
The transport
A hosted MCP server uses the Streamable HTTP transport. The MCP specification (2025-06-18) defines two standard transports, stdio and Streamable HTTP, and requires the server to expose a single endpoint path that accepts POST and GET, such as https://example.com/mcp. The client sends every JSON-RPC message as a POST, and the server answers with plain JSON or an event stream.
In practice, a hosted server is one URL plus one credential, with nothing to run locally, and it works from Claude Code, Cursor or any other MCP client.
Discovery and calls
Per the MCP tools specification, the client lists what a server offers with tools/list and invokes a tool with tools/call. Each tool carries a name, a description and an input schema. When you ask for "a 5-second clip of a ceramic mug rotating on a walnut table", Claude reads the tool descriptions, fills in the arguments and makes the call. Anthropic describes the trigger rule in its MCP connector docs: Claude calls an MCP tool when the request maps to that tool's described capability, whether or not you name the tool, and it does not call tools to answer general knowledge questions.
What comes back
A tool result can carry text, an image (base64 data plus a MIME type), audio, a resource_link or an embedded resource, and optionally a structuredContent object that follows an output schema. The spec asks servers that return structured content to also include the same data as serialized JSON in a text block. For generation, this is the part that matters. A well-built server returns a link to the file and a small structured record of the job (model, duration, resolution, cost), and keeps the video itself out of the conversation.
Step by step: connecting a hosted MCP server
The Claude Code MCP docs give the general form for a remote server: claude mcp add --transport http <name> <url>, with an optional --header "Authorization: Bearer your-token" for servers that authenticate with a static key.
- Create a key. One key per project or per agent, so you can revoke one without breaking the others.
- Pick a scope. Local is the default and lives in ~/.claude.json for the current project only. Project writes the server into a .mcp.json file meant to be shared through the repo. User makes it available in every project on your machine. Set it with -s or --scope.
- Add the server. For Aitachyon: claude mcp add --transport http aitachyon https://aitachyon.com/api/mcp --header "Authorization: Bearer ait_..."
- Verify. claude mcp list shows every configured server and claude mcp get aitachyon inspects this one. Inside a session, /mcp shows the connection status and the tools exposed.
- Make a cheap first call. Ask for one image before any video. It tests the key and your permission settings for a few cents.
Servers that use OAuth follow a slightly different path: add the server without a header, then sign in with claude mcp login <name> or through /mcp inside Claude Code. Underneath, the MCP authorization spec builds on OAuth 2.1 with PKCE, dynamic client registration (RFC 7591), protected resource metadata (RFC 9728) and RFC 8707 resource indicators. Both paths end with a bearer token on every request.
Three configuration traps
- Project scope and literal keys do not mix. A .mcp.json file is designed to be committed. Put a bearer token in it, push, and the key sits in your git history. Keep keyed servers on local or user scope.
- Secret-looking environment variables read as empty. The Claude Code docs note that variables whose names contain TOKEN, SECRET, PASSWORD, KEY or AUTH read as empty when used in a remote server's URL or headers. A header built from AITACHYON_API_KEY goes out blank and the server rejects the call. Pass the key on the add command.
- Keys stay out of the URL. The authorization spec requires access tokens in the Authorization header on every request and forbids them in the query string, where they would land in proxy and server logs.
Permissions: decide what the agent may spend without asking
A generation tool spends money on every call, so it deserves more thought than a file-read tool. The MCP tools spec says there should always be a human in the loop able to deny tool invocations, and that clients should show tool inputs before calling. Claude Code implements this with permission rules that target MCP tools through the mcp__ prefix.
- mcp__aitachyon in an allow rule lets the whole server run without prompts.
- mcp__aitachyon in a deny rule blocks it for a project, useful in repos where nobody should generate media.
- mcp__* as a deny rule blocks every MCP tool at once.
Two details save debugging time. Claude Code skips any mcp__ rule written with parentheses and lists it in the invalid-settings dialog and in claude doctor. To match a specific parameter value on an MCP tool, pass a deny rule through --disallowedTools.
A decision rule for approvals
- First week with a new server: approve every call by hand and read the model, duration and resolution Claude filled in, since those set the price.
- After ten calls in a row with sensible arguments: allow image tools, keep video tools on approval. Images cost cents; a video batch costs dollars.
- Unattended runs: allow everything, but only with a dedicated key that has its own spend alert and can be revoked in one click.
Asking for media in plain words
Claude picks reasonable defaults, and those defaults are often longer, sharper or louder than the shot needs. Every video request should pin four things: the model (or the trade-off you care about), the duration, the resolution or aspect ratio, and whether you need audio.
Before and after
Before: "Make a video of our mug for the homepage."
Model, length, format and audio are all left to the agent, which may turn on native audio you will mute anyway.
After: "Generate one 5-second 16:9 clip on Kling v3, no audio: a matte white ceramic mug rotating slowly on a walnut table, soft window light from the left, shallow depth of field. Save it to public/media/hero-mug.mp4 and report the ref and the cost."
The second prompt fixes the price before the call. At the time of writing, Kling v3 on Aitachyon costs $0.16 per second silent and $0.32 per second with native audio, so this clip is 5 × $0.16 = $0.80. Adding "with audio" makes it $1.60 for sound that will autoplay muted on most homepages.
A reusable request template
Paste this into your project's CLAUDE.md so every request carries the same constraints:
- Asset: image, clip or voiceover, and how many
- Model: a named model, or "cheapest that supports X"
- Spec: duration in seconds, resolution, aspect ratio (9:16, 16:9, 1:1)
- Audio: none, native, or a separate voiceover
- Content: subject, action, setting, light, camera
- Output: target path and filename pattern
- Budget: "stop and ask if the batch exceeds $X"
- Report: "list each file with its ref, model and cost, then the total"
The last two lines are the ones people skip, and they are what leave the numbers in your terminal history next to the files.
Picking a model per shot, with the trade-offs stated
With one account behind one MCP server, the model becomes a per-shot argument instead of a per-subscription commitment. The prices below are Aitachyon's per-call prices at the time of writing; the full grid is on aitachyon.com/models and as JSON at https://aitachyon.com/api/pricing.
Video
- Hailuo 02, $0.085/s. The cheapest video option on the list. Use it for drafts, hook tests and anything where volume matters more than polish.
- Kling v3, $0.16/s silent, $0.32/s with native audio. Audio doubles the price, so decide per clip. The trade-offs against Seedance are in the Kling v3 vs Seedance 2.5 comparison.
- Veo 3.1 Fast, from $0.19/s. For reference, Google's Gemini API pricing page lists Veo 3.1 Fast at $0.10 per second at 720p, $0.12 at 1080p and $0.30 at 4K, Veo 3.1 Standard at $0.40 per second, and a Lite tier from $0.05. If Veo is the only model you will ever call, going direct to Google is cheaper per second. An aggregator earns its margin with the other models on the same key and the per-job cost record. More in the Veo 3.1 Fast API guide.
- Wan 2.7, $0.19/s. A useful second attempt on a shot another model keeps getting wrong.
- Seedance 2.5, $0.20/s at 480p, $0.44/s at 720p. Resolution more than doubles the price, so draft at 480p and re-render only keepers at 720p.
Images
- Seedream 4.0 at $0.057 per image and FLUX.2 [pro] at $0.057 to $0.086: the low end, fine for product shots and backgrounds.
- Nano Banana, $0.13 per image. Nano Banana vs FLUX.2 Pro for product images covers when the higher price is worth paying.
- gpt-image-2, $0.10 to $0.40 per image. The widest price range, so pin the size in the prompt.
Voice
ElevenLabs voiceover runs $0.043 to $0.086 per 450 characters at the time of writing. The underlying models differ in limits and language coverage: ElevenLabs documents eleven_v3 at 5,000 characters and 70+ languages, eleven_multilingual_v2 at 10,000 characters and 29 languages, and eleven_flash_v2_5 at 40,000 characters and 32 languages. For ad-length scripts the character limit rarely binds, but the language list does once you localise. The per-minute math is in the ElevenLabs API voiceover cost breakdown.
The arithmetic of three real batches
Per-second and per-image pricing makes a batch predictable, provided you multiply before you press enter. All figures use Aitachyon's prices at the time of writing.
Batch 1: 20 hook clips for creative testing
Twenty 5-second openers, each testing a different first line from a list like these hook formulas.
- Hailuo 02: 20 × 5 s × $0.085 = $8.50
- Kling v3, silent: 20 × 5 s × $0.16 = $16.00
- Seedance 2.5 at 720p: 20 × 5 s × $0.44 = $44.00
The sensible sequence: draft all twenty on the cheapest model, keep the three that read best, re-render only those on the model you would ship. Three winners on Seedance at 720p add 3 × $2.20 = $6.60, for a total of $15.10 against $44.00 for rendering everything at the top tier.
Batch 2: 50 product shots
- Seedream 4.0: 50 × $0.057 = $2.85
- FLUX.2 [pro]: 50 × $0.057 to $0.086 = $2.85 to $4.30
- Nano Banana: 50 × $0.13 = $6.50
Batch 3: a week of shorts
Seven shorts, each built from three 5-second Kling v3 clips (silent), three Seedream keyframes and a 900-character voiceover.
- Clips: 15 s × $0.16 = $2.40 per short
- Keyframes: 3 × $0.057 = $0.171 per short
- Voiceover: 2 × 450 characters at up to $0.086 = up to $0.172 per short
- Per short: about $2.74. For the week: about $19.20
The same arithmetic across every model is in AI video generation API pricing in 2026.
Long renders, output limits and getting files back
Output size
Claude Code has a hard budget for what a tool may return. Per the Claude Code MCP docs, it warns when MCP tool output passes 10,000 tokens and caps it at 25,000 tokens by default, adjustable with the MAX_MCP_OUTPUT_TOKENS environment variable (for example 50000). A server that returned raw base64 video would hit that ceiling on the first clip. A link and a short structured record keep each result small and leave the context window for your actual work.
Timeouts
Video takes longer than a typical tool call. The same docs let you set a per-server tool timeout in .mcp.json, in milliseconds with a minimum of 1000, and note that HTTP and SSE connections idle out after 5 minutes by default. If a long batch dies partway, check these before blaming the model. Ask for renders in small groups and have the agent report after each group, so a dropped connection only loses the status of the group in flight.
Stable refs as a handoff
On Aitachyon every generated file gets a stable ref (img_, scn_, vo_ and so on) that code can fetch with GET /api/generations/{ref}. That gives a clean line between the agent and your build:
- In Claude Code, generate the asset and ask for the ref in the report.
- Record the ref in a manifest file in the repo, next to the prompt that produced it and its cost.
- A build script or CI step fetches each ref over the plain HTTP API with the same bearer key.
- When you regenerate a shot, change one ref in the manifest and the pipeline picks it up.
The manifest doubles as a cost ledger per shot. The longer version of this pipeline, with scripting on top, is in automating video production with AI agents.
Running it without Claude Code
The same hosted server can be called from your own backend. Anthropic's MCP connector lets the Messages API reach remote MCP servers without a separate MCP client. It is in beta under the header mcp-client-2025-11-20, supports allowlisting and denylisting tools, accepts OAuth bearer tokens and several servers per request, and is not eligible for zero data retention. The tool allowlist is the useful part here: expose image tools to a customer-facing feature and keep video tools off until you have priced them in.
Spend control when an agent holds the key
An agent that can call a paid tool in a loop can also spend money in a loop. A misread instruction, "make variations" taken as fifty instead of five, is a realistic failure. The controls that help, ordered by how early they catch it:
- Budget in the prompt. "Stop and ask if the batch will cost more than $10" is a rule the agent can check against per-second prices.
- Approval on video tools, as set in the permissions section.
- One key per agent or project. Aitachyon tracks spend per API key, raises an alert when a key spends unusually fast, and revokes a key in one click.
- Prepaid balance. When it runs out, calls stop. That is a ceiling a runaway loop cannot cross.
- Automatic refunds on failure. A failed render is refunded to the cent, so retries do not quietly double-charge you.
Anthropic adds the baseline in its note on remote MCP servers: they are third-party services, so connect only to servers you trust and review each one's security practices and terms.
FAQ
How do I add an MCP server to Claude Code?
Run claude mcp add --transport http <name> <url>, adding --header "Authorization: Bearer <key>" for key-based servers. Use -s to choose local, project or user scope, then confirm with claude mcp list or /mcp inside a session. For OAuth servers, add without the header and sign in with claude mcp login <name>.
Can Claude Code generate video on its own?
Claude Code hands the rendering to a tool on a connected MCP server, which runs the video model. Once such a server is connected, you ask in plain words and Claude fills in the tool arguments, makes the call and hands back the result.
Why does my MCP server reject the key I put in an environment variable?
Claude Code reads environment variables whose names contain TOKEN, SECRET, PASSWORD, KEY or AUTH as empty when they appear in a remote server's URL or headers, so the header goes out blank. Pass the key directly on the claude mcp add command and keep the server on local or user scope.
How much does a 5-second AI video clip cost?
It depends on the model and settings. At Aitachyon's prices at the time of writing, 5 seconds costs about $0.43 on Hailuo 02, $0.80 on Kling v3 silent, $1.60 with native audio, and $1.00 or $2.20 on Seedance 2.5 at 480p or 720p.
Sources
- Anthropic (Claude Code docs): Connect Claude Code to tools via MCP
- Anthropic (Claude Code docs): Configure permissions
- Model Context Protocol: Specification 2025-06-18, Transports
- Model Context Protocol: Specification 2025-06-18, Tools
- Model Context Protocol: Specification 2025-06-18, Authorization
- Anthropic (Claude Platform docs): MCP connector
- Anthropic (Claude Platform docs): Remote MCP servers
- Google AI for Developers: Gemini API pricing
- ElevenLabs: Text-to-speech models
If you want this setup without five separate accounts, Aitachyon puts the video, image and voice models above behind one key, one hosted MCP server and one prepaid balance, with every job itemised by model and cost. The one-line Claude Code command and the HTTP API are on the developers page.
Related articles
Video Ad Approval: The Meta and TikTok Policy Traps That Get You Rejected
A media buyer's guide to video ad approval—the exact phrases, visuals, and gated categories that get Meta and TikTok ads rejected, and how to reframe.
TutorialsHow to Make AI Ads That Don't Look AI (500M-Impression Data)
AI ads that don't look AI beat both human and obvious-AI creative across 500M impressions. The exact tells to remove, ranked by what gives you away.
TutorialsAI UGC Ads: How to Make Them Look Like Real User Content
How to script, cast, and post-process AI UGC ads so they read as organic creator footage instead of getting clocked as fake on paid social.
TutorialsFrom URL to finished video ad: the two-minute pipeline
Paste a URL, get a finished video ad in about two minutes. A step-by-step look at the pipeline: brand scrape, three scripts, voice, avatar, captions.
TutorialsTestimonial Video Ads With No Customers Yet: 4 Honest Ways
Make testimonial video ads at launch with beta quotes, founder narration, and demo proof, without fabricating reviews the FTC now bans.
TutorialsA/B Testing Video Ads: Test One Variable and Reach Significance on a Small Budget
A/B testing video ads on a small budget: isolate one variable, hit real sample thresholds, and read significance without fooling yourself.
Free tools to try
Free AI image generator
Describe what you want and get a high-quality AI image in seconds. A free AI image generator, no account needed to preview, keep your first image when you sign up.
Try it freeFree toolFree AI product photo generator
Generate clean, studio-style product photos for your store and listings in seconds. Crisp lighting and tidy backgrounds, free to try with no account, keep your first shot on signup.
Try it freeFree toolFree background remover
Remove the background from any image in seconds and get a clean, transparent cutout. A free background remover, no account needed to preview, keep your first cutout when you sign up.
Try it freeStop describing your brand. Paste your URL.
Aitachyon reads your whole brand from your website, then creates videos, images, carousels, posts and banners, on-brand, every format, every feed.