Cover image for How to Make a Promo Video with Claude Code, Free
Claude Code
AI Video
Kokoro TTS
FFmpeg
Marketing
Prompt Engineering

How to Make a Promo Video with Claude Code, Free

Shihab Shahriar Antor
12 min read

TL;DR

Make a 30-second promo video — animation, voiceover, music, sound effects — with Claude and Claude Code. No paid APIs, no licensed music. Every prompt included.

To make a promo video with Claude Code: write a timestamped scene sheet first, generate the animation from it, then hand Claude Code the silent video plus your script and ask it to build the voiceover, score and sound effects locally. No paid APIs, no stock music, no editor. Here is the video that came out, and every prompt that made it.

The finished video. Voice, music and every sound effect were generated locally, at zero cost. Watch on YouTube

That is the promo for LetX, a collaborative LaTeX editor. The animation, the voice, the music and roughly ninety sound effects were all produced by prompting — and the audio side cost exactly nothing to run, because none of it calls an API.

I am not a video editor. I have never opened After Effects. The useful discovery was that you do not need to: the skill that produces a good promo video this way is writing a precise brief, and that is a skill you already have if you use Claude.

The one thing that matters more than the tools

Most people open this kind of project by asking for a video. That is the mistake, and it is worth being blunt about why.

A video is thirty seconds of decisions — what appears, when, for how long, and what sound lands on it. If you ask for all of that in one go, the model picks defaults for every one of those decisions, and defaults are what "AI slop" is actually made of. It is not that the model cannot do better. It is that nobody told it what better meant.

So the first thing you produce is not a video. It is a scene sheet: five scenes, what happens in each, and the exact seconds they start and end. Everything downstream keys off that document. The voiceover is written into its gaps. The music picks a tempo that lands on its cuts. The sound effects are placed against its boundaries. Get the scene sheet right and the rest of the pipeline is mostly mechanical.

The promo above has cuts at 3.0, 8.0, 19.0 and 25.0 seconds. Every single audio cue in it is defined relative to those four numbers.

The five steps

Each step below is a real prompt. Replace anything in angle brackets and paste it — they are written to be taken, not admired.

The five steps

Click a step to see exactly what to paste. Every prompt is a fill-in-the-blanks template — replace anything in <angle brackets>.

Where: ClaudeTime: ~10 min

Do not ask for a video yet. Ask for the script. The single biggest quality jump comes from making Claude commit to a second-by-second plan before any pixels exist — because that plan is what every later step keys off.

You end up with: A scene sheet: 5 scenes, what happens in each, and the exact seconds they start and end.

Paste into Claude
You are a direct-response video director. I want a 30-second vertical
promo (9:16, for Reels / Shorts / TikTok) for my product.

PRODUCT: <name> — <one sentence on what it does>
AUDIENCE: <who>
THE ONE FEELING I WANT: <e.g. "relief", "this is unfair how easy it is">
THE ONE ACTION: <e.g. visit letx.app and drag a project in>

Before writing anything, answer these three questions in one line each:
1. What is the single most painful thing my audience does today?
2. What is the one moment where my product obviously wins?
3. What would make someone stop scrolling in the first 2 seconds?

Then give me a 5-scene sheet with exact timestamps totalling 30.0s.
For each scene: the seconds, what is on screen, and why it earns the
next scene. Bias hard toward the first 3 seconds — if the hook does not
land, nothing after it matters.

Constraints:
- No stock-photo energy. No emoji. No generic "boost your workflow" copy.
- Every claim must be one I can actually verify. Flag any number you
  invent so I can replace it with a real one.

Two of those steps deserve expanding, because they are where the quality actually comes from.

Write the voiceover as a timetable, not a script

Ask for a script and you get prose. Ask for a script with a start time and a maximum length for every line and you get something a machine can place on a timeline without guessing.

This matters more than it sounds. In the LetX promo there is 22.5 seconds of speech inside 30 seconds of video. That gap is not laziness — it is what makes the video feel produced. Every line has air on both sides, and no line collides with the one after it or steps on a hard cut. You only get that reliably if the timing is part of the script rather than something you fix afterwards.

The other half of this step is spelling. Text-to-speech reads characters, not intent. Written normally, a speech engine says "dot tex" for .tex and stumbles over LetX entirely. So the script gets a second column written for the phonemizer: .tex becomes "dot tek", LetX becomes "Let X", letx.app becomes "Let X dot app". It looks wrong on the page and sounds right in the video, which is the only place it matters.

Then hand the whole thing to Claude Code

This is the step that surprised me. I expected to be shopping for royalty-free music and a sound effects library. Instead the working instruction turned out to be, roughly: synthesize the music and the sound effects from scratch, and put every hit on an exact frame.

That sounds far-fetched until you see what comes back. Drag the slider below — this is the actual cue sheet that got rendered, not a diagram of one.

30 seconds of designed sound

Drag the slider. This is the actual cue sheet Claude Code generated and rendered — every impact, riser and UI tick, keyed to the frame it lands on. Nothing here was hand-placed in an editor.

The Hook
The Problem
The Product
Proof
CTA
0.00s2.20s30.00s

Scene

1. The Hook

Earn the next 27 seconds

Voiceover

silence — 22.5s of speech in 30s

Firing now (1)

  • THE SLAMSub drop + crack + tail. The single most important frame in the video
ImpactTransitionUI / SFXMusic

There are three reasons this beats downloading a music track, and none of them are about being clever:

  1. Nothing is licensed, so nothing can be claimed. There is no attribution to track and no content ID to trip over when you post it.
  2. Every cue lands on the exact frame you asked for, because the sound is generated against your timeline rather than trimmed to fit it. The riser that ends precisely on the cut at 8.0 seconds is what makes that cut feel deliberate.
  3. It is reproducible. The generator uses a fixed random seed, so re-running it produces a byte-identical result. When you change one scene you re-render everything and only that scene moves.

You do not need to understand any of the signal processing to ask for this. You need to describe what you want to hear, and check the result.

Now build the prompt for your own product

Fill in your own details and this rewrites the opening prompt for you. It is pre-filled with the LetX answers so you can see the shape of a good response before replacing it.

Build your own prompt

Fill this in for your product. The prompt below rewrites itself as you type — then copy it straight into Claude. Pre-filled with the LetX answers so you can see what good looks like.

This becomes your first 3 seconds. Be specific and slightly embarrassing — recognition is what stops the scroll.

This is the money shot in the middle of the video.

#0A0A12
#9333EA
#22D3EE
Your prompt — ready to paste
You are a direct-response video director. I want a 30-second vertical
promo (9:16, for Reels / Shorts / TikTok) for my product.

PRODUCT: LetX — a collaborative LaTeX editor that runs in the browser
AUDIENCE: researchers, PhD students and academics who write papers in LaTeX
THE PAIN THEY LIVE WITH: emailing .tex files back and forth and losing track of which version is current
THE MOMENT I WIN: two people typing in the same document at the same time, and it just works
THE ONE ACTION: go to letx.app

Before writing anything, answer these in one line each:
1. What would make researchers, PhD students and academics who write papers in LaTeX stop scrolling in the first 2 seconds?
2. What is the one shot that proves "two people typing in the same document at the same time, and it just works" without me claiming it?
3. What am I tempted to include that I should cut?

Then give me a 5-scene sheet with exact timestamps totalling 30.0s,
structured as: hook / problem / product / proof / call to action.
For each scene give me the seconds, what is on screen, and why it earns
the next scene. Bias hard toward the first 3 seconds.

=== HARD TECHNICAL REQUIREMENTS (needed for MP4 export) ===
- Wrap everything in a single <Stage width={1080} height={1920} duration={30}>.
- Drive ALL motion from useTime(). Never use CSS @keyframes, CSS
  transitions, Math.random(), or Date.now() — every frame must be a
  pure function of time, or the export will not match the preview.
- Any on-canvas control gets data-export-hide.

=== CANVAS & SAFE ZONES ===
1080 x 1920 (9:16). Reels, Shorts and TikTok overlay their own UI.
Keep ALL text and key visuals inside:
  vertical:   y = 260 to y = 1560
  horizontal: x = 70 to x = 950   (right edge holds the like/share rail)
Background art may bleed to the edges; text may not.

=== BRAND ===
Background #0A0A12   Primary #9333EA   Accent #22D3EE
Feel: premium dark developer tool. Confident, not playful. Think Linear or Vercel.
Generous negative space. No drop shadows, no gradients on text,
no emoji, no stock illustration.

=== MOTION PRINCIPLES ===
- Ease everything with a smooth cubic ease-out. Nothing linear,
  nothing bouncy.
- Cuts are hard. At most one wipe in the whole video.
- Something must move in every single frame — no dead holds over 1.4s.
- The final frame must be legible as a still. It becomes the thumbnail,
  and it must show "LetX" and "letx.app".

=== HONESTY CONSTRAINT ===
Every claim must be one I can verify. If you write a number I have not
given you, flag it explicitly so I can replace it with a real one or
cut it. Do not invent user counts, ratings, or institutions.

The two fields that carry the most weight are the painful thing they do today and the moment you obviously win. Almost every weak promo video fails at one of those two, and no amount of animation polish rescues it. For LetX the pain was emailing thesis_final_FINAL_v3.tex to a supervisor, and the win was two cursors typing in one document. The video is essentially just those two ideas with three seconds of setup and five seconds of call to action.

Why the voice is local, and the dead end I hit first

Skip this section if you do not care, but it will save you an afternoon.

I started out planning to use an API for the voice, because that is what every tutorial does — usually ElevenLabs. I had an OpenRouter key already, so I tried to route it through there first, and at the time it had nothing that could speak. So the promo above was voiced entirely locally.

That has since changed, and it is worth knowing about. OpenRouter now serves openai/gpt-audio, which will read a line for you, and google/lyria-3, which generates actual music. When I built the 95-second explainer that followed this promo, I used gpt-audio for the voice: the full narration came to $0.094 per render, which is cheap enough that re-rendering to tune timing is effectively free. It is noticeably more natural than a local model.

Two things to know if you take that route. Audio output requires stream: true — a normal request fails with Audio output requires stream: true — and the audio arrives as base64 chunks you concatenate yourself. And there is no speed parameter, so when a line overruns its slot you re-render it with a pacing instruction rather than stretching it.

So pick your trade-off honestly: local is free, private, and unlimited; the API sounds better for about ten cents a render. The rest of this post works either way. Everything below still runs with no API key at all.

The macOS say command is right there and free, and it is unusably robotic — the good system voices cannot be scripted.

What worked was Kokoro-82M, an open-weight speech model small enough to run comfortably on a laptop. It produces the warm, unhurried delivery in the video above, it runs entirely offline, and it costs nothing per render — which matters more than it seems, because it means you can re-render the whole soundtrack twenty times while you tune the timing without watching a bill go up.

Three things I got wrong with it, so you do not have to:

  • It needs Python 3.12. On 3.14 the dependency chain fails to build at all.
  • Its speech-data dependency ships incomplete and needs files copied into place before it will run. Claude Code diagnosed and fixed this from the error message faster than I could have read the issue tracker.
  • Never stretch a take to fit its slot. If a line runs long, re-render it slightly faster. Time-stretching is precisely what makes synthetic voice sound synthetic, and it is the default thing an editor would do.

When it goes wrong

Every problem below actually happened while making this video. Pick your symptom.

When it goes wrong

Every one of these happened while making the video above. Pick the symptom you have.

Why it happens

Almost always one of two things: the text was stretched to fit a time slot, or it was spelled for humans instead of for the speech engine. Stretching audio is what produces that flat, synthetic quality — and an engine reading “.tex” says “dot tex” like a filename, not “dot tek” like a person.

The fix

Re-render the line at a slightly different speed instead of stretching it, and respell anything the engine will mispronounce. Generate one line at a time so each gets natural sentence-final intonation, rather than one long flat take.

Paste into Claude Code
The voiceover sounds robotic. Fix it without changing my script:

1. Never time-stretch a take to fit its slot. If a line overruns, nudge
   the TTS speed parameter and re-render it clean.
2. Render each line as its own take so sentence-final prosody is
   natural, then place them at their timestamps.
3. Respell anything the phonemizer will get wrong — read my script back
   to me and flag every word where the spelling and the pronunciation
   disagree (file extensions, product names, URLs, acronyms).
4. Trim the dead air the engine leaves at the start and end of each
   take, then add an 8ms fade so nothing clicks at the seams.

The one I would flag hardest is the loudness setting, because it fails silently. A mix that sounds perfect on your laptop can arrive distorted on Instagram, because the platform re-encodes it and a master sitting at the ceiling gets pushed over. The fix is a specific target — -14 LUFS with a true peak of -1.5 dBTP — reached with a two-pass normalisation. My first attempt used a single pass, which is dynamic, and it overshot to +0.4 dBTP. That clips the moment any platform touches it.

The part nobody warns you about: your own numbers

This is the section I would keep if I had to cut everything else.

A language model will write you a confident, specific, entirely invented statistic, and a promo video is the worst possible place for one. It is the artifact that gets embedded, reposted and screenshotted, and the claims in it outlive any correction you publish.

The LetX video makes three claims. Before it shipped, each was traced back to something queryable — verified institutional email domains, not a self-reported profile field. That distinction mattered: the optional country field was blank for most users and would have supported a far larger, completely unusable number. The house rule became publish the round figure below the verified count, so every claim is defensible and none of them expires as the product grows.

We still shipped one mistake. An on-screen card reads a number slightly above what the data supports, and it is on the fix list. The voiceover deliberately routes around it — it never speaks that figure. If you look closely at the video above you will not hear a single number that was not checked, and that was not an accident.

Round down. Always round down.

What it costs

Being precise, since "free" gets thrown around loosely:

ItemCost
Voice — local modelFree, unlimited re-renders
Voice — OpenRouter gpt-audio (optional upgrade)~$0.09 per 95s narration
Music and sound effectsFree, generated — nothing licensed
Assembly and exportFree, open-source tooling
Your Claude subscriptionWhatever you already pay
Per-render marginal costZero, or about ten cents with the paid voice

The disk cost is real though: the speech model pulls in a machine-learning runtime that lands around 1.1 GB. Budget for that once.

Time-wise: about 90 minutes end to end, most of it spent on the scene sheet and rewriting voiceover lines that did not sound right out loud. The actual audio render takes about two minutes and is fully repeatable.

Steal the whole thing

Every prompt on this page, the full scene sheet, the voiceover script with its timings, and the setup notes are in one repository:

github.com/shihabshahrier/claude-code-promo-video

Clone it, replace the brand tokens and the script, and run it. If you make something with it I would genuinely like to see it.

And if you write LaTeX — the thing this promo is actually for — LetX is free, runs in the browser, and lets you drag an existing project straight in. That is the whole pitch, which is why it fit in thirty seconds.


Related: four loop patterns for Claude Code and how LetX was built.

Frequently Asked Questions

Do I need video editing experience to make a promo video with Claude Code?

No. The skill that produces a good result is writing a precise brief, not operating an editor. The workflow in this post never opens a timeline UI — the video is generated from a timestamped scene sheet, and the audio is generated against those same timestamps. I have never used After Effects.

Can I use OpenRouter for the text-to-speech?

Yes — use the openai/gpt-audio model, which costs about $0.09 for a 95-second narration and sounds more natural than a local model. Two gotchas: audio output requires stream:true, and there is no speed parameter, so re-render an overlong line with a pacing instruction rather than stretching it. If you would rather pay nothing, a local open-weight model such as Kokoro-82M runs offline with no per-render cost.

Why generate music instead of using a royalty-free track?

Three reasons: nothing is licensed so nothing can be claimed or content-ID matched, every cue lands on the exact frame because the audio is generated against your timeline rather than trimmed to fit it, and a fixed random seed makes the whole render reproducible when you change a scene.

How long does it take to make a 30-second promo video this way?

About 90 minutes end to end for the first one, with most of that spent on the scene sheet and rewriting voiceover lines. The audio render itself takes roughly two minutes and is fully repeatable, so iterating on timing is cheap.

What loudness should a promo video for Reels, Shorts or TikTok be?

Target -14 LUFS integrated with a true peak of -1.5 dBTP, and reach it with a two-pass linear normalisation. A single-pass normalisation is dynamic and can overshoot above 0 dBTP, which clips as soon as the platform re-encodes your upload.

Why does the AI voiceover sound robotic?

Usually because the audio was time-stretched to fit a time slot, or because the text was spelled for humans rather than for the speech engine. Re-render the line at a slightly different speed instead of stretching it, and respell anything the engine will mispronounce — file extensions, product names and URLs especially.

Written by

Shihab Shahriar Antor — AI Engineer & Founder of Shahriar Labs. Creator of LetX, QuantumSketch, and more.

Share this mission log