I put off making YouTube videos for years because I didn't want to learn editing software. That was the whole blocker. Not the camera, not the talking, not the ideas. The timeline.
So I stopped trying to learn it. Now I film, say one sentence to Claude, and come back to a cut I can review. Audio leveled, dead space gone, text landing on the words I actually said, chapters written from the transcript.
The tools underneath are free. You install them once and stop thinking about them.
The part that matters most isn't the code
Everyone wants the scripts. The scripts are the easy half.
What makes this work is that one shoot lives in one folder, and the folder always has the same shape. Footage in one place, the script and the on-screen graphics in another, the video project in a third. The AI never has to guess where anything is, and it never has to ask.
Get the folder wrong and every clever script downstream turns into a conversation about paths.
The pipeline
- Film. That's your job and it stays your job.
- Normalize the audio to -14 LUFS, -1.5 dBTP. That's YouTube's own target, so hitting it means YouTube won't re-level you on upload.
- Transcribe with word-level timestamps. Not sentences. Words. This is the part everything else hangs off.
- Assemble in a Remotion project. One composition for a short, a transition series for a long.
- Place the text. You write what should appear and which spoken phrase it lands on. A script matches the phrase against the transcript and resolves the timing.
- Review it. Not optional. More on that below.
- Render.
What a skill actually is
It isn't an app and it isn't a plugin. It's three things sitting in a folder:
- Instructions. One text file, written in plain English, that your AI reads.
- A template. An empty video project it copies fresh for each video. You never touch it.
- A few helper scripts. The mechanical parts. Run once, forget them.
That's it. Mine is about 22 KB of English. The reason it works isn't that it's clever, it's that it's written down.
The part I still do
This does not run unattended, and I'd be selling you something if I said otherwise.
I watch it before it renders. I call the cuts. I make the graphics ahead of the shoot, because deciding what goes on screen is the actual creative work and I don't want that handed off. The AI removes the part I was never going to enjoy. It doesn't remove the judgment.
If you build something that claims to skip that step, you've built the wrong thing.
Take the prompt, not my files
I'm not giving you my setup. My folders aren't your folders, and my paths would break on your machine about four minutes in.
Here's the prompt instead. Paste it into whatever AI already has access to your files. Its first move is to interview you. Where your projects live, what OS you're on, what you already have installed, what your brand colors are. It refuses to build anything until you've answered.
That's deliberate. Forcing my folder structure onto you is the most common way this goes wrong.
You are helping me build an AI video editing pipeline. I edit YouTube videos by talking to you
instead of opening an editor, and I want the same setup on my machine.
Do not just generate files. Work through this with me, ask before assuming, and build it into my
existing folder structure rather than inventing one.
## Step 0: ask me these before you build anything
1. Where do my video projects live? Show me what you find and confirm before creating anything.
2. Mac, Windows, or Linux? Shell?
3. Do I already have ffmpeg, Node, Python 3, and Whisper installed? Check, do not assume.
4. Am I making vertical shorts, landscape long-form, or both?
5. What are my brand fonts and colors for on-screen text? If I do not have any, say so and we
will pick together.
Do not proceed until I have answered. My folder layout will not match the one this playbook came
from, and forcing a structure on me is the most common way this goes wrong.
## What we are building
Two related pipelines that share the same helper scripts.
**Short pipeline:** one talking-head clip becomes a finished vertical video with on-screen text
timed to my spoken words.
**Long pipeline:** several filmed segments become one landscape video with cross-dissolves,
chapters, and a description written from what I actually said.
The mechanics are deterministic scripts. The creative calls stay conversational between us.
## The pipeline, in order
1. **Normalize the audio.** Loudness to -14 LUFS, true peak -1.5 dBTP, LRA 11. That is the
YouTube standard and hitting it means YouTube will not re-level me.
2. **Transcribe with word-level timestamps.** Whisper with word timestamps on, flattened to a
simple list of `{ word, start, end }`. Word-level is the whole trick. Sentence-level
timestamps are not precise enough to land text on a spoken word.
3. **Assemble.** Scaffold a Remotion project. One composition for a short, a transition series
for a long.
4. **Place the text.** I write a plan of what text appears and which spoken phrase it should
land on. A script fuzzy-matches each phrase against the transcript and resolves the timing.
5. **Review.** Remotion Studio opens and I scrub it. This is not optional. See the note at the
bottom.
6. **Render.** H.264, CRF 18, preset slow, AAC 192k, 30fps.
## The numbers that matter
| Setting | Value | Why |
|---|---|---|
| Loudness | -14 LUFS, -1.5 dBTP, LRA 11 | YouTube's normalization target |
| Video | H.264 CRF 18, preset slow | Visually lossless, reasonable file size |
| Audio | AAC 192k, 48kHz | Clean, universally supported |
| FPS | 30 | Match the source or everything drifts |
| Cross-dissolve | 15 frames | About half a second, reads as intentional |
| Short frame | 2160x3840 (9:16) | Vertical 4K |
| Long frame | 3840x2160 (16:9) | Landscape 4K |
## The traps (these cost me real time, do not rediscover them)
- **Long-form: normalize audio only, copy the video stream** (`-c:v copy`). Re-encoding 4K for
eight minutes wastes an hour and gains nothing, because there is no orientation change.
Shorts are the exception, they need the full re-encode.
- **Never add a crop filter to fix portrait footage.** Phone and camera video is often portrait
stored in a landscape container with a 90 degree rotation tag. ffmpeg applies it automatically
when no crop is present. Adding a crop double-rotates and wrecks the framing. Always read the
rotation metadata first and report it.
- **A cross-dissolve overlaps both clips**, so the finished timeline is shorter than the sum of
the clips: `total = sum(clips) - 15 x (number of transitions)`. Overlay timing and chapters
must use transition-aware offsets, not raw cumulative durations. Getting this wrong pushes
every overlay and every chapter out of sync, and it drifts further the longer the video runs.
- **Anchor text to what I actually said, not to my script.** The take always drifts. Read the
transcript back to me before placing anything.
- **Flag any fuzzy match below 0.5 confidence** and make me resolve it by hand instead of
guessing. Pick distinctive anchor phrases, three to eight words, no common filler.
- **Check for overlapping text after placing.** If each overlay's end is computed independently,
two overlays in the same screen position can stack. Scan for one ending after the next one
starts and trim the earlier one.
- **Pin every Remotion package to the same exact version.** A caret range on a sub-package will
quietly install a newer build than the core and then fail to bundle.
- **Never copy node_modules between machines or operating systems.** Native bindings break.
Install fresh.
- **Hard-link media into the project's public folder, never symlink.** Remotion returns 404 on
symlinked files.
- **Whisper builds differ.** Some do not support naming the output file, so check what mine does
before you write the script around it.
- **Do not let the render port collide** with whatever else I run locally. Ask me.
## For long-form, also
- Chapters get timestamped from the FINAL render. If I trim anything in Studio, re-derive them.
Everything downstream of a cut shifts.
- Overlay density is lower than a short. One per beat. Do not caption everything.
- Thumbnails can be still frames rendered out of the same project over a real frame from the
footage.
## What I want back
- The helper scripts, written for my OS and my paths
- A project template I can copy per video
- A written description of the workflow, in my own folder structure, that I can hand back to you
next time so you already know the setup
- A list of what I still need to install, with the commands
## One thing to be honest with me about
This does not run unattended. It gets me to a reviewable cut fast, and then I still watch it
before it renders. I decide the cuts, I write the text, I check that the words landed where they
should. If you build me something that claims to skip that step, you have built the wrong thing.The useful part isn't the steps. It's the eleven traps in there that cost me real time to find: the transition math that drifts your overlays further out of sync the longer your video runs, the crop filter that double-rotates phone footage, the dependency that quietly installs a mismatched version and breaks the bundle. Those transfer to any setup. File paths don't.
What's next
The next video is the half nobody shows: how I decide what to make before the camera turns on. Scoring titles, building the outline, making the graphics ahead of the shoot. That turns out to matter more than any of the editing.
If you run the prompt, tell me what broke. That's genuinely the part I want to hear about.
Built with: Remotion 4.0 · ffmpeg · OpenAI Whisper · Claude Code


