In partnership with

Hey all, happy Friday,

First, the big news from our house: James took his first steps. Sunday, out on the Upper East Side, he stood up between Maddie and her mom and walked from one to the other. I was in the next room on a call and missed the lad’s penguin walk, which is exactly what happens when you book a call against a one-year-old. I got in fast enough to catch the next ten, him grinning like he'd been holding out on us.

He's cruising everywhere now: stands up, grabs whatever's nearby, gets from point A to point B faster than feels fair.

Surgery week did have a short sequel. We were thrilled to send him back to daycare on Monday, and a few hours later, they sent him home with a fever. It was just a post-op spike, nothing more, and he was back Tuesday morning fully unbothered. Tubes working, kid thriving, parents aged a little (I hear people our age do Botox these days to remove these lines?).

Today we drove up for a wedding and split the convoy: Linus and I took the car, Maddie and James took the train. We both drew the short straw. A four-hour drive became six and a half on a parked stretch of I-95, Linus riding co-pilot with his head out the window. We made it through with some Ezra Klein and a Hidden Brain, and the second we cleared the jam, we put on Phish and really let it rip.

The train ran its own program according to Maddie: James paced the aisle the whole way, beaned a stranger in the head with his pacifier, then (because the universe is kind) befriended the same man over peekaboo before stealing, and to his credit, returning the guy's chips.

We fly to France in two months. I cannot wait.

My sister and her family are up here with us, too, so I'll be on and off on Friday between the good parts.

Also,

If you’re not using Wispr Flow with Claude Code yet, you’re wasting time. Talk to text is a fundamental game-changer, and I wouldn’t be putting this here if I didn’t use it every day. Sometimes I will speak to it like a drunken sailor, and it comes back like I am F Scott Fitzgerald…

Your prompts are leaving out 80% of what you're thinking.

When you type a prompt, you summarize. When you speak one, you explain. Wispr Flow captures your full reasoning — constraints, edge cases, examples, tone — and turns it into clean, structured text you paste into ChatGPT, Claude, or any AI tool. The difference shows up immediately. More context in, fewer follow-ups out.

89% of messages sent with zero edits. Used by teams at OpenAI, Vercel, and Clay. Try Wispr Flow free — works on Mac, Windows, and iPhone.

Anyway,

Almost every brand I talk to has the same thing buried in a Google Drive folder: a folder of raw clips named something like “Product B-Roll_FINAL_v2.”

  • Phone footage of the product

  • Old creators that show the product

  • A few close-ups

  • Somebody’s hands using the thing.

Nobody knows what to do with it, so it sits there going stale while the same brand runs ads that look worse than what’s already in the folder. (I am now a person with strong feelings about unused b-roll. This was not the plan for my life.)

So here’s the question I wanted to answer with nothing hidden: how do you use AI to turn that folder into one finished vertical 9×16 ad, without an editor, without editing software, and without guessing whether it’s any good?

The answer is an assembly line, and the secret ingredient at the end is an AI paid to be a jerk. I ran it on a real client’s folder (I’m keeping them anonymous) and watched it climb from the client’s own “20” to a 93. Below is exactly how you run the same thing on your own clips, start to finish. After that, if you’re curious, what it’s doing while it runs.

Here’s how you run it yourself

If all you do this week is the five steps below, you’ll have a finished vertical ad by the end of it. No editor, no software you have to learn, no guessing.

1. Install the two free things. You need Claude Code (the tool that runs the prompt and does the work) and ffmpeg (the free, no-interface video tool that does all the cutting). On a Mac, ffmpeg is one line in Terminal: brew install ffmpeg. You also need an ElevenLabs account for the AI voice and music. The free tier is enough to try the whole thing. Honestly just ask Claude Code to do all of this for you.

2. Make one folder and drop four things in it. Call it whatever you want. Inside, put a clips/ folder with your raw b-roll (any number of .mp4 or .mov files, they do not have to be good). Put a brand/ folder with your logo, a few product photos, and a one-page note listing your fonts, your colors, and any words you’re not allowed to say. Add a .env file with one line, your ElevenLabs API key. Then drop in the two files from the kit, PLAYBOOK.md and auditor-rubric.md. That’s the whole setup.

3. Open the folder in Claude Code and paste the prompt. The prompt is the file called START_HERE.md in the kit. Copy it, paste it in, hit enter. That’s the last technical thing you do.

4. Answer five questions, then approve twice. First it interviews you: what’s the product, the one thing to remember, who’s it for, your fonts and colors, any claim you can’t legally make. Then it writes the ad as nine lines, asks you to approve them, and proposes one clip per line, asking you to approve that too. Those two approvals are your whole job. Everything else runs on its own.

5. Let the mean AI do the grading. When the ad is built, a separate AI scores it out of 100 and refuses to pass anything under 90, handing back exact fixes and rebuilding until it clears the bar. Then it gives the finished ad to you for the final call. You shipped an ad and never opened a video editor.

If all you wanted was the ad, you can stop right here. What follows is for the curious: what the assembly line is actually doing while it runs.

For the enthusiastic readers, here are the nine steps the AI is running:

1. Write it as nine beats, not a paragraph. A beat is one idea, one line, one shot. Don’t sit down to “write an ad,” that’s how you stare at a blank page. Write nine: two hooks, what-it-is, a demo, a proof point, two benefits, a spec, a close. Every later station keys off them, each beat gets its own clip, caption, voice line, and slice of time. One rule the AI has to follow: an early draft bragged “4.7 stars” with no source, so we cut it from the voiceover and the caption. Your copy AI has to refuse claims you can’t back up. An unsourced number is a legal risk, not a clever line.

2. Speak each line, then time the video to the voice. This is the one that trips up everyone: you cannot guess how long a spoken line will be. Pencil in “3 seconds,” the voice takes 3.8, and the words drift off the picture like a bad dub. So flip the order. Generate each line with ElevenLabs, measure its real length, trim the silence, and set that beat’s on-screen length to match. Speak first, time second.

3. Add music, and keep it out of the voice’s way. ElevenLabs makes the instrumental too: describe the vibe (“clean corporate instrumental, no vocals”) and the length. Then the rule that makes it sound professional: duck the music under the voice so it dips whenever the voice talks and comes back in the gaps. We park it about 22 decibels down. If you ever actually notice the music, it’s too loud.

4. Pick one clip per beat (I assumed this was my job; the AI did it). This is the part I braced to do by hand, the hour of opening clips and muttering "nope." It didn't happen. The AI proposed a clip for every beat, told me what each one showed, and left me exactly one move: scan the mapping and veto anything wrong. Each beat needs a clip that shows what the line says, and the machine enforced two non-negotiables while it picked. Only the advertised product is allowed on screen (it tossed clips with a competitor's label or a stray marker pen, because you don't run an ad that accidentally advertises someone else's product). And no leftover text baked into the footage (you can't remove burned-in text, only crop around it). The product's own label is fine. That's the product, not contamination.

5. Reframe to fill the frame, and color-grade. Two moves do about 80% of the visual work. Fill the frame; don’t add bars. Your footage was shot wide; the lazy fix is shrinking it and padding the gaps with blurry bars, the dead giveaway of a rushed ad. The right fix is “cover”: blow the clip up until it fills the tall frame edge to edge and crop the overflow. Then color-grade every clip, because raw product footage comes out cold, flat, and near-gray. Nudging saturation and contrast until it looks alive is most of why the client’s “20” stopped looking like a 20.

6. Burn the captions onto the screen. Most people watch muted, so the captions carry the message. Time each one to the same beat boundaries as the voice, style it to brand, and keep it in the safe zone, the middle band of the screen clear of the app’s like and share buttons along the right and bottom.

7. Stitch and combine. One pass joins the nine cropped, graded, captioned beats end to end and layers the audio: voice at full volume, music ducked underneath. Out comes a single 1080x1920 file. You had nine pieces; now you have one ad. Whether it’s a good ad is the next station’s problem, and that station is the whole reason this works.

8. Hand it to an AI paid to be mean. The secret weapon. Everything above is assembly; this is judgment. You hand the finished ad to a separate auditor AI with one instruction: find everything wrong with this. Don’t be nice. It scores the ad out of 100 against a written checklist, because an encouraging reviewer is useless and you stopped seeing the flaws an hour ago.

Two things make this real and not a vibe check. Hard fails cap the score: one disqualifying thing (a competitor’s product in frame, leftover text, a warped object) caps you at 60 no matter how nice everything else is. The gate is 90: below it, the auditor hands back a ranked to-do list with timestamps and the exact fix for each. You fix the top one, rebuild, re-score, loop. Here’s how the climb actually went: the raw folder was the client’s “20,” the first full assembly already jumped to 91, and the audit loop then carried it through 88 and 89 to a passing 93.

The jump from 89 to a passing 93 came from one fix the auditor named precisely. Our team had watched this ad something like forty times. It looked fine. The auditor flagged a single frame where a puff of vapor washed the shot near-white, a split-second nobody catches at full speed, named the exact beat, and told us to pull the out-point in by a fraction so it lands on a crisp frame instead. We did. It passed. So I am now taking notes from a machine that had never seen the product before, and the machine was right. You don’t eyeball that. The machine measures it.

9. A human makes the final call. The auditor gets you to “objectively clean,” not “done.” A person still decides feel, rhythm, and whether it’s actually on-brand, then ships. But notice what changed: the human reviews a 93, not a 40. Machines on the labor, humans on the decisions. That’s the whole point of the assembly line.

The one guardrail I’d never skip

Nothing renders until the brand basics are validated in that session: fonts, exact colors as hex codes, logo rules, banned words, official product photos, all checked from the original source, not from memory or a folder that may have drifted. The accuracy guardrail and the hostile auditor are two different jobs: one keeps you from shipping something wrong, the other from shipping something boring. You need both, at opposite ends of the line.

The full station-by-station build, the exact ffmpeg moves, and the five gotchas that each cost me an hour are all in the kit, so I’m not retyping them here. The worst gotcha, for a taste: back-to-back captions can overlap by a single frame, and I spent an hour convinced I’d broken causality before I found the one-line fix.

TLDR

  • You don’t need an editor to turn raw clips into a vertical ad. You need an assembly line with a conscience.

  • Break it into nine beats: one idea, one line, one shot each. Let AI write, speak, and score it.

  • Generate the voice first, then time the picture to it (your guess is always wrong). Fill the tall frame by cropping, never bars. Hand the cut to a hostile AI auditor with a hard 90 gate.

  • Our real run climbed from a “20” to 93. The last point came from a fix only the machine could measure. Grunt work to the machine, taste to you. That order is the trick.

Next week I’m going to try to automate the one manual step, assuming I don’t spend three days arguing with myself about whether “clean” means sans-serif or just “not visibly dirty,” which is genuinely the kind of thing I’ll do instead of the thing I said I’d do. I have a folder of unfinished projects that can confirm this.

But if I can turn a brand’s clip folder into a tagged, searchable library, every clip labeled with what it shows, the footage-picking station could stop being hand-work and start being an AI match, one line to its best clip. That’s the difference between making one ad and making a hundred.

Reply and tell me what’s sitting in your clip folder. Better yet, run this on it and send me the auditor’s first-pass score, the ugly first number, before the fixes. The first number is the honest one.

Have a great weekend.

Will

Recommended for you

View all
caret-right