Table of Contents

Hey all, happy Friday,
I'm sure you all missed me last week. I was supposed to go to Vegas to see Phish at the Sphere, but then James gave me a sinus infection. Not the end of the world, right?
Wrong.
Being the genius that I am, I tried to clear my clogged ear by holding my nose and blowing (is there a name for this? I swear an ENT told me this was a good method at some point. Anyway, screw that guy), which forced the fluid to push into my ear and collapse my eardrum.
So I couldn't fly to Nevada, as that stunt left me (temporarily) deaf in my left ear.
Now, some might argue that being deaf at a Phish show is a good thing. Some others might disagree.
We will just let bygones be bygones.

Said bygones
I took the week off from the newsletter anyway. The truth is, I didn't want to give anyone whiplash.
“Oh my god, I thought we were on break, I thought I didn't have to pay attention.” Etc.
Onto important matters:
Maddie and I have a babysitter this evening (Thursday), and we're going to see The Devil Wears Prada 2. I've never seen the original. Honestly, I'm just excited to get the Nitehawk popcorn. It rips.
James’ update: he is very close to walking (even his teachers say so!). He scoots with one hand on the wall down the hallway like he’s had a little too much merlot (I just rewatched that scene from Sideways, and honestly, I tend to agree. Who likes Merlot? Also, did you know Giamatti’s rant about the velvety red caused a massive decline in Merlot sales in the U.S. at the time? Wonders never cease).
He has very long hair (I feel like he’s giving Owen Wilson a run for his money), he still loves the swings, and he still enjoys a good fever now and again.
Throwing food off the high chair to his pal Linus, who, I will admit (I think this is the final stage of denial), is now noticeably heavier, has officially become a pastime.
And you know what? The consideration for doggy Ozempic aside, our cattle dog pup is ostensibly more loyal, and honestly, that’s the most important thing.
Consider them Calvin and Hobbes.
Anyway,
I generated an ad earlier this week. Looked clean. On-brand headline, Miracle Balm in frame, warm parchment background, the whole thing.
Then I ran it through my grader.
Four gates flagged:
The headline could run for any premium beauty brand if you covered the logo
The proprietary truths Jones Road owns (Bobbi as founder, original no-makeup makeup, multipurpose product, anti-anti-aging positioning) were nowhere on the ad
The grader called it "REGENERATE," LOW chance of performing in the auction, and gave me specific line edits to fix it.
Three rounds later, the same brief had Bobbi's name in the headline, two proprietary truths in the sub-line, and a verdict of SHIP, HIGH chance.
It went from a generic ad to a uniquely Jones-Road ad in about 15 minutes.
The thing nobody is doing is pre-launch QC on AI ads.
Here's why:
Image generators are great at making things that look like ads.
They are bad at noticing the specific things that kill performance once the ad hits the auction:
Safe zones get violated
Text vanishes at thumbnail
Logos appear on UGC formats they're supposed to hide from
Stats get fudged
Headlines could run for any brand in the category
None of these are obvious until you're staring at a CPM you can't explain.
I shipped a slate of fourteen AI ads two weeks ago without running the grader.
Hubris.
Three of them had headlines I would never have signed off on if I'd looked once more.
The performance showed.
The grader has paid for itself ten times over since then. Which is a low bar because I built it for free, but still.
What I do internally is run every ad through a creative grader before launch.
I want to send you the simplified version, the prompt that builds a version customized to your specific brand, and a real example of the workflow in action.
Generate ad. Grade ad. Apply specific fixes. Regenerate. Re-grade. Repeat until the verdict says SHIP.
Most marketers do step one and stop. The grader does steps two through five for you in 30 seconds per round.
Here is the actual run from my Jones Road test this morning. Same brief, four rounds, every round graded.

Round 1. Headline: YOUR LINES DON'T NEED MORE PRODUCT. THEY NEED BETTER PRODUCT. This was the version I almost shipped. I was so sure this was the one. I have never been more wrong about an ad I almost shipped on a Tuesday. The grader flagged four fails: safe zone, brand fingerprint, voice violations from the ALL CAPS load, and proprietary-truth leverage. Its reasoning on brand fingerprint: cover the logo, this could run for Augustinus Bader, Westman Atelier, RMS, or any premium clean beauty brand. No Bobbi voice. No founder fact. No category-claim Jones Road can make that nobody else can. Verdict: REGENERATE. Chance: LOW.
Round 2. Tightened the headline to YOUR LINES NEED LESS. NOT MORE. with a sub-line listing the four uses. Down to three fails. Headline still applied to other premium beauty brands. The proprietary-truth gate was still empty even though the sub-line listed product features. Verdict: FIX AND RESHIP. Chance: MEDIUM.
Round 3. Twist on JR's actual tagline. Headline: YOUR LINES, ONLY BETTER. (a mash-up of "your skin, only better" with "lines" instead of "skin"). Sub: Bobbi's Miracle Balm. The original no-makeup makeup. Down to one fail. Brand fingerprint cleared because the sub-line names Bobbi and the proprietary category. Only a borderline safe zone issue left. Verdict: FIX AND RESHIP. Chance: MEDIUM.
Round 4. Made the headline uniquely ownable by naming the founder and the category outright. Headline: BOBBI INVENTED NO-MAKEUP MAKEUP. Sub: This is her Miracle Balm. One pot, four uses. All ten gates pass. Brand fingerprint cleared because no other brand can credibly claim Bobbi invented this. Proprietary-truth leverage cleared because two of JR's strongest claims are in the copy. Verdict: SHIP. Chance: HIGH.
The grader's note on the final round: "This ad leverages multiple proprietary truths and demonstrates unmistakable brand fingerprint that no competitor could replicate. The combination of founder authority, category ownership, and practical product benefit creates a compelling value proposition with authentic Jones Road DNA. Ready to ship."
That is the difference between Round 1 and Round 4. Round 1 was a competent ad. Round 4 is an unmistakably-Jones-Road ad. Same brief, same brand context, same fifteen minutes.
Why this works
The grader does what your eye does not. Your eye gets used to your own work. You stop seeing what's missing because you've been staring at seventeen versions of the same Miracle Balm bottle for forty minutes. I have looked at 200 AI beauty ads in the last two weeks. My eye is useless at this point. The grader is not.
The grader catches what your brand voice agent cannot. That's the whole point. Voice agents are good at scoring tone. They are bad at noticing what is MISSING. The proprietary-truth gate is what catches "this could be any premium beauty brand." Your voice agent passes that copy because nothing technical is wrong with it. The grader catches it because it asks the question your voice agent does not: would this be ownable to anyone but you.
Same workflow on a totally different brand
Just to make sure this wasn't a Jones-Road-only thing, I ran the same iterate-with-grader loop on Ridge (the wallet brand) this afternoon. Different category, different voice register, different banned words, different proprietary truths. Same workflow.

Round 1. "The slim wallet you'll actually carry. Premium leather. RFID-blocking. Built for everyday." Looked clean to me. The grader called it REGENERATE / LOW with six fails. The biggest one was Gate 7 (Banned Language): Ridge's grader specifically flags the word "wallet" used alone, because any slim-wallet brand can run that copy. The brand wants "wallet" paired with a material or a use case ("carbon fiber wallet," "EDC wallet," "front-pocket wallet") so the ad reads as Ridge and not as Bellroy or Bryker Hyde or any of the other twenty slim-wallet brands. I forgot this rule. The grader did not.
Round 7. After a few rounds of grader-driven tightening, landed on: "TWO MILLION CARRY THE HERITAGE EDITION. Oil-waxed leather. Patented dual-track expansion." This was the unlock moment. "Two million" is Ridge-proprietary (no slim-wallet competitor can claim that scale). "Heritage Edition" replaced "wallet" entirely as the noun. "Patented dual-track" is the product's actual technical differentiator. Nine of ten gates passed. The grader's verdict: "Strong proprietary truth leverage and excellent brand voice. The 'two million' social proof + 'patented dual-track' tech + material callout creates excellent credibility architecture. Fix the placement and this is a strong performer for warm audiences."
Round 13. Same locked copy, rendered at native 1080x1920 with maximum top-buffer prompt directives. Every quality gate passes. Ridge brand fingerprint intact. The composition has clear top headroom, the product is hero, the CTA is anchored bottom. Production ready.
The path on Ridge took more rounds than Jones Road because Ridge has a stricter banned-word rule (the "wallet alone" trap kept catching me). That's the value of the brand-specific gates. They encode the rules I'd otherwise have to remember on every brief. The grader doesn't forget.
A practical note: across both Jones Road and Ridge, the grader was conservative on Gate 1 (Safe Zone) even when the layout looked reasonable to my eye. That's by design. The grader is a pre-launch gate, not a final stamp. When it flags placement, you either over-specify pixel positions in the next regenerate or do a 10-second crop in Figma. Either way, you ship with the safe zone respected.
How to build your own Agent (10 minutes)
Click below to download the meta-prompt. The file walks you through the four uploads (your brand bible, 5-10 best ads, 3-5 ads you killed with reasons, your banned words list), then the prompt itself. Paste it into Claude with your uploads, get back ONE custom grader prompt for your brand that you save and use forever after.
Save it down into your folder and call it a day.
How to use it once you have it
You generate a static. Before you put it in a deck, before you upload it to Meta, before you show it to a client, you paste your grader prompt into a fresh Claude conversation and drag in the ad image. 30 seconds later, you have a verdict.
If it says SHIP, ship.
If it says FIX AND RESHIP, the grader names the specific gates that failed and the specific fixes. Apply them to your prompt or your copy. Regenerate. Re-grade. Most ads I run hit SHIP after one or two iterations.
I ran it multiple times for this newsletter as an example.
TLDR
Stop shipping round-one AI ads. The model produces things that look fine and miss the specific failure modes that kill performance.
Build a custom grader once. Upload your brand context, 5 to 10 best ads, 3 to 5 killed ads with reasons, and a banned words list. Paste the meta-prompt. Save what comes back.
Run every ad through your grader before launch. Apply whatever fixes it names. Iterate until the verdict says SHIP.
I tested this on Jones Road this morning. Round 1 was REGENERATE, LOW chance. Round 4 was SHIP, HIGH chance. Fifteen minutes. Then I ran the same workflow on Ridge to make sure it generalizes. It did.
Reply with your custom grader if you build one. I want to see what brand-specific gates other people surface. Most of mine came from shipping the same headline structure three times in one month before someone on the team finally pointed it out.
Will