Hey all, happy Friday,

Picking up where last week’s traffic jam left off: we made it to Rhode Island, and James got a full weekend with his cousins, who he loves with his entire body. They spent most of it at the water table, which went great right up until the water table got flipped onto James, who ended up soaked and deeply offended.

At the beach, he committed to a strict all-sand diet. We fought it for about an hour, then accepted that this is just who he is now. He and his cousin also ran coordinated raids on my mom (whom the kids call Yaya), who loved every second of it (or at least that was my sister’s and my interpretation 😉).

Stack on a cousin's level supply of ice cream, and we got the full cycle on repeat: sugar rush, laughing attack, crash.

Linus, meanwhile, treats leaving the city like parole, so he spent the entire weekend chasing bunnies. He does have an electric fence, so I think the bunnies were more taunting him than anything else.

The place we stay also has a thriving skunk community, and Linus has been skunked three times in his career. He is black and white himself, so I assume he sees them as family. (If your dog ever gets skunked, genuinely, hit us up. We are experts now. It is not tomato juice.)

This trip: zero skunkings. Growth.

A highlight of the week (aside from obviously hanging with Maddie, James, and Linus): Wednesday night, I went to a bar to watch the Knicks game. I am not a Knicks fan. I am not a sports guy at all, which you have probably gathered from my rants about ad pipelines. My lifetime sample size of complete basketball games is unlikely to be double digits, and so when I say that comeback was the best game I have ever watched, please discount it heavily.

We don’t have ESPN at home, so the current plan is a suspiciously well-timed free trial before game 5 on Saturday. Meanwhile, Maddie has been listening to the playoffs on the radio (it comes on right after FDR’s fireside chat), which I am told is the most romantic possible way to consume basketball, so, relatively speaking, she is the household superfan.

And two of our best friends just had their baby a few days ago, both die-hard Knicks fans, so their kid’s first weeks on earth are doubling as the most memorable sporting event of their lives (presumably).

Pretty good timing, kid.

Anyway,

Anthropic shipped a new Claude model this week. Every time this happens, my feed fills up with the same question: is the new one better? And every time, I think that’s the wrong question.

Nobody at a company asks, “Is the new hire better?”

They ask, “Better at what?” Better at strategy? Better at copy? Better at catching the mistake before the client does?

So I ran the interview I’d actually run. I took a real ad account, handed the same raw export to both models, and made each one do every job in our pipeline, alone, start to finish.

Then I made them grade each other.

Mostly blind.

We’ll get to the “mostly.”

The math converged. The strategy did not.

Both models, working alone, drew the same picture: an account with a handful of hero ads carrying most of the spend, firing a good chunk of the budget at one single persona, and an expensive gluten-free joke ad quietly losing money. Both even picked the same dark-horse product (a chelada with lime and sea salt) and the same calendar moment (July 4th) without being told.

Then they walked in opposite directions.

Fable bet on structure. The category leans hard on offer ads, and this account is not running enough of them, so it built one (“EIGHT STYLES. 24 CANS. $54.99.”). The account’s best recent statistics are native screenshots, so it built the next surface in that franchise: a group text that asks “what should I bring Saturday?”

Opus bet on story. It dug up the one piece of ammunition nobody had ever used: this little independent brewery’s double IPA took Gold at the awards where the biggest corporate-backed competitor took bronze, same category.

Ten finished ads, five per model, one generation attempt each, no do-overs. Every single one came back with clean, correctly spelled text on the first try, which would have been the headline of this email a year ago, and today is just the baseline.

Opus on top, Fable on bottom

Four jobs, judged separately

Here’s where it gets fun. I didn’t want one winner.

I wanted a hiring decision per seat: who analyzes the account better, who writes better copy, who writes better image prompts, and who QCs better. The judging panel: both models grading each other (blind wherever possible), plus a neutral third model as the measurement instrument on the objective stuff, plus me, the human, fact-checking the judges.

Seat one, analysis: Fable, narrowly, both judges. Honestly this seat was close to a tie on competence. The judges recomputed over a hundred claims from both analysis docs against the raw data and found basically everything checked out. The gap was interpretive: Opus argued the account “ignores Relief” as an emotion because the benchmark says the category over-indexes on it. Except Relief is already this account’s top-earning emotion, by a comfortable margin. Fable funded the thing that works. Opus tried to discover a thing that was already discovered.

Seat two, copy: Fable, clearly, both judges. This is the one that matters most, and the reason is one single quote. Stay with me.

One of these quotes is not real

Both models independently landed on the same concept for one slot: the gluten-free buyer is underserved, the reviews prove it, so let a real customer say it on the ad. Identical strategy. Then:

Fable pulled an actual review, verbatim: ”Being GF and NA is pretty niche and hard to find.” Slightly awkward, clearly human, real.

Opus wrote: ”Finally, a gluten-free beer I can actually enjoy.” Cleaner. Punchier. Honestly the better line.

It does not exist.

No customer ever said it.

And the ad renders it in quotation marks with a “Verified buyer” credit underneath.

Fable on left, Opus on right

It gets worse before it gets better.

Opus also trimmed words out of a second real review while keeping it inside quotation marks.

None of this is dumb.

The fabricated line is what a real customer WOULD say, and a human copywriter would absolutely pitch it.

The difference is that a human copywriter pitches it as a headline. The model dressed it as evidence. In a regulated-adjacent category, with the FTC’s opinions about testimonials, that’s not a style note.

That’s the difference between an ad and an exhibit.

Seat three, image prompts: Fable, by a nose, 98.2 to 96.6 per the neutral audit. Ten for ten on clean text on both sides. The entire gap is one fact: Opus rendered the wrong award category name onto an ad, and the error didn’t come from the image model. It was born in Opus’s own copy doc and faithfully executed from there. The pixels did exactly what they were told. That sentence is most of what you need to know about AI image pipelines in 2026.

The QC seat, where it stopped being close

Both models then received the same 10 finished ads, fully blind this time (just anonymized PNGs), and served as production gatekeepers using our standard grading rubric.

Opus’s QC pass scored the fabricated-quote ad an 88 and the founder ad an 87, flagging an “attribution-integrity wobble” and noting, of the founder portrait, that someone “should confirm it’s the real Joe.”

Reader, it was not the real Joe.

It was a fully synthetic man in a baseball cap, generated by the image model, captioned as the founder of a real company.

Fable’s QC pass hard-failed both ads to zero. It checked the quote against the review corpus and called the fabrication a fabrication. It diffed the second quote word-for-word and caught the rewording. It flagged the synthetic founder likeness as unshippable.

Then it verified the things that SHOULD pass: the kit price math, every award claim, and the verbatim review of the other gluten-free ad. Catching problems is half of QC. Not inventing problems is the other half, and the one false flag of the round actually belongs to Opus’s QC, which dinged the brand’s own signature tagline as off-brand. The signature tagline. It’s in the brand book as the one sanctioned use of the word “cheers.”

To be fair to Opus, its QC found one real thing nobody else caught: two ads in the batch cite the same medal under two different category names, a cross-ad consistency catch that’s genuinely hard. It can see. It just graded on a curve when the stakes called for a checklist.

The part where I have to rat on my own experiment

Three confessions, because receipts cut both ways.

One: the text-judging blind leaked. Both models had signed their analysis docs like proud kindergartners, model name right in the header, and my anonymization renamed the files without scrubbing the contents. So for the text seats, the judges knew.

The reason I still trust the result: Opus, fully able to see which work was its own, voted against itself on both seats. Whatever bias survived that leak was pointing the wrong way to help it.

Two: the pixel blind held, and it produced my favorite useless stat of the experiment. Asked to sort the ten anonymous ads back into their two authoring teams, both judges managed 6 out of 10. A coin gets 5. The models have fingerprints in prose and apparently none in pixels.

Three: the least reliable instrument in this entire experiment was me. I eyeballed a render at thumbnail size, read “2025 World Beer Cup” as “2021,” and confidently logged a dating error for the fidelity audit to confirm. The audit came back: the ad was right and the human was wrong. I’m keeping that one in the email as a public service.

The verdict

One account, one run, five ads a side. I’m not writing a benchmark paper. But every seat pointed the same direction, and the direction has a shape:

Fable runs the pipeline. Analysis, copy, prompts, QC. It treats brand law like law, traces every claim to a source, and its QC pass is the first one I’ve seen that I’d let near a client deliverable without me hovering.

Opus stays in the room. The Gold-versus-bronze wedge was, per every judge including Fable, the single best strategic idea anyone had all day. The best hooks of the round were Opus’s.

If I’m hunting for the angle nobody’s run yet, I’m still asking Opus first.

I’m just never again letting its output skip the gate.

If you’re running this at home

Three takeaways, in order of how much money they’ll save you:

1. Strip the thinking out of the data before you hand it over. If your export contains anyone’s grades, tags, or recommendations, every model you test will hand you back the same slate, and you’ll conclude AI is a commodity when actually your input was. Raw rows in, real differences out.

2. Hire per seat, not per model. “Which model is best” is a benchmark question. “Which model do I trust with which job” is a business question. In our run the answer split: one model for the pipeline, the other for the ideas. Yours may split differently, which is exactly why you test it on YOUR account instead of reading someone’s leaderboard.

3. Never let AI copy ship without a claim gate. The most expensive sentence in this experiment sounded like the best one. The detail that matters: the model that wrote the fabricated quote scored its own ad 88 out of 100. A fresh session running the auditor prompt scored it zero. Same intelligence, different seat. Use the gate.

Run this on your own account

The expensive part isn’t the tokens (well, actually, it is sort of is… Fable is expensive!), it’s reading two strategy docs side by side and admitting which one you’d actually pay for.

If you run a test, reply and tell me which seat surprised you. For me, it was QC. I expected the new model to write nicer sentences. I did not expect it to be the only one in the building that checks whether the customer is real.

Have a great weekend,

Will

P.S. Game 5 is Saturday. If the free trial works, I will be watching basketball number twelve-ish of my life. If the Knicks win the title the same week a new Claude wins our bake-off, my friends’ newborn is officially the luckiest kid in New York.

Recommended for you

View all
caret-right