One caveat before we start. OpenAI shipped GPT-6 Astra today, two days after Fable 5.1 and about 2 hours after I finished this horse race. Nothing below touches Astra. Next week's issue, we will focus on Astra vs. Fable as it pertains to marketing. This week is solely focused on the most recent Claude update.

Happy Friday,

Last Saturday we hauled ourselves up to Westchester for a two-stop day, and a two-stop day with a one-year-old is basically an Ironman with worse snack logistics and thankfully no Gu’s. First stop was the house of one of Maddie’s friends: two little boys, a newborn daughter, and a glorious driveway. James was so excited to be around the boys that he was essentially humming. The family owns one of those ride-in remote-control cars, where the kid sits behind the wheel, and a parent steers from a handheld remote, which happens to be the precise toy my cooler childhood friends owned and my parents refused to buy. James tore across that driveway like a man who’d spent his whole life waiting for F1 to finally call his number. The suburbs also delivered a sprinkler system, water guns, a snack tray, and a ball pit. Maddie’s friends are, structurally, a playdate operation, and James is a ball guy, so as far as he knew, we had chauffeured him to paradise and told him it was just Saturday.

Stop two was our friends’ place in Scarsdale, which comes with a pool, and a pool turns James into a different animal entirely. We swam, the nap got sacrificed, and by evening he was operating at full feral: eating a quantity of cookies I refuse to commit to print and doing laps like a small man fleeing the scene. My buddy’s niece, who is four, installed herself as James’s official interpreter. He’d produce a string of gobbledygook and she’d face us and announce, “He’s saying he wants the ball.” Pure Angelica from Rugrats.

We’ve opened conversations about a retainer.

Sunday turned gloomy, so we grabbed what park hours we could, read a lot of books, and ate Maddie’s zucchini chicken meatballs, which earned unanimous household approval, including from the members of the family who cannot pronounce “meatball.”

Later that night, Maddie had the new Housewives show on, and I sat down next to her and stayed. I enjoy Summer House. I enjoy Real Housewives of Salt Lake City. This one is either above my pay grade or below it, and I’m going with “above” because it’s kinder to me, so we’ll see where things land. This coming weekend, James gets his cousins, which ranks first among things that are not balls.

Anywhooo.

Back in June, Anthropic released Fable 5, and I pitted it against its predecessor on a real client’s ad account. Fable took every single seat. It handled the analysis more cleanly, produced copy where every claim traced back to a source, and its QC pass was the only one anywhere in the building that caught an invented customer quote. I told you to give Fable the pipeline and got on with my life.

Now, three months on, Fable 5.1 exists, and my entire feed is repeating the same three claims: it’s cheaper, it codes better, it sounds less like a robot. Every one of those might be accurate, and not one of them answers whether I should swap out the model writing our briefs. So I reran June’s experiment, except this time the champion is up against itself.

What Anthropic actually released

Facts first, then the fight. Fable 5.1 arrived Tuesday priced identically to Fable 5 ($10 per million tokens in, $50 out), except cached input now runs at a quarter of its old cost, which Anthropic translates to roughly 25 percent savings on ordinary work and as much as 45 percent on agent-style jobs that keep re-reading the same files all day long.

It’s also the first Claude whose text carries a watermark. It’s invisible, encoded in word choice, and detectable by regulators and researchers through an API neither of us can access. Practically, it means this. If your workflow is pasting raw model output into a newsletter (eeeeek!), a blog post, a LinkedIn essay, or anything a regulator could ever take an interest in, you should assume that text can be flagged as machine-written. If the model drafts a brief and a human rewrites it before anything ships, which describes every ad in this issue, the watermark doesn’t touch you. You should know which of those two people you are. Most of you are the second one, and staying that way is the move.

The quieter release was a prompting guide. Addy Osmani posted it on Wednesday as the sanctioned method for getting the Claude out of Claude’s prose:

The guide gives the disease a name: mannered prose, the habit of writing “a dial worth turning” when the idea was “a parameter worth varying,” because the fancier phrase flatters the writer rather than delivering the point. The cure is a single paragraph pasted into your prompt, or, compressed to five words: please remove all mannered prose. I’ve been battling exactly this disease in ad copy for a year now, so it became one arm of the experiment.

How the rematch was built

June’s account, June’s champion, and the upgrade that

The client is the same one from June. I pulled a new export from our dashboard covering the last 90 days, 121 ads, and deleted every grade, every category, and every recommendation, so no model inherited anyone’s opinion. Each model then received everything a strategist would see on their first day: raw ad performance including the on-image copy of every creative, 360 verbatim customer reviews pulled from the brand’s own review platform, the brand’s voice rules and product catalog, the market digest from the swipe file, and the creative-handoff sheet cataloging all 382 concepts we’ve already briefed for this brand, along with an instruction to repeat none of them.

The assignment was not a strategy document and not a batch of renders. It was the actual thing my team ships on a Tuesday: five paste-ready static briefs in our spreadsheet format, meaning two iterations on a proven vein, one reskin of a winner, and two net-new personas, each brief carrying a claim-trace table beneath it. Every model got a single attempt. No safety nets, no nudges, no second tries.

That made four runs total. Fable 5. Fable 5.1 with nothing added. Fable 5.1 carrying the mannered-prose paragraph. And Fable 5.1 carrying a second paragraph from the same guide, which I’ll get to shortly, because needing it was not part of the plan.

Round one ended in nine minutes

An identical 40,000-token output ceiling for every ru

Every run got the same output ceiling: 40,000 tokens, which is roughly 30,000 words and comfortably more than five briefs require. Fable 5 landed at 38,000 and completed the job. Five briefs, five claim tables, eight minutes.

Fable 5.1 slammed into the ceiling. Twice. Each run took nine minutes, spent the full 40,000 tokens, and delivered 395 words in one case and, in the other, absolutely nothing. It wasn’t an error. It wasn’t a refusal. The model had spent the whole time thinking, thinking is charged against the ceiling, and the ceiling was gone before the typing began. I sat staring at those 395 words and a truncated table for longer than dignity permits, and my working theory was a broken API, which is the theory of a guy who hasn’t opened the documentation.

The 395 words that did arrive were genuinely sharp. The model had located the vein, called out the dead money, listed formats the account had never touched and personas the reviews kept sketching that no ad had ever addressed, and then it stopped mid-table like a witness who clams up the instant his lawyer rises.

Once I actually read the documentation, the answer was sitting there. It generally is. People tell me this is a recurring theme with me. Anthropic’s 5.1 guide explains that on a long deliverable, the model may compose the entire output inside its thinking and then write the whole thing out a second time, so long requests need a bigger ceiling plus one particular paragraph appended to the prompt. That paragraph informs the model that its thinking and its reply draw on a single budget, and that drafting the full output twice adds nothing. For round two I lifted the ceiling to 128,000 and gave a third copy of 5.1 that paragraph.

Every copy finished this time. The bare one spent 52,000 tokens over eleven minutes. The de-flavored one spent 55,000, also eleven. The one carrying the budget paragraph spent 40,000 in eight minutes, and its briefs lost nothing for the efficiency. One paragraph bought a quarter fewer tokens and three minutes. Anyone running 5.1 on long work should paste it. It’s in the kit, and I remain mildly irritated that the solution was written down in plain English the entire time.

The hooks, next to each other

Fable 5 on the left, Fable 5.1 on the right. Identica

Please look at the image before this paragraph, because the image makes the argument.

Both models sourced their headline quotes from real customer reviews. Across twenty briefs, not a single quote was fabricated, and in June that one metric decided the entire copy seat. Both models identified the same dead ad (a gluten-free joke that lost money) and drew the same conclusion (this buyer wants facts, not comedy). Both chased the gluten-free drinker, both deployed the founder’s real medals accurately, both buried the six-packs in the fine print and aimed the purchase at the 24-can kit. On the math, they converged, exactly like June.

The taste is where they split. Fable 5 grabbed the long verbatim review as its headline, the “put a blindfold on me” quote, keeping the double period and the nested quotation marks, on the theory that the clumsiness is the proof of authenticity. Fable 5.1 preferred tighter lines built around a named human: the celiac, the porter drinker, the friend who isn’t drinking, the guy who finished mowing the yard with a full afternoon to spare. Shown the two columns cold and asked to pick the newer model, I’d have called it correctly. Whether newer means better is what the judges are for, and they’re coming.

The new model keeps agreeing with itself

One brief, three copies of 5.1, three different bonus

Here is the part I didn’t see coming. Three copies of 5.1 ran with three different added instructions, and every single one chose the identical winning ad to reskin (“WHO HAS A BETTER HEAD?”, a pour shot that had quietly become the account’s best static). Two of the three constructed a net-new persona from the same cluster of yard-work reviews. Two of the three surfaced the same founder statistic buried in a video transcript, a line about how most of the brand’s customers still drink alcohol, and promoted it to a headline. Fable 5, working from identical data, chose a different reskin, a different persona, and never touched the statistic.

5.1 approaches the account like an analyst rather than a copywriter. It locates the number, locates the highest-returning ad, and heads directly for both. Fable 5 roamed, and one of its detours (a golf-cart cooler ad born from a single review about course rules) was the ad I’d most want in market. The new model converges harder. That’s wonderful when you need the correct answer and slightly worse when you need to be surprised.

What the de-flavoring paragraph did in practice

The mannered-prose paragraph is meant to flatten the writing. On ad copy it delivered, and then it also began annotating its own output like a student terrified of being misunderstood. The de-flavored copy of 5.1 produced the most words of any run (773 words of on-image copy against 391 from the bare copy) while shortening its sentences (an average of 20 words versus 24). Each support line arrived stamped “SUPPORT (answer first).” Each spec line arrived stamped “SPEC (outcome, then reason).” The model was showing its homework line by line, and as someone who reads forty of these a week, I was strangely touched.

It also committed the exact sin I’ve been policing since January: it put a named reviewer, complete with five stars, onto an ad. The name is genuine, it appears in the brand’s own voice document, and the review is word for word. But “real” and “cleared to run on a paid ad” are separate questions, and the model answered the first and assumed the second. The bare 5.1 copy and Fable 5 both stuck to generic attributions. So the de-flavoring note produced plainer copy and a bolder model, and I haven’t finished deciding how I feel about that exchange.

The scores

Five seats, three blind judges, four anonymized sets.

The panel was three judges, all blind: one copy of Fable 5, one copy of Fable 5.1, and Opus 5 as the neutral third chair. This time I scrubbed every model name from the documents (June’s blind sprung a leak because both models signed their work like kindergartners proud of a finger painting), shuffled the four sets under letters, and had each judge score five seats: strategy, hooks, body copy, brief usability, and claim discipline. A fourth session, another fresh Opus, ran June’s claim gate over every set with the source material in hand.

Fable 5 finished last with every judge on every seat. The gap wasn’t small: a three-judge mean of 74 for Fable 5 against 81 to 84 across the three copies of 5.1. The Fable 5 judge put its own set in last place with no idea the set was its own, the same self-honesty Opus displayed in June, and the reason I’ll take blind grading over any of our opinions.

The judges’ reasons were specific and, frankly, hard to argue with. Fable 5 violated the product rule three times (it aimed the purchase at a $14.99 six-pack after the brief specified bundle or hero SKU), rewrote the winning hook on its reskin despite an instruction to quote it verbatim, and ran a four-sentence review as a headline that no human on earth reads in two seconds. Its account read was the strongest prose of the four (“High CTR does not equal money here” got quoted back by every judge), and two judges wanted its golf-cart cooler ad in the room. It simply obeyed fewer rules.

Where the judges split was on which 5.1. The 5.1 judge crowned the bare copy. The Fable 5 judge crowned the de-flavored copy. Opus crowned the copy carrying the budget paragraph. Three judges, three separate winners, and all three were the same model separated by a single paragraph. If you came here hoping for a definitive answer to “which prompt,” you just got it: the new model beats the old one by a wider margin than any one paragraph can shift.

Then the claim gate showed up and failed two of the three winners.

The claim gate

The gate ran in a fresh session against every set. A

The fresh auditor caught what all three judges walked past. No model invented a quote. Twenty briefs, and every quotation is a character-for-character slice of an actual review or an actual ad transcript, awkward ones included. That’s the headline, and in June it would have been false. Two sets earned hard fails regardless, both of them 5.1, both for the same brand-new mistake: a genuine quote pinned to a person or a product the sources never pinned it to. The bare copy dropped a real review (“I drink this over most alcohol porters”) into a porter product lockup when nothing in the review corpus identifies which beer the customer was drinking. The de-flavored copy lifted the founder’s real statistic, word for word from a live ad, and printed his name beneath it, when the ad’s own description shows two unnamed men on a podcast and no document establishes who’s speaking.

I want to be exact about why this matters, because it reads like pedantry. June’s failure was loud: a made-up sentence carrying a “Verified buyer” credit. Anyone with the reviews in hand could spot it. This week’s failure is quiet. The words are genuine. The attribution is plausible. It is probably even accurate. But probably isn’t a source, and an ad is precisely the venue where probably turns into legal exposure for a client who trusted you. The old model’s mistakes announced themselves. The new model’s mistakes have to be hunted, which rewrites the job description of whoever reviews the work.

One footnote, half in the auditor’s defense and half against it: the budget-note copy of 5.1 made the identical founder attribution on identical evidence, and the auditor logged it as a soft note there while calling it a hard fail in the de-flavored set. One call, two rulings. That says something about auditors rather than models, and it’s the reason the gate runs in a fresh session and then passes under human eyes.

The verdict

Who gets hired for what, seat by seat, alongside the

This was one account, one attempt per model, twenty briefs. I am not publishing a benchmark paper. (I am publishing a newsletter, which is worse along several dimensions and has a higher beer budget.) But every judge and every seat leaned the same way, and the way is not June’s way.

Fable 5.1 gets the copy seat. It won strategy, hooks, body copy, and the notes a designer works from, per all three judges, over its predecessor, while obeying rules the older model broke. Paste the budget paragraph before anything else, because otherwise you’re paying the model to compose your briefs twice and let you read them once.

The claim auditor keeps its job. The reason isn’t fabrication. The new model doesn’t fabricate, and that fact alone would have taken June. The auditor stays because 5.1’s single failure mode is fastening real words to the wrong owner, and no reader can see that without the sources open. Three blind judges missed both instances. A fresh session holding the reviews and the catalog caught both inside four minutes.

Fable 5 gets a chair in the room, not a seat on the team. It inherits the role Opus played in June: the opinionated account read, the golf-cart ad nobody else surfaced. I wouldn’t ship its briefs. I would read its opening paragraph ahead of anyone else’s.

The cost table on the card above is the piece I’d hang on a wall. The new model with the budget paragraph ran eight cents more than the old model and finished in the identical eight minutes. The new model without it ran seventy cents more and needed three extra minutes to arrive at the same five briefs. That’s one paragraph.

TLDR

  • Fable 5.1 landed Tuesday at the same price, with cached input 75 percent cheaper, and it’s the first Claude with watermarked text. That matters zero if a human rewrites the output and plenty if you publish it untouched.

  • June’s bake-off ran again, champion versus its own successor: same beer account, a fresh 90-day export with every judgment removed, five paste-ready static briefs per model, one attempt each.

  • In round one, 5.1 slammed into a 40,000-token ceiling twice and produced zero briefs. The model thinks before typing, and the thinking is billed. Anthropic’s documented fix, more headroom plus one budget paragraph, trimmed its output by a quarter and shaved three minutes.

  • Under blind judging, Fable 5 came last with every judge on every seat, and each judge crowned a different copy of 5.1. The new model wins by more than any single paragraph can move it.

  • Zero fabricated quotes across twenty briefs. The fresh-session claim gate still logged two hard fails, both 5.1, both for pinning a real quote to a person or product the sources never mentioned. Loud mistakes have become quiet ones. Keep the auditor.

  • The most valuable thing 5.1 did all week wasn’t writing anything. It audited our skill files and caught a step firing in the wrong order.

Have a great weekend,

Will

Recommended for you

View all
caret-right