
Hey all, happy Friday,
We spent the long weekend out on Long Island with my sister, her husband, and their two daughters. My dad and my stepmom hosted the whole crew. The drive out happened Thursday night. James slept for the full ride, moved into a crib without so much as cracking an eye, and stayed asleep until morning. I am reporting this because it has never once happened before, and I have no reason to believe it will happen again.
Friday meant the beach, obviously. It was an absolute stunner of a day. James parked himself in the orange bucket, which is the same orange bucket I sat in as a kid, a fact I have brought up before and will continue bringing up until the bucket physically fails. Putting him next to my two nieces is a fascinating exercise in contrast. The girls are methodical. He is a wrecking ball. The older of the two sisters methodically opened an ice cream stand made of a plastic ice cream cone and sand, and James marched over saying "ice cream, ice cream," grabbed her ice cream, and walked off. Shortly thereafter, we determined it was time to take him into the ocean, where his standing policy is to stay in the water until his lips turn blue, then be outraged when someone carries him out mid-shiver.
Back at the house, my dad and stepmom had converted the attic into a playroom, and James and the younger cousin passed a solid hour sprinting and launching themselves into a corner full of stuffed animals while the older cousin worked on a puzzle. James contributed by attempting to eat the pieces. She handled it like an absolutely incredible older cousin and was very patient with him. What’s more, when they were not playing, she ensured that James never waddled down the stairs on his own.
He also got halfway through Toy Story 4 because his cousins had it on, which stands as a personal record. His attention span, as should be obvious by now, is not where he shines.
Linus went full maniac out there. He wanted to be outdoors every waking second, because bunnies exist, and since James's two favorite words all weekend were "outside" and "ball," the two of them formed a united front.
Sunday brought a second beach day, this one cloudier, with more orange bucket time and more toys that did not technically belong to him. Monday we made the drive back, nailed the timing, and got a rare outdoor lunch, just the four of us, at Bar Toto in Park Slope. You should go if you have never been. After that, we hit the playground, which has a water feature, which is where James spent every last minute.
I am confident somewhere under his tummy he is hiding gills.
Anywhooo.
Two launches landed five days apart. One of them made all the noise.
OpenAI released GPT-6 Astra on September 3. Greg Brockman announced we are now in the AGI era. The benchmark charts basically went vertical. By lunchtime, every feed I follow was nothing but computer-use demos, the model filling out spreadsheets and clicking around websites like an intern who skipped lunch. The list price matches the Claude model from last week's issue (ten dollars in, fifty out, per million tokens), the context window clears a million tokens, and the launch post covers spreadsheets, exploits, and math benchmarks. Writing appears zero times in it. I checked.
Then OpenAI released ChatGPT Images 2.5 on Monday, September 8. Two API models, Flare and Sunburst. The token prices match GPT Image 2. The claims: sharper detail, closer fidelity to reference photos, edits that only touch what you asked them to touch, up to half the wait. It hit fal.ai that same day. It hit my feed basically nowhere.
The gap between those two receptions matters if you make ads for a living. Astra is a model you talk to. The image model renders what you actually put in front of customers. Every static image our agency has shipped since April has come out of GPT Image 2 (that was the month I held Nano Banana's funeral, and some of you attended). A better renderer means every ad we make changes. A quietly worse renderer also means every ad we make changes. So I ran the only kind of test I trust: the model's actual daily job, using copy I already had.
Quick housekeeping before anything new, because this is now the third issue riding on one account. Back in June, Anthropic dropped Fable 5, and I pitted it against Opus 4.8 using Go Brewing's raw Meta export. Fable swept every seat, and the auditor flagged the other model for fabricating a customer quote. Then last week, Fable 5.1 showed up, so I ran it against Fable 5 on the same account with a fresh export. The newer model took every seat with every judge, no quotes got invented this time, and the only things the auditor flagged were genuine words attributed to the wrong person. This week I am back on the same account, with the same judging harness and the same claim gate. From issue to issue, the models change, their failure modes change, and everything else stays put.

The verdict card from last week, if you skipped it: Fable 5.1 stole the copy seat from Fable 5, the claim auditor earned its keep, and the budget paragraph shaved three minutes and a quarter of the tokens.

Two launches, five days apart, and the quiet one is the renderer. Four numbers frame the test: 30 renders, 0 typos on the new models, 60 blind judgments, roughly a dime per render. All five ads run along the bottom.
The bakeoff
Fable 5.1 wrote five paste-ready static briefs last week for Go Brewing, the non-alcoholic craft brewery whose account I have now borrowed for three straight issues. The formats were deliberately varied: a seven-element offer stack, a testimonial quote card with one phrase highlighted, a twelve-element Us vs Them grid, a whiteboard Venn diagram drawn in handwritten marker, and an iOS reminder card hovering in front of a fridge stocked with six different cans. The copy stayed exactly as written. It is the same copy the blind judges scored last week.
Every brief became a single prompt. That prompt is the one our production scripts have carried since July: a composition block explaining that Instagram's interface will eat the top 27 percent and bottom 27 percent of the canvas, a brand block carrying the typefaces and the exact orange, a copy block instructing "render EXACTLY these text elements and nothing else", the scene, and a closing final-checks paragraph. Three reference images rode along with each prompt: the brand spec card, the font card, and one real product render.
The identical prompt, references, canvas (1088 by 1920), and quality setting ("high", the ceiling GPT Image 2 supports, so no model got a knob the others lacked) then went to three fal endpoints: openai/gpt-image-2/edit, openai/gpt-image-2.5/flare/edit, and openai/gpt-image-2.5/sunburst/edit. Two attempts per ad. Thirty renders total. All thirty came back on the first attempt. Zero errors, zero refusals.
I need to flag one fairness issue before any numbers. That prompt spent seven rounds in July being tuned specifically against GPT Image 2. The line-A line-B phrasing, the entire "the empty bottom strip will feel wrong and that is correct" paragraph, all of it only exists because GPT Image 2 kept welding the button to the canvas floor. The 2.5 models received a prompt built for the model they replaced. Everything they did, they did with a handicap.
The timer
Going in, I had three predictions. Speed did not make the list.
GPT Image 2 averaged 96 seconds a render, with the slowest at 117. Flare averaged 35. Sunburst averaged 36. OpenAI promised "up to 50 percent lower latency" and, on this workload through fal, the new models finished in roughly a third of the old time. The genuinely surprising part: Sunburst, which OpenAI positions as the precision model and which launch coverage pegged at one and a half to two times slower than Flare, ran no slower than Flare at the high setting. Not once in the whole set. Its fastest render took thirty-two seconds and its slowest took forty-four.
Scale that up to a batch. A thirty-ad wave used to mean 48 minutes of render time. It now means 18. That is the gap between kicking off the wave and leaving for coffee, versus kicking it off and just sitting there while it finishes. For someone whose job is grading renders, that turns out to be an actual workflow shift rather than a nice-to-have.

Wall-clock seconds per render, averaged across ten renders apiece: 96 for GPT Image 2, 35 for Flare, 36 for Sunburst. The same Us vs Them ad appears from each model with its time below it.
The typo audit
Back in April, the top reason a render got binned was a typo. "You'lll." A mangled "Naperville." A price sprouting an extra digit. Our grading rubric literally contains the line "typos are the number one render failure", and every grader agent I run checks the pixels against the locked copy letter by letter.
Two blind judges ran that letter-by-letter check on all thirty renders. Sixty total judgments. These were not friendly ads, either. One carried a twelve-element cross-versus-check grid. One carried an accented word (CLICHÉ) plus curly quotation marks. One was entirely handwritten dry-erase marker. One featured an iOS card with a blue "Okay" button and six distinct can labels in the background. On the two new models, across all of it, the judges logged zero copy errors. Forty judgments, forty clean.
The old model took two soft misses out of twenty, and neither one was a spelling error: the orange highlight on the porter quote quit one word too early, and the support line on the Venn ad ducked behind the glove holding the can.
I want to size this claim honestly. It is enormous. Six months ago I devoted a full issue to GPT Image 2 being the first model that could get a headline out without a typo most of the time. The re-roll budget existed precisely to cover "most of the time."
If "most of the time" has quietly turned into "every time," then a whole piece of the grading pipeline this newsletter has documented since spring is hunting a failure that no longer occurs.

Clean-copy rate over twenty blind judgments per model: 90 percent for GPT Image 2, 100 percent for both 2.5 models. Each model's offer-stack ad is shown, seven text elements apiece, every letter correct on all three.
The failure that survived
So what did the judges actually penalize? The exact same thing they have penalized since July. Button placement.
Instagram lays its own interface over the top 14 percent and the bottom 18 percent of a 9:16 ad. Any text inside those bands is invisible to a scrolling human. Our prompt states this in pixels, twice, and then restates it in the closing paragraph. Every model drifts low regardless. GPT Image 2 dropped text or a button into the dead zone on 10 of 20 judgments. Flare did it on 7 of 20, with the fridge ad as the serial offender: both attempts jammed the customer quote down to 89 percent of the canvas, directly beneath the bar. Sunburst did it on 2 of 20, both times on the seven-element offer stack, where one attempt parked the button at 86 percent.
On its passing ads, Sunburst placed the bottom element between 68 and 71 percent of the canvas. That is exactly the target the prompt specifies. No model before this one has hit that target without a re-roll.
The ruler remains the entire job, though. A ten percent dead-zone rate is a re-roll budget rather than zero, and the model we currently ship on sits at fifty.

Hard-fail rate, defined as text or a button inside Instagram's bars: 50 percent for GPT Image 2, 35 percent for Flare, 10 percent for Sunburst. One GPT Image 2 render appears with both bars painted over it, and the customer quote sits fully beneath the bottom one.
The tally
Every render went before two judges, Claude Opus 5 and Claude Fable 5.1, filed under a random six-character id. Neither judge saw a model name. Neither judge saw the other's marks. Each one transcribed the text, tallied errors, estimated the top and bottom text pixels as a percentage of canvas, graded the can label, scanned for faces (zero appeared, on any render, mercifully), and answered the single question that counts: would you hand this to the client untouched, on a scale of zero to one hundred, where ninety and above means no edits needed.
Three numbers came out.
GPT Image 2: 36. Flare: 50. Sunburst: 70.
The two judges landed inside nine points of each other on every model and inside one point on Sunburst (Opus gave 70, Fable gave 70). They also converged on the cause. Nearly the whole gap traces to the dead-zone failures. A render whose quote sits under Instagram's bar scores close to zero regardless of how pretty it is, and the old model produced twice as many of those as Flare and five times as many as Sunburst. Set those failures aside and grade layout and craft alone, and the field tightens: 71, 77, 80. On label fidelity, meaning does the rendered can match the real can, the three sit basically level at 7.5, 7.7, and 7.8 out of 10, which lines up with what my own zoom-in showed: all three recreated the porter can down to the tiny cap on the label and the "dark brew with rich vanilla" line, and neither detail was anywhere in my prompt.
Sunburst took every ad except one. Its single loss, the Us vs Them grid, went to Flare by three points.
I realize I have now spent four paragraphs on which robot painted the better beer, and that this qualifies as either extreme rigor or extreme embarrassment. Both, most likely.

Would-ship scores, averaged across both judges: 36, 50, 70. The per-judge rows show Opus 5 and Fable 5.1 within nine points everywhere. The per-ad grid marks which ads each model botched.
What the quality settings actually buy
Do the pricier settings change anything real? GPT Image 2.5 stacks two tiers above high, named xhigh and max, and I wanted to see where the money actually goes. So the offer-stack ad went back through Flare at "low" (the basement) and Sunburst at "max" (the penthouse), sitting beside the "high" render the main test used.
Every word came out letter-perfect at all three tiers, including at low, which would have read as science fiction in April. What low lost was the brand itself. The heavy condensed headline face from the font card came back as a wide generic sans. The subhead came back in a rounded system font. The fine print shrank past anything a phone could resolve. The words made it. The typography did not. At max, which took 61 seconds instead of 32, I cannot spot a difference on a phone screen. High is the setting.

The identical ad at low, high, and max. All three tiers spelled every word right. Low dropped the brand typeface, which the zoom strip along the bottom compares side by side. Max needed twice high's time and looks identical to it.
Try test one against your own brand. The kit holds the render harness, the locked prompt blocks, and the exact grading prompt I used. You get five briefs, three models, thirty renders, and a bill of about three dollars. Grab the kit here:
The bug I shipped myself
One more finding, because it cost me a full re-run and it will cost you one if you copied our scripts from past issues.
Our production scripts request fal's size preset named portrait_16_9. I assumed that resolved to 1080 by 1920, since that number sits in every prompt we have written since April. It does not. All three endpoints returned a 608 by 1088 image. My entire first sample came back at that size, and I examined three of them, complimented the typography, and noticed nothing. Thirty renders, zero typos, wrong dimensions. I am a professional.
The whole wave got re-run with the canvas passed explicitly, 1088 by 1920 (fal insists on multiples of sixteen), and those re-runs are what appears in this issue. Go check your output size. I had been reciting pixel coordinates to a model in a prompt while asking it to render at one third of them. Since July. On paying client work.
Then I got a key
At two o'clock Wednesday, one hour after writing the paragraph above about not having an OpenAI key, I quit being precious and made one. So the section that would have lived here, the one about the test I could not run, no longer exists, and the test itself sits here instead. It grew.
The brief grew with it, because a rematch across five ads is a coin flip and I wanted a map instead. The swipe file our iteration tracker maintains holds seventeen canonical static formats, running from the boring end (headline, offer, testimonial) to the ones the market digest currently has on watch (a comic strip, a Venn diagram, a cyclical narrative, whatever that turns out to be). Twenty-three formats altogether. The new brief demands ONE net-new static per format: a persona and angle the account has never touched, anchored to a real review or a real Reddit thread, with a render-ready copy block and scene written in the precise grammar our image prompts speak, plus a claim trace beneath each one. Both writers got last week's raw account export, the same brand law, the same registry of everything previously briefed, and one fresh document listing the seventeen product renders that actually exist, so nobody scripts a can we cannot draw.
Two writers, identical documents, identical brief, one attempt each, matching output caps: Claude Fable 5.1 and GPT-6 Astra.
The copy rematch
Astra wrapped in seven minutes on 29,000 output tokens. Fable 5.1 needed fourteen minutes and 79,000, with 28,000 of those spent thinking. Both handed over all twenty-three briefs, and both produced a machine block that parsed on the first pass, which beats most humans I have ever briefed.
I handed both writers the budget paragraph from last week's kit, the one that warns the model that its reasoning and its output draw from a single pool. That paragraph is the only reason Fable 5.1 crossed the finish line: last week, running without it, two copies of 5.1 slammed into the ceiling and produced nothing. This week, with it in place, Fable 5.1 still burned 28,000 tokens on thinking before typing anything. Astra burned 2,000. I am not knocking either model. I am telling you this because it is the sharpest illustration I have of how two models spend the exact same budget in completely different ways.
Then came the blind judging, and this round I patched the thing that nagged me last week: every judge had been a Claude. This time four judges sat, two per vendor: Fable 5, Fable 5.1, Opus 5, and Astra itself. Each received the identical source documents plus both anonymized sets and graded five seats: slate and strategy, hooks, body copy and voice, render readiness, and claim discipline. Two claim auditors sat on top of that, one Opus and one Astra, in clean sessions holding nothing except the sources.
Three of the four judges sent the Fable set forward. The fourth judge was Astra, and it forwarded its own work.
The three Claude judges stayed within ten points of one another on both sets (Fable 5.1's set: 79, 68, 74; Astra's set: 69, 63, 65). Astra graded its own set at 76 and Fable's at 44, a 25-point spread pointing the opposite way from everyone else. I will not pretend to know whether that is a model recognizing its own sentences or a model with genuinely divergent taste. I will show you the seats and let you call it.
Among the three judges who authored neither set, Fable 5.1 carried slate and strategy (80 to 63), hooks (69 to 44), body and voice (76 to 59), and render readiness (80 to 73). Astra carried exactly one seat, claim discipline, and carried it decisively: 89 to 65.
The hooks number tells the entire story by itself. Judges run each headline through nine binary tests, and the Fable set averaged between 5.6 and 7.0 out of 9 depending on the judge, while the Astra set averaged 3.6 to 5.0. Here is what that gap looks like on the page. Fable's headlines: "THE $15 MOCKTAIL. OR A SIX PACK FOR $14.99." "FOR PEOPLE WHO KNOW WHAT TETTNANG IS." "Cooler cleared the starter shack. Back nine still ahead." Astra's headlines: "WHERE DID THE PREP GO?" "WHAT WAITS AFTER THE SHIFT?" "WHICH SHELF IS THE PILSNER SHELF?" Astra phrased eighteen of its twenty-three headlines as questions. Fable phrased two.
The claim seat tells the whole opposite story. Both auditors handed Astra's set zero hard fails. Zero. Across twenty-three briefs, every quotation matched character for character, every price traced back to a catalog line, and no award appeared anywhere (a choice the Opus auditor labeled "the single smartest decision in the batch"). Fable's set drew three hard fails from the Opus auditor and six from the Astra auditor, most of them one error committed twice: it priced a 24-pack of Burn It Down at $59.99 when the catalog only carries that beer in six-packs, and it invoked my own asset document's "where the catalog lists one" clause as though the clause were a permission slip. The Astra auditor additionally flagged a verbatim review rendered in ALL CAPS as an altered quote, which is exactly the strictness I want from an auditor and exactly what I would not want from a writer.
So the June pattern holds, and last week's pattern holds, and now the pattern crosses vendors: the stronger advertiser plays looser with facts, and the safer writer turns in twenty-three questions. Every judge, Astra included, gave the same verdict on the ideas room: the Fable set belongs in it. Every judge, including the three who chose Fable, gave the same verdict on the remedy: fifteen minutes correcting two prices before anything renders.
Two echoes from last week showed up that I did not engineer and honestly could not have. Last week, three separate copies of 5.1, all working from identical data, landed on the founder's line about most of the brand's customers still drinking alcohol, and two of them spun a persona out of the yard-work reviews. This week, working from a different brief with twenty-three formats to fill, Fable 5.1 marched right back to both wells: the statistic turned into its Statistics ad, and the mower turned into its Before and After. Astra had the same data in front of it and ignored both. The new Claude is not a wanderer. It locates the richest vein in the pile and keeps mining it, which is precisely what last week's convergence card said would happen and precisely why running it twice gets you repeats.
Astra owns the second echo. Last week, the de-flavored copy of 5.1 dropped a real named reviewer, five stars and all, onto an ad, and I pointed out that "real" and "cleared to run on a paid ad" are two different questions. This week Astra pulled the identical move on its Testimonials brief: an actual customer name, lifted word for word from the brand's own voice document, sitting right under the quote. Both auditors waved it through, because nothing about it is fake. I still would not ship it without the customer's permission, and you should not either. The trap changed vendors. It is still a trap.

Four blind judges, two per vendor. Three forward the Fable 5.1 set; Astra forwards its own. Seat by seat, Fable takes strategy, hooks, body, and render readiness; Astra takes claim discipline 89 to 65 with zero hard fails from either auditor.
Go through both writers' briefs on your own. The kit contains all forty-six of them, untouched, sitting beside the four judges' verdicts and both claim audits, which means you can push back on the judges while the receipts sit right in front of you. Grab the kit here:
The format map
The copy test settled who writes better. It left open the question I actually woke up with, which was whether the new renderer is better, and better at what. Five ads is a coin flip. So all forty-six briefs went through both renderers, GPT Image 2 and 2.5 Sunburst, identical prompt, identical three references, one attempt apiece. Ninety-two renders. Three blind vision judges this round, Opus 5, Fable 5.1, and Astra, each grading every render against its locked copy: transcription, safe zone, label fidelity, faces, plus two fresh scores, format execution (does this actually read as a comic strip, a Venn diagram, a news screenshot) and would-ship.
Sunburst: 68. GPT Image 2: 50. All three judges concurred, the OpenAI judge included (Opus 71 to 50, Fable 67 to 53, Astra 65 to 47).
Sunburst took seventeen of the twenty-three formats. The widest margins showed up wherever the layout carried the most cargo. Offer / Promotion jumped from 4 to 70, because the old model shoved the button under Instagram's bar on almost every attempt and the new one did not. Product Progression jumped from 1 to 66. News, bullet points, comic strip, UGC, before-and-after, and press all climbed thirty to forty points. Those are the formats where the old model handled the crowding by pushing everything downward.
GPT Image 2 kept six formats, and five of them are the leanest layouts on the sheet: the packshot, the headline, the testimonial (dead tie at 79), the founder quote, and social proof. Give the old model three text elements and a can and it does fine. Sunburst's one genuine defeat came on the cyclical narrative, 51 to 73, where it drew the prettier loop and dropped the words in the wrong spots.
One format collapsed on both models: the six-can multi-product grid, at 5 and 6. Six cans, six labels, a headline, a support line, and fine print will not squeeze between 27 and 73 percent of a phone screen regardless of who renders them. That is a brief problem rather than a model problem, and it is the identical problem the tracker's own format notes have flagged since July.

Would-ship by format, both writers' briefs, three blind judges. Sunburst takes 17 of 23. The largest gains sit on the element-heavy formats; the old model keeps the lightweight ones.
The twist
Now the part I never saw coming, and it recasts the copy result.
The copy judges said Fable 5.1 wrote the stronger briefs. The pixel judges said Astra's briefs came out of the renderers looking better: 64 to 54 across both models. Identical briefs. Different question.
The explanation lives in the grammar. Fable's briefs ran sharper and denser. It positioned several headlines above the floor the prompt defines, packed more elements into each ad, and dropped check and cross glyphs into rendered copy. Fed to the old renderer, those briefs landed text under Instagram's bar on 43 percent of renders. Fed to Sunburst, the identical briefs failed 12 percent of the time and scored 66. Astra's leaner briefs suited either renderer: 58 on the old one, 69 on the new one.
Which means the better writer depends more heavily on the better renderer. And across all ninety-two renders, exactly one species of hard failure appeared: the lowest text element sinking below 82 percent of the canvas. No typos. No faces. No fabricated labels. Six months ago that failure list ran four entries long. It now runs one, and that one is a geometry problem, which makes it a prompt problem, which makes it my problem.

All four writer-by-renderer pairings. Fable's briefs climb from 42 on the old renderer to 66 on Sunburst. Astra's climb from 58 to 69. Every single hard fail was a dead-zone miss.
Point the sweep at your own account. The kit includes the twenty-three-format brief, the asset list template that keeps a model from briefing a product you cannot render, and the two-lane runner. You just drop in your own data and your own keys. Grab the kit here:
What we do differently on Monday
Change the renderer. Every static runs on Sunburst starting Monday. It costs the same tokens as our current model, waits a third as long, produced zero typos across forty blind judgments, and fails the dead zone one fifth as often. Flare handles drafts and thumbnails where the layout never ships, since it runs at the same speed and I have yet to find a job it wins.
Keep measuring. The safe-zone grader survives. The letter-by-letter pass survives too, for now, until a few hundred error-free renders pile up and I trust the model enough to demote the check. Ten percent dead-zone is a re-roll budget, not a retirement party for the ruler.
Rewrite the prompt. The composition block was engineered to fight a model that no longer lives in our pipeline. Sunburst hit the exact bottom-element target on every passing ad, which suggests the "aim at 65 percent and it overshoots to 73" arithmetic baked into our prompt may now be dragging it lower than necessary. And since the sole remaining failure in ninety-two renders lives at the bottom edge, the real fix is a hard cap on element count per format before rendering, not one more paragraph of begging.
Leave copy with Claude, plus a gate. Fable 5.1 writes the ad I want to run, and Astra writes the ad I can ship without calling a lawyer. Until a single model manages both, the pipeline reads: Fable drafts the brief, a fresh-session claim gate runs before render, and someone spends fifteen minutes checking prices. Astra earned the second auditor chair outright: it graded harder than anyone in the room and it was correct about the six-pack.
Put the budget paragraph in. I gave you this advice last week, and now I have a second data point backing it up. The kit includes it again. Skip it, and long jobs on 5.1 can die before a single word appears; include it, and the model still thinks for ages, but the job actually ends.
Set the canvas explicitly. Pass the size by hand. Read the sidecar. This one embarrasses me. (Three months. Client work. A prompt saying 1920 to a model drawing 1088.)

Move the renderer to Sunburst, keep the safe-zone ruler, rework the composition block, patch the size preset in the script, and the Astra copy test happened: three of four blind judges forward the Claude set.
TLDR
This makes three issues on one account: June (Fable 5 beat Opus 4.8), last week (Fable 5.1 beat Fable 5), this week (Fable 5.1 beat Astra on copy, and the renderer changed).
GPT-6 Astra shipped September 3 and ChatGPT Images 2.5 (Flare and Sunburst) shipped September 8. The quieter launch is the one that matters for ad production.
The test: the same five Claude-written briefs, one prompt, identical references, three renderers, thirty renders, two blind judges.
Speed: GPT Image 2 needs 96 seconds a render, Flare needs 35, Sunburst needs 36. A thirty-ad wave falls from 48 minutes to 18.
Spelling: the new models logged zero copy errors over forty blind judgments. The old model missed twice out of twenty.
The surviving failure is the safe zone: text under Instagram's bars on 50 percent of GPT Image 2 renders, 35 percent of Flare's, 10 percent of Sunburst's.
Blind would-ship from both judges: 36, 50, 70. Sunburst takes over Monday. The grader stays.
Round two, once the key existed: twenty-three net-new briefs per writer, one per swipe-file format, Fable 5.1 against Astra on identical data. Three of four blind judges (two per vendor) pick the Claude set. Fable takes hooks 69 to 44; Astra takes claim discipline 89 to 65 with zero hard fails.
Ninety-two renders of those briefs: Sunburst 68 to GPT Image 2's 50, all three judges aligned. Sunburst takes 17 of 23 formats, widest on the element-heavy ones. Every hard fail was bottom text under Instagram's bar. Zero typos anywhere.
Have a great weekend,
Will