Skip to content

44 Agents Tried to Break My Case Studies. Three Succeeded.

AI WORKFLOW
44 Agents Tried to Break My Case Studies. Three Succeeded.
Conner Crowe

Last month I finished five case studies for this site. Real accounts, real numbers, each one traced back to the source report. Before publishing I ran my usual pass: the voice check, the fact check, the read-aloud. It came back clean. Then I pointed 44 AI agents at the same five case studies and told each one to assume the copy was wrong and prove it. Between them they raised 38 findings. Three would have embarrassed me.

The lesson wasn’t that AI makes a good editor, though it does. The lesson was about my own clean pass. I wrote the copy, I checked the copy, I signed off on the copy, and an adversarial reader still found a client-data leak, an overclaim, and a number I couldn’t trace, all in work I had just called finished. Your own review of your own work is the least reliable review you run.

How the 44 agents were set up

The structure is the part that matters, because “ask ChatGPT to check it” doesn’t do this. I split the review into six dimensions: factual accuracy, client confidentiality, voice, internal consistency, overclaiming, and links. For each dimension I ran a set of independent agents, and every one got the same instruction. Refute by default. Don’t confirm the claim, try to break it. Assume a number is wrong until you can’t prove it is. Assume the client can be identified until you have checked every figure.

Forty-four agents in total, each hunting for what wasn’t right, none able to see the others’ work or lean on my confidence that the copy was done. I couldn’t talk them out of a finding, because they never heard me call it finished.

The three that landed

Thirty-eight findings came back. Most were small: a link to tighten, a sentence that read templated, a stat that needed its source line. Three were not small.

One case study about a consent-tracking fix said two platforms had been reporting zero conversions. The account’s own history showed one of them had already been corrected before my window. I had overstated the break. Scoped it down to what was true.

One case study printed a client’s exact cost per click. Anonymized everywhere else, and I had still left a real dollar figure a competitor could use. Changed it to a relative multiple.

One live figure on a results page, a shopping-campaign lift, I couldn’t trace back to a report when the agent pushed me to. If I can’t source it, it doesn’t ship. Pulled it and replaced it with a number I could stand behind.

None of those were lies. Each was the kind of thing that happens when you’re close to your own work and reading for confirmation instead of for holes.

Why your own pass misses them

When you review your own writing, you’re not checking it. You’re re-experiencing writing it. You remember what you meant, so you read what you meant, not what is on the page. You remember tracing the number, so the number looks sourced even where the citation is missing. Confidence is the problem, and you have the most confidence in the work you just finished. Catching an AI tell in your own prose is the easy version of this. Catching a false number you believe is the hard version.

An adversary has none of that. It didn’t write the sentence, it doesn’t know what you meant, and you told it to assume you got it wrong. It reads what is there. That’s why refute-by-default matters more than the model behind it. A friendly reviewer confirms. An adversarial one breaks. Only the breaking finds the leak.

I did it again this week

The home-services cost-per-lead post on this site went through the same thing before it shipped. Three agents, refute by default, and they caught two problems my own pass had waved through. The draft was about to cannibalize an existing page of mine, competing with it for the same search, and I had a number backwards. Both fixed before anyone saw them. I don’t publish account work here without running it now.

What I did not claim

Not that the agents wrote or fixed anything. They found, I decided and rewrote. And not that this makes the case studies perfect. It makes them checked by something other than the person most motivated to believe they were done. That’s a lower bar than perfect, and a much higher one than a self-review.

The receipts

Five case studies, one 44-agent adversarial workflow across the six dimensions above, 38 findings raised and dispositioned, three rated high severity and fixed before publish. The five are live on the results page, and each one carries a “what I did not claim” section. That habit is the same instinct, written into the copy itself: name the thing you’re not saying, so the thing you’re saying can be trusted. A claim that survived an adversary is worth more than one you only reviewed yourself.

Keep going

Free PDF: The Voice Audit Checklist. The by-hand version of the voice dimension, the one I still run first. No email gate.

What’s next

The five case studies at the results page are the ones that came out the other side of all 44 agents. The leaked CPC is gone, the overclaim is scoped down, the number I couldn’t trace is replaced. If you’re deciding who to trust with your accounts, the section to read isn’t the win. It’s the “what I did not claim.” That line is the difference between someone showing you their work and someone showing you only the parts that flatter it.

Want a review like this on your account?

Want this kind of review
on your account?

Thirty minutes on the phone. Same person on the call as on the work. Walk out with a clear set of next steps.