Everyone’s optimizing content for AI visibility. New research shows the cost
Generative Engine Optimization has develop into the new website positioning quick. Every content group needs their work cited inside ChatGPT, Perplexity, and AI Overviews, and a rising trade of instruments guarantees to assist. Most of the recommendation treats this as a single-round recreation: tweak a doc, examine if it will get cited, repeat.
A new paper accepted to COLM 2026 asks a special query. What occurs as soon as each creator in a market runs that very same playbook, again and again, in opposition to the similar AI rating sign?
What the researchers really constructed
Researchers from UC Berkeley and Zhejiang University constructed CHASE, a simulation that iterates 4 phases throughout 20 rounds:
- Rank
- Discriminate
- Rewrite
- Evaluate
Documents that rank properly keep as they’re. Documents that miss get rewritten towards no matter options the present spherical rewards.
The setup runs throughout six domains, three advice classes (retail, video video games, books), and three question-answering classes (net, information, debate).
To keep away from a mannequin favoring its personal writing type, the group cut up the roles throughout three separate mannequin households: Gemini 3.1 Flash-Lite ranks paperwork, OpenAI’s GPT-5.4-mini rewrites them, and Claude Haiku 4.5 judges high quality independently.
Before trusting the simulation, the researchers checked whether or not their rating sign really predicts actual quotation habits in AI-generated answers. It does, with a rank-citation AUC of 0.853 throughout all six domains, that means a doc ranked greater is very more likely to be the one an AI system really cites.
The core discovering: profitable drifts away from high quality
Across all six domains, the alignment between “hits the options that win the rating” and “scores properly on an unbiased high quality judgment” received weaker each spherical.
The decline ranged from a gentle 0.018-point drop in Books to a pointy 0.107-point drop in Web content, averaging 0.068 throughout the board.
That quantity wants a plain-English translation. Early on, paperwork that ranked properly and paperwork that have been independently judged good have been largely the similar paperwork. By spherical 20, that overlap had shrunk in each single area examined.
Here is the element price sitting with: total doc high quality barely moved, and even rose barely in the retail area. Ranking success merely stopped being a reliable signal of it. Optimizing for the ranker and optimizing for the reader had cut up into two completely different jobs.
Ruling out the apparent rationalization
A skeptical reader’s first query ought to be whether or not that is simply what occurs when AI rewrites textual content repeatedly, no matter any rating sign concerned.
The researchers examined this immediately.
In what the researchers name a frozen-document management, paperwork stayed unchanged whereas rating continued, and the inhabitants stayed the similar spherical after spherical, confirming that static content stays put by itself.
In a random-target management, the place the rewriting mechanism stayed, however the goal options have been chosen at random as a substitute of pulled from the rating sign, the high quality hole opened up much more slowly than in the actual experiment.
The drift traces particularly to chasing what the ranker rewards, slightly than to rewriting itself.
The researchers additionally audited 3,472 accepted rewrites for integrity. Roughly 93% handed each examine cleanly, with fabricated claims displaying up in a mere 3.4% of instances. The move fee held regular throughout all 20 rounds.
Whatever is driving the quality-ranking cut up, sloppy AI writing degrading the corpus is a poor rationalization for it.
The sample seems completely different relying on the place you compete
- Structural convergence. In retail, a small set of structural options, issues like phrase rely and formatting, more and more outline what wins. Winning paperwork begin to look alike.
- Signal instability. In debate content, the rating itself proved comparatively unstable spherical to spherical, with much less consistency wherein options predicted success.
- Feature dominance. In information, rating success concentrated round one or two options carrying disproportionate weight, a brittle setup if that specific ranker’s preferences shift.
What this implies for content and advertising groups
- Track the hole alongside the win. A rising distance between “hits the present rating profile” and unbiased reader or buyer suggestions is the main indicator right here, and it shows up properly earlier than any seen high quality downside does.
- Know which regime your class sits in. A content technique constructed round one or two dominant options works solely so long as the ranker retains rewarding those self same options. Diversifying the alerts a bit of content depends on reduces that publicity.
- Keep the writing-quality query separate from the incentive query. The audit right here shows AI-generated rewrites can maintain up mechanically wonderful whereas the underlying incentive nonetheless reshapes what will get rewarded behind the scenes.
- Treat any single rating sign as a shifting goal, even whereas the underlying mannequin stays fixed. The ranker on this research stayed operationally stable and fully mounted for all 20 rounds, and the ecosystem nonetheless drifted purely from creators adapting round it.
The sample echoes a theme that shows up throughout the mistakes AI leaders keep making with agentic deployments: a system that appears steady in a demo or an early spherical can nonetheless drift as soon as actual incentives and actual opponents get entangled.
The uncomfortable a part of this research is how acquainted it feels. Search Engine Optimization went by means of the similar arc: early wins for genuinely helpful content, adopted by a gradual drift towards no matter the algorithm occurred to reward that 12 months.
CHASE suggests AI-visibility optimization is on the similar monitor, simply shifting quicker and with far much less visibility into what the ranker really needs.
