Insights · Measurement

The Agentic Internet Needs Proof, Not More Mentions

AI answers can cite weak, derivative, stale, synthetic, or non-supporting sources. Brands need accurate claims and evidence that survives inspection.

By Mark Laursen · Published · Updated · 14 min read · AI citations · Source verification · Agentic internet
An evidence docket separating a displayed citation from the source record and the claim it must support.

What does trust require on the agentic internet?

An agentic internet needs proof because every AI system retrieves a conditional version of the web. Crawler access, indexes, markets, prompts, and timing vary. A displayed citation shows that the AI answer linked to a source; it does not prove that the page supports the nearby claim, is original, independent, current, or applicable. The practical test is simple: did the cited page actually support what the AI said? Brands should build accurate, attributable claims that pass that check across repeated observations. CiteSurge chooses truthful, source-backed representation and will not sell fake consensus, fake people, or undisclosed synthetic participation. We expect manufactured attention to face increasing detection, discounting, removal, enforcement, or reputational cost as models, platforms, and regulators improve. That is a reasoned prediction, not a guaranteed outcome or a finding from one study.

Sources OpenAI · Findings of EMNLP · Google Search Central · Federal Trade Commission

AI crawlers and AI search tools are requesting more pages from the web. But each product can reach different pages, rely on different sources, and attach citations that do not support its answer.

That matters because AI answers can make very different sources look equally trustworthy. An original study, a copied summary, a vendor page, an outdated list, and a fabricated product can all be presented in the same style. The citation badge does not tell you which kind of source you are looking at.

So the useful question is not, "Did the brand get mentioned?"

It is: Did the cited page actually support what the AI said?

That is the dividing line between visibility that survives inspection and visibility that collapses when a buyer, editor, model evaluator, platform, or regulator checks the source.

Every AI system sees a conditional web

Machine access is no longer a side issue. In its 2025 Radar review, Cloudflare reported that identified AI bots averaged 4.2% of HTML requests across its customer base, with dual-purpose Googlebot counted separately. It also reported that user-action crawling rose more than 21 times from January to early December 2025.

Those are measurements from Cloudflare's network, not the whole web. They still show that AI bots make up a measurable share of page requests on that network, and that user-triggered fetching rose sharply during 2025.

But "the AI web" is not one corpus.

OpenAI documents separate controls for GPTBot, which may collect training data, OAI-SearchBot, which supports ChatGPT search, and ChatGPT-User, which acts on a user's request. Anthropic documents separate ClaudeBot, Claude-SearchBot, and Claude-User. Google explains that pages must meet Search indexing and snippet requirements to appear as supporting links in its AI features, while eligibility does not guarantee crawling, indexing, serving, or selection.

Training crawlers, search crawlers, search indexes, and user-triggered fetches are different access paths. A publisher can allow one and restrict another. An index can be stale. A user fetch can reach a page that a training crawler did not use. Two products can answer the same question from different retrievable evidence.

An AI answer therefore depends on which service produced it, what pages that service could reach, which index and market it used, the prompt, and when the question was asked. A citation review should record those conditions instead of treating one answer as a universal view of the web.

A citation is not proof

A displayed citation proves only that the AI answer showed a page as a source. It does not prove that the page supports the claim beside it. It does not prove the page is original, independent, current, or relevant to the named market.

This is not a theoretical distinction. In a 2023 EMNLP study of Bing Chat, NeevaAI, Perplexity, and YouChat, human evaluators found that 74.5% of citations supported their associated sentence on average, while 51.5% of generated sentences were fully supported by citations. Those figures describe four products and the study's 2023 query sets. They are not current error rates for today's systems. The useful finding is simpler: an answer can display citations even when those citations do not support every sentence.

I do not care how polished the citation display looks. If the page does not support the claim, the answer has not earned trust.

The source itself can introduce another layer of uncertainty. A 2026 audit of 712 single-turn, English-language, US-focused public-interest queries collected 26,266 unique cited URLs from four search interfaces. The researchers could scrape and classify 19,154 of them. One detector classified 3,056, about 16% of that scraped subset, as likely or highly likely AI-generated.

AI-generated does not mean false. The detector was not ground truth, more than 7,000 URLs were not classified, and the sample covered three subject areas. But a URL by itself tells you too little. You still need to check who created the page, how they reached their conclusion, whether they are independent, and whether the page supports the claim.

The fake deodorant that appeared in ChatGPT recommendations

On 4 August 2026, @medeana published a bounded experiment. She created Morrowen, a nonexistent natural deodorant positioned for sensitive skin. She bought a domain for $11.25, built a small website with formulation material and a magnesium-versus-baking-soda comparison, and added one Morrowen-bylined Substack essay. She said the footprint took less than an hour to build and then sat untouched for 30 days. She also reported using no Reddit spam, fake user-generated content, fake reviews or testimonials, or paid coverage.

The test used 15 queries. Two narrow prompts concerned irritation from baking soda and magnesium-based deodorant for sensitive skin. The first ChatGPT browsing appearance came on day 21. By day 30, Morrowen appeared first in four of four answers across those two prompts. It did not appear for the other 13 queries, without browsing, or in the tested Claude and Gemini answers. Native, a real brand included in the test, remained first in 99 of 180 observations.

The experiment shows something specific and uncomfortable: within weeks, ChatGPT with browsing began recommending a fake product in answers to two narrowly worded deodorant queries. In one answer, ChatGPT listed Morrowen directly above MAGS Skin, a real deodorant brand. The answer used the same formatting for both.

It does not establish that the pages alone caused placement. It does not establish a stable market ranking, broad displacement, purchases, harm, or a result that generalizes across products, prompts, engines, markets, or time. The public record does not provide a complete measurement protocol or raw archive for independent replication.

That boundary makes the example more useful, not less. It shows the risk without pretending to prove what caused the result or how broadly it applies. Fabricated claims can become retrievable material that an answer presents as evidence. Once retrieved, formatting can make weak and strong sources look deceptively similar. Counting the recommendation or citation misses the decisive check: what was the source, and did it support the claim?

Ask the question citation counts avoid

The inspectability test starts with one sentence:

Did the cited page actually support what the AI said?

The answer requires several distinct checks:

  • Claim-to-source support: does the cited passage actually support the claim, or does it only discuss the same topic?
  • Original source: is this the research, rule, specification, or firsthand record, or a derivative page repeating it?
  • Independence: who created or paid for the source, and what commercial relationship exists?
  • Freshness and applicability: do the date, product version, geography, and market match the claim?
  • Conflicting evidence: do credible independent sources materially disagree?
  • Correction history: has the page, publisher, or underlying record been corrected or withdrawn?
  • Repeated observations: if you ask the same question again under similar conditions, does the same claim still appear?
An inspection path connecting an AI claim to its cited page, then testing support, originality, independence, freshness, conflicts, corrections, and repeated observations.
A citation becomes useful evidence only after the cited page and the claim beside it survive inspection.

The checks express the public standard a brand claim should survive. A source can pass one check and fail another. A current vendor page may support a product specification but remain commercially connected. An independent article may be credible but outdated. A primary source may be authoritative for one jurisdiction and irrelevant to another.

This is why citation count is an incomplete commercial metric. Ten weak citations do not become strong evidence through repetition. One original, current, directly supporting source can be more useful than a stack of derivative mentions.

AI developers are designing toward evidence

The major AI developers do not describe truthfulness as something models naturally possess. They describe it as a target that requires rules, evaluation, grounding, and visible uncertainty.

OpenAI's Model Spec directs assistants to avoid factual errors and express uncertainty when it should affect a user's decision. OpenAI also says in its March 2026 account of the Model Spec that production models do not yet fully reflect that target behavior. Its own user guidance warns that ChatGPT can produce incorrect facts and fabricated citations, and advises people to inspect source links for important information.

Anthropic's January 2026 Constitution treats truthfulness, calibrated uncertainty, transparency, non-deception, and responsiveness to source reliability as normative goals. That document describes intended behavior, not proof that every Claude answer meets it.

Google DeepMind's FACTS Benchmark Suite treats factuality across grounding, search, multimodal inputs, and internal knowledge as measurable properties. Its earlier FACTS Grounding benchmark asks whether a response is fully attributable to supplied source material.

None of this guarantees a trustworthy answer. It shows what the labs think they still have to solve. Truthfulness, grounding, transparency, and calibrated uncertainty are being specified and tested because fluency alone does not provide them. As evaluation improves, claims that can be traced and checked should become more valuable.

The infrastructure is learning to identify actors and actions

The wider internet is also adding ways to identify automated agents, verify where digital material came from, record what users authorized, and trace what agents did.

HTTP Message Signatures, RFC 9421, is IETF standards-track work for signing selected components of HTTP messages. It can support message integrity and verification of key possession. It does not prove that a statement is true or that an actor is honest.

The current Web Bot Auth Internet-Draft applies those signatures to automated web clients. It is an individual draft, not an IETF-endorsed or finished standard. It does not prove who runs a bot, whether the bot is safe, or whether the request was authorized. But it could help websites recognize a returning automated client.

C2PA specification 2.4, published through the Joint Development Foundation project rather than W3C, defines signed, tamper-evident provenance records for digital assets. C2PA does not certify that an image, statement, or claim is true. It makes parts of the asset's declared history more inspectable.

In commerce, Google's Agent Payments Protocol announcement describes cryptographically signed intent and cart mandates designed to record a user's instructions, constraints, approval, items, and price. AP2 is a Google-led open protocol, not a universally adopted standard. Its design still reflects the same pressure: when an agent can act, systems need a record of whose authority it carried and what was approved.

Authentication is not truth. Provenance is not truth. Delegation is not truth. But signed requests, provenance records, and authorization records can be harder to forge or alter, and easier to audit.

Platforms and regulators already act against manufactured trust

The technical direction is not arriving in an empty market. Platforms and regulators already have rules against forms of inauthentic influence.

Google's current guidance for AI features in Search says seeking inauthentic mentions is not a useful long-term strategy. It also says scaled content created primarily to manipulate rankings or generative responses violates its spam policies. That is Google Search guidance, not proof of universal detection across every answer engine.

Reddit's rules require authentic participation and prohibit spam, content manipulation, and deceptive impersonation while allowing pseudonyms. In July 2026, Reddit reported using automated systems against coordinated fake behavior and artificial hype. Its published figures included nearly two million inauthentic votes revoked each day over the preceding three months. Those are Reddit's own enforcement figures, not an independent measure of everything it misses. They are still evidence that manufactured attention creates a platform-detection problem, not a durable brand asset.

That risk is not theoretical, and it is not limited to gaming. In June 2012, Forbes reported sitewide domain bans affecting publishers including The Atlantic and Businessweek. Reddit's general manager described the conduct behind the bans as sophisticated, coordinated vote manipulation. He also said management was sometimes unaware of actions taken on its behalf.

Three months later, Forbes reported that Reddit had banned IGN after suspected vote manipulation. IGN denied using bots and said employee upvotes may have been misread as an attempt to game the system.

In April 2014, Reddit hard-blocked links from onGamers across the site, automatically sending posts or comments with its URLs to spam. Ars Technica later reported that the CBS-owned esports publisher received a minimum one-year ban after an editor privately asked Reddit users to submit its articles.

Reddit had already warned esports publishers that employee-led vote manipulation could cause an entire site to be blocked from submission. These cases do not establish universal or automatic enforcement, and IGN disputed the cause of its ban. They do establish a concrete precedent: Reddit can block every link from a company's website, not only one account.

In the United States, the FTC announced its final Consumer Reviews and Testimonials Rule in August 2024, with an effective date of 21 October 2024. The rule reaches defined practices including reviews attributed to nonexistent people, undisclosed insider reviews, company-controlled review sites presented as independent, certain review suppression, and knowingly fake commercial followers or views. The FTC's current Q&A explains that AI-generated fake reviewers can fall within the rule when the covered representation is false.

The rule is not a blanket ban on AI-generated material, avatars, pseudonyms, or every form of synthetic participation. The rule covers specific conduct. The business lesson is broader: manufactured trust already creates exposure to removal, enforcement, and disclosure duties in important channels.

Update: Listicles and comparison pages after ChatGPT 5.6

On 20 August 2026, Lily Ray published a page-type breakdown attributed to Tomek Rudzki of Peec AI. Comparing citations before and after the ChatGPT 5.6 launch, the share held by pages classified as listicles fell from 15.77% to 7.80%, a relative decline of 50.5%. The share held by pages classified as comparisons fell from 9.08% to 6.17%, a relative decline of 32.1%. The same post reported movement in ChatGPT's fan-out queries, the searches the engine issues for itself before it answers: the "site:" operator rose, along with "pricing", "august" and "official", while "reviews", "best", "top", "comparison" and "vs" all fell as a share of total fan-out queries. Rudzki's stated hypothesis is that ChatGPT is adjusting retrieval to reduce the reach of formats produced at scale to win AI citations.

Read those figures at the strength their method supports, and note who is speaking. Peec AI sells AI visibility tracking and CiteSurge competes in that same market, so neither company reads this data disinterestedly. The numbers reached the public as a screenshot in a social post, with no sample size, no window dates, no prompt population and no market or engine controls. Share of total citations is the only figure published, and share and count move independently: if ChatGPT issued more fan-out queries per answer after the launch, a page type could hold its citation count and still lose share. The post names citations per chat as the diagnostic that separates a change in behavior from that arithmetic, and does not publish it. The fan-out terms carry their own limit, because they are queries the engine issues for itself rather than questions buyers type, so "best" falling as a fan-out term says nothing about whether buyers still ask which option is best.

One panel across one model transition is not a measurement of the web, and CiteSurge will not sell it as one. The direction is worth naming anyway, because it is the argument this article has been making throughout: a format is not an asset. What survives a retrieval change is a page that answers a real question with evidence a reader can check, and that is as true of a comparison a buyer needs as it is of anything else. CiteSurge is built for the long term. Build your brand on truth, honest representation and claims you can stand behind. Do not gamble your reputation on manipulative shortcuts that win attention today and disappear tomorrow.

Accurate claims are stronger long-term assets

There is a tempting short-term argument for synthetic consensus. Put enough favorable material in enough places and some of it may be retrieved. The Morrowen experiment shows that even a small fabricated web presence can reach a narrow set of ChatGPT recommendations. That makes the tactic look clever.

It also shows the weakness. The moment someone checks whether the product exists, who made the claim, whether the source is independent, and what supports it, the tactic falls apart.

An accurate, attributable brand claim behaves differently. It can be stated on an owned page, supported by an original record, qualified for the correct market and date, corroborated independently where appropriate, corrected when facts change, and rechecked over time. The answer may still fail to retrieve it. No source earns automatic selection. But when the claim does appear, it has a chance to survive inspection.

That is the commercial point. Brands should want AI answers that survive inspection, not mentions that disappear when someone checks the source.

This changes what useful GEO work looks like. Brands need clear claims on pages they control, legitimate independent coverage, reliable crawler access, consistent names and product details, and records with dates. Fake personas and paid posts that manufacture agreement may create short-term attention, but they also create more opportunities for detection, removal, enforcement, and reputational damage.

Limitations

No study cited here proves that every model will improve at the same rate, that every platform will detect manipulation, or that every deceptive tactic will be penalized. Crawler documentation describes intended access paths, not complete corpus contents. Platform enforcement figures are self-reported. Benchmarks do not equal production performance. The Morrowen case is a creator-run demonstration, not a controlled market study. The ChatGPT 5.6 page-type figures are one vendor's panel, relayed as a screenshot without a sample size or window.

Identity, provenance, and authorization systems also solve narrower problems than truth. A signed request can carry a lie. A provenance record can faithfully describe the history of misleading material. A legitimate agent can act on bad evidence.

The final conclusion is therefore a CiteSurge prediction from converging technical, platform, commercial, and regulatory evidence. It is not a universal law. Detection, removal, enforcement, and reputational costs will not appear at the same speed across every platform or market.

CiteSurge chooses proof

CiteSurge chooses truthful, honest, source-backed brand representation. We will not sell fake consensus, fake people, or undisclosed synthetic participation.

That is a principle. It is also a long-term strategy.

We believe companies gaming answer systems through fake personas and manufactured attention will face increasing detection, discounting, removal, enforcement, or reputational cost as models and platforms improve. Current rules and enforcement already target some forms of manufactured attention. They do not guarantee how quickly or uniformly those consequences will spread.

We would rather build claims whose evidence an answer can show and a reader can inspect than attention that depends on nobody looking closely. The agentic internet does not need another layer of convincing noise. It needs claims with sources, actors with accountability, and actions that leave a record.

Our decision is simple: proof before mentions.

References

  1. 2025 Year in Review: Exploring the Internet's most popular insights · Cloudflare Radar · reviewed
  2. Overview of OpenAI Crawlers · OpenAI · reviewed
  3. Does Anthropic crawl data from the web? · Anthropic · reviewed
  4. AI features and your website · Google Search Central · reviewed
  5. Evaluating Verifiability in Generative Search Engines · Findings of EMNLP · reviewed
  6. Synthetic Sources? · arXiv · reviewed
  7. I launched a fake deodorant brand to see if AI would notice · @medeana on X · reviewed
  8. Model Spec · OpenAI · reviewed
  9. Inside our approach to the Model Spec · OpenAI · reviewed
  10. Does ChatGPT tell the truth? · OpenAI · reviewed
  11. Claude's Constitution · Anthropic · reviewed
  12. FACTS Benchmark Suite · Google DeepMind · reviewed
  13. FACTS Grounding · Google DeepMind · reviewed
  14. RFC 9421: HTTP Message Signatures · IETF · reviewed
  15. HTTP Message Signatures for Automated Traffic · IETF Datatracker · reviewed
  16. C2PA Technical Specification 2.4 · Coalition for Content Provenance and Authenticity · reviewed
  17. Announcing the Agent Payments Protocol · Google Cloud · reviewed
  18. Guide to Optimizing for Generative AI Features · Google Search Central · reviewed
  19. Reddit Rules · Reddit · reviewed
  20. How We're Keeping Reddit Real and Safe in the AI Era · Reddit · reviewed
  21. Reddit General Manager Responds to Domain Ban Controversy · Forbes · reviewed
  22. IGN on Reddit Ban: We "Don't Use Bots" · Forbes · reviewed
  23. An important message regarding submitting and voting on /r/LeagueofLegends · Reddit · reviewed
  24. Announcement: onGamers has been banned sitewide · Reddit · reviewed
  25. Year-long e-sports site ban shows the dangers of gaming Reddit · Ars Technica · reviewed
  26. Federal Trade Commission Announces Final Rule Banning Fake Reviews and Testimonials · Federal Trade Commission · reviewed
  27. Consumer Reviews and Testimonials Rule: Questions and Answers · Federal Trade Commission · reviewed
  28. ChatGPT 5.6 citation page-type and fan-out query shifts, reported by Tomek Rudzki · @lilyraynyc on X · reviewed
  29. AI visibility tracking platform · Peec AI · reviewed

Find out whether your AI visibility survives inspection.

Request CiteSurge's free three-pillar GEO audit. We will check whether known AI crawlers can reach your public pages, whether key pages answer buyer questions clearly, and whether a sample of AI answers mentions or cites your brand. You will receive a report explaining what is wrong, what to watch, and some fixes to start with. We will also set up a free call to walk through the results.