We respect your privacy.

We use strictly necessary cookies to keep you signed in and to protect against CSRF. With your permission we also use a small amount of first-party analytics to improve the product. We do not sell your data and we do not use third-party advertising trackers. See our cookie policy and privacy policy .

← All posts

A robots.txt block does not unpublish you

Crawlmind Engineering··5 min read

A robots.txt block is a request that crawlers stop fetching your pages from now on; it does not remove what an AI engine has already fetched, indexed or cached.

That distinction sounds obvious when written down. In practice, many teams treat a Disallow line as an off switch for AI answers: add the rule, wait a day, and expect ChatGPT or Perplexity to stop quoting the page. A study from researchers at Duke and Carnegie Mellon tested that expectation directly, and the result is worth knowing before you change a crawler policy for legal, commercial or reputational reasons.

#How the study worked

The paper, Identifying AI Web Scrapers Using Canary Tokens (Seiden, Ren, Zhang, Kim, Liu and Wenger; revised September 3, 2026), built dynamic websites that served a unique token to every visiting scraper. The researchers then asked production chatbots about those sites. If a chatbot repeated a token, the token identified exactly which scraper had fetched the content that reached the answer.

The setup covered 22 AI chatbots and 20 websites, with 4,042 unique visitors recorded across the sites, and queries run in March and April 2026. Four systems (DeepSeek, Hunyuan, GLM and Liquid) never returned a token, which left 18 systems with usable data. The study measured real-time retrieval at answer time, not pretraining.

#The bot that answers is often not the AI bot

The first finding is a map of which crawler feeds which chatbot. According to the paper, ChatGPT content came through OAI-SearchBot, Copilot through Bingbot, Gemini through Googlebot, and Claude through Brave's crawler. Perplexity returned content fetched by both PerplexityBot and Googlebot, and the authors note that relationships such as Qwen and Perplexity returning Googlebot-fetched content were not publicly documented.

Six systems fetched pages with ordinary browser user agents rather than a named bot, including ERNIE, Grok, Qwen and Kimi. Kimi rotated through a large list of user agent strings while scraping.

The practical reading: a robots.txt file that names GPTBot, ClaudeBot and PerplexityBot covers only part of the path from your server to an AI answer. Several engines answer from a search index built by a crawler you probably allow on purpose, and some fetch without announcing themselves at all. We covered the identity side of this in your crawler allowlist trusts a string.

#What happened after the block

In the second stage, the researchers split the sites. Ten were taken offline. The other ten stayed up with the most restrictive robots.txt possible, User-agent: * followed by Disallow: /. One week later they queried the chatbots again.

The headline result, in the authors' words: "Of the 18 AI chatbots for which we obtained User-Agent information, 12 continued to return content in both blocking conditions." Duck.ai was the only chatbot that stopped returning the content in both conditions. Among chatbots that had recited content fetched by Googlebot, Bingbot or Brave's crawler, seven of eight kept doing so after the sites had been offline for a week.

The authors conclude that taking a site offline or disallowing crawling does not appear effective at stopping chatbots from returning content "if that content was already indexed prior to blocking."

Two caveats keep this in proportion. The sample is 20 purpose-built sites, not a population of real publishers. The follow-up window was one week, so the study does not show how long cached content survives. It does show that the survival period is longer than most teams assume when they flip a rule.

#Why a block cannot do what people expect

Robots.txt governs fetching. It says nothing about copies that already exist. An answer engine that retrieves from its own index, or from a partner's search index, can keep serving the stored version until that index drops or refreshes the document. Your new rule only takes effect the next time a compliant crawler checks it, and even the vendor side builds in delay: OpenAI says it can take about 24 hours from a robots.txt update for its systems to adjust.

There is a second, less intuitive problem. Blocking a page can prevent removal. Google's own documentation states that for a noindex rule to be effective, the page must not be blocked by robots.txt, because a crawler that cannot fetch the page never sees the instruction. A site that disallows a page it wants gone from answers may freeze the old indexed copy in place.

Token scope adds more confusion. Google states that Google-Extended "does not impact a site's inclusion in Google Search", and AI Overviews and AI Mode are part of Search. For those surfaces, Google points site owners to nosnippet, data-nosnippet, max-snippet or noindex. OpenAI's documentation separates its bots the same way, noting that for ChatGPT-User, "robots.txt rules may not apply" because a user starts the fetch.

#What to do instead

Decide what you are trying to achieve before you edit the file, because each goal needs a different tool.

To stop future training collection, robots.txt rules for training tokens such as GPTBot and Google-Extended are the correct and sufficient lever. They were never meant to affect answers.

To keep a page out of AI answers going forward, let the relevant search crawler fetch it and serve a removal signal it can read: noindex, or nosnippet and data-nosnippet for Google surfaces. Keep the page crawlable until the engines have processed the change, then block if you still want to.

To remove content that must disappear, such as legal, privacy or pricing errors, change or remove the page itself and return a 404 or 410 so recrawls find nothing to store. Use the engines' own removal channels where they exist, and expect a lag measured in more than days.

To confirm it worked, re-run the specific prompts that surfaced the content and log the answers over several weeks. Checking your server logs tells you whether bots stopped fetching. It does not tell you whether answers stopped quoting you, and the canary study shows those two things can diverge for at least a week.

The longer-term fix sits outside any single site. The IETF AI Preferences working group is standardizing a vocabulary for how content may be used by AI systems, separate from whether it may be crawled. Until something like that is adopted and honored, treat robots.txt as a rule about tomorrow's fetches, not about what engines already hold.

Related field notes

Share or discuss

Field notes in your inbox

New posts, no spam. Roughly monthly. Unsubscribe with one click.