We respect your privacy.

We use strictly necessary cookies to keep you signed in and to protect against CSRF. With your permission we also use a small amount of first-party analytics to improve the product. We do not sell your data and we do not use third-party advertising trackers. See our cookie policy and privacy policy .

← All posts

Slowing Googlebot also slows your AI answers

Crawlmind Engineering··5 min read

Crawl-rate throttling is how a site tells a crawler to slow down, and for Google the only lever left is the HTTP status code your server returns.

On October 6, 2026, Google restructured its documentation on reducing the Google crawl rate and added examples for the Retry-After header (Search Engine Roundtable). The mechanism is not new. Google already treated 500, 503 and 429 responses as a signal to back off. What the update does is put the whole emergency playbook on one page, including the costs. Those costs matter more now that the same index feeds AI Overviews and AI Mode.

#What the updated page says

The guidance is short and specific:

  • Return 500, 503 or 429 to reduce crawling urgently.
  • With 503 or 429, add a Retry-After header, either as seconds (Retry-After: 120) or as an absolute UTC date, following the HTTP semantics in RFC 9110.
  • Treat this as short-term only. Google frames the safe window as hours up to 1 to 2 days.

The warning sits next to the instructions. If Googlebot sees these status codes "on the same URL for multiple days, the URL may be dropped from Google's index." Even inside the safe window, Google lists the side effects: fewer new pages discovered, existing pages refreshed less often ("prices and product availability may take longer to be reflected in Search"), and removed pages staying in the index longer.

There is no softer control. Google does not support the crawl-delay line in robots.txt, and the Search Console crawl rate limiter tool was deprecated on January 8, 2024. Google's reasoning at the time was that Googlebot already slows down when a server returns errors or responds slowly, and that the old tool could take over a day to apply.

#Why this is now an AI visibility question

Google's documentation on AI features says a page needs to be indexed and eligible for a snippet to be shown as a supporting link in AI Overviews or AI Mode, with no additional requirements. There is no separate AI crawler for those surfaces. Googlebot's view of your site is the raw material.

That turns each of Google's listed side effects into an AI answer problem:

  • Fewer discovered pages means a new comparison page or pricing page is not available to cite until crawling recovers.
  • Slower refresh means the AI answer keeps quoting the old price, the old feature list, or the old policy for longer.
  • Removed pages staying indexed means a retired page can keep being used as grounding after you have taken it down.
  • Errors lasting multiple days can drop the URL from the index, and an unindexed page is not eligible for AI features at all.

A throttle you add during a traffic spike and forget to remove is a slow, quiet AI visibility loss. Nothing in an AI answer tells you that the source is stale because your edge returned 503 to Googlebot for a week.

#Other AI crawlers listen to different signals

The pressure to throttle is real. Wikimedia reported in April 2025 that bandwidth for multimedia downloads had grown 50% since January 2024, and that bots made up about 35% of pageviews but at least 65% of its most expensive traffic. Many smaller sites see the same pattern in their own logs and respond with blanket rate limits.

Blanket limits ignore that each crawler accepts a different control:

  • Googlebot reacts to status codes and response time. crawl-delay is ignored.
  • Bingbot honors crawl-delay in robots.txt, according to Bing's webmaster blog, which recommends the lowest value possible to keep the index fresh.
  • ClaudeBot supports the non-standard crawl-delay extension, per Anthropic's crawler help article.
  • OpenAI documents its crawlers and user agents but says nothing about crawl-delay, 429 or Retry-After. It notes that robots.txt changes can take about 24 hours to affect search results, and that ChatGPT-User visits pages when a user asks a question and is not used for automatic crawling.

The last point deserves attention. A user-triggered fetch happens while someone is waiting for an answer. If your rate limiter answers that request with 429, the engine has nothing to read from your page at that moment. A crawl-style throttle applied to a user-style fetcher does not delay anything. It just removes you from that answer.

#Throttle by bot, briefly, and with an end date

None of this argues against protecting your servers. It argues for being precise about it. A practical setup:

  1. Identify crawlers before you limit them. Verify Googlebot and other major crawlers by their published IP ranges or reverse DNS instead of trusting the user agent string. Impersonators are the traffic you can block outright.
  2. Use the control each crawler reads. 503 or 429 with Retry-After for Googlebot. crawl-delay for crawlers that document support for it. Do not use 401 or 403 to slow crawling: Google's robots.txt documentation says 4xx codes other than 429 have no effect on crawl rate.
  3. Never throttle robots.txt itself. Google's spec says that if robots.txt returns a server error, Google stops crawling the site for the first 12 hours and then uses the last good copy for up to 30 days. A rate limiter that catches /robots.txt can halt all Google crawling.
  4. Exempt user-triggered fetchers or limit them separately. Their volume tracks real user questions, and the cost of refusing one is a missing citation in a live answer.
  5. Put an expiry on every emergency rule. Google's own window is 1 to 2 days. A rule without an end date becomes the multi-day error pattern that gets URLs dropped.
  6. Log status codes per crawler. Count 429 and 5xx responses by bot and by URL each day. If a key page has returned errors to Googlebot for more than a day, treat it as an AI visibility incident, not just an infrastructure note. We covered this logging approach in reading your AI crawler log by status code.
  7. Annotate throttle windows. When citations or AI answer accuracy shift, the first question should be whether you told the crawler to stay away during that period.

#What to take from the update

Google did not change how it crawls on October 6. It made the emergency brake easier to find and was candid about what pulling it costs. For teams that track AI answers, read that list of costs as a list of ways AI Overviews and AI Mode fall behind your site. Throttle the crawler that is actually causing the load, use the signal that crawler understands, and take the throttle off as soon as the load drops.

Related field notes

Share or discuss

Field notes in your inbox

New posts, no spam. Roughly monthly. Unsubscribe with one click.