We respect your privacy.

We use strictly necessary cookies to keep you signed in and to protect against CSRF. With your permission we also use a small amount of first-party analytics to improve the product. We do not sell your data and we do not use third-party advertising trackers. See our cookie policy and privacy policy .

← All posts

An AI SEO diagnosis needs an evidence trail

Crawlmind Engineering··5 min read

An AI SEO diagnosis is a model's claim about why a page or site is underperforming, and a new benchmark shows that many models will make that claim with confidence even when the data in front of them cannot support it.

That is a practical problem for anyone who pastes a Search Console export, a crawl report or a log sample into a chatbot and asks what went wrong. The answer reads the same whether the evidence was decisive or missing. The difference only shows up later, when the fix does not work.

#What the benchmark tested

On October 8, 2026, iPullRank published WARRANT-SEO, a benchmark built to check whether a model notices when an SEO case cannot be solved from the information given. The study ran 27 models against 27 SEO scenarios, each written in three versions:

  1. The full case, including one decisive piece of evidence that points to a single cause.
  2. The same case with an added option to say the evidence is insufficient. Taking that exit here counts as a failure, because the evidence is there.
  3. The case with the decisive line removed, so the only correct answer is that there is not enough information.

The design rule matters. According to the study, every scenario must have a decisive fact that uniquely supports the right diagnosis, and removing that fact must leave at least two plausible explanations. A model that still picks one cause in version three is guessing.

Each of the 81 questions was run three times per model, which produced 6,561 responses, of which 6,427 were scored. The response set was collected in August 2026, and the full scenario pool is kept private to limit memorization.

#The gap between knowing and judging

With the decisive evidence present, models were correct 99% of the time. With it removed, they correctly recognized the case as unsupported only 51.5% of the time. In the other half of those cases they committed to a specific cause anyway.

Results varied sharply by model. Seven models, including GPT-5.5, Gemini 3.7 Flash and Claude Opus 5, posted a perfect WARRANT score of 1.000. Mistral Large scored 0.500: it answered every evidence-present case correctly and never once said the evidence was missing. Claude Sonnet 5 scored 0.728 and Claude Haiku 4.5 scored 0.556, so behavior was not uniform even within one lab.

The authors also ran a separate technical SEO knowledge test on the same models. Knowledge and judgment were strongly correlated overall, with a Spearman correlation of 0.803, yet some models split. Mistral Large scored 0.808 on knowledge and 0.500 on judgment. A model can know what crawling, indexing and rendering mean and still fail to see that a given export says nothing about which one broke.

Price did not track judgment either. The study reports that a cheaper model reached the same perfect score as models costing several times more per thousand calls.

#One sentence in the prompt changes the answer

The most useful finding for practitioners is about the prompt, not the model. In a side experiment on 22 dated factual questions, the authors added one sentence telling models to say so when their training data did not cover the period asked about. Abstentions rose from 20.7% to 85%. Wrong answers fell from 14.5% to 1.2%. Correct answers fell too, from 64.8% to 13.8%, and GPT-5.5 went from 95.5% correct to 0%.

So the instruction did not make models smarter. It moved the threshold at which they committed. That is a reminder that the thing you are evaluating is the whole setup (model, prompt, context and tools), not a model name on a leaderboard.

The study also found that some models gave the same unsupported answer on every run. Repetition looks like confirmation. It is not. A model that is consistently wrong for the same reason will agree with itself every time you ask.

#Google is saying the same thing about AI output

This lines up with a change Google made to its own guidance the same month. On October 1, 2026, Google updated its documentation on generative AI content to say it is critical to manually fact-check and review all AI-generated content before publishing. The added text explains that generative models do not retrieve facts but predict a likely sequence of words, as reported by Search Engine Roundtable. The guidance also extends review to titles, meta descriptions, structured data and image alt text.

That guidance is about published content, but the logic carries over to diagnosis. A model's explanation of a traffic drop is also generated text. It deserves the same check before anyone acts on it.

#What a defensible diagnosis looks like

Google's own guide to debugging Search traffic drops lists several distinct causes: algorithmic updates, technical issues, security issues, spam issues, seasonality and site moves. Each one leaves a different fingerprint. A site-wide drop points to the Page indexing report. A drop in a group of pages points to URL inspection. Falling clicks with stable impressions points to titles and snippets. The guide also suggests Google Trends to check whether the drop belongs to the site or to the whole query space.

That is the shape an AI-assisted diagnosis should take. Before accepting one, ask the model for four things, which mirror the evidence trail the WARRANT authors recommend:

  • The claim. One named cause, stated plainly.
  • The supporting observation. The specific row, date, status code or report that points to that cause and not another.
  • The competing causes. What else could produce the same symptom, and why the model ruled it out.
  • What is missing. The data that would confirm or kill the diagnosis, such as a year-over-year comparison, server logs for the affected URLs, or the date of a template change.

If the model cannot fill in the second line with something you can open and verify, treat the answer as a list of hypotheses. That is still useful. Brainstorming possible causes is a good use of a model. Deciding which one is true is a different task, and it needs evidence.

#Apply the same rule to tools

The same standard works for any audit tool, AI or not. A finding should name the URL, the check that failed and what was observed, so the person reading it can confirm the issue without trusting the tool. When a cause cannot be established from crawl data alone, such as a traffic drop that may be seasonal, the honest output is a statement of what the crawl can and cannot show.

The WARRANT results suggest the best models are already capable of that restraint. Whether you get it depends on which model you use and how you ask. Ask for the trail every time, and treat any diagnosis without one as a guess.

Related field notes

Share or discuss

Field notes in your inbox

New posts, no spam. Roughly monthly. Unsubscribe with one click.