We respect your privacy.

We use strictly necessary cookies to keep you signed in and to protect against CSRF. With your permission we also use a small amount of first-party analytics to improve the product. We do not sell your data and we do not use third-party advertising trackers. See our cookie policy and privacy policy .

← All posts

Model swaps are the new core updates

Crawlmind Engineering··6 min read

A model swap is a change to the language model behind an AI search surface, and it can reshuffle which sites get cited as sharply as a Google core update, without any change to your pages.

Most AI visibility reporting assumes the engine is a fixed instrument and your content is the variable. That assumption broke several times this year. ChatGPT changed its default model at least twice, and Google now ships new Gemini models into AI Mode every few weeks. Each change moved citations, and none of them came with a note to site owners.

#What the swaps did to citations this year

The measured effects are large enough to swamp most content changes.

Fewer sites per answer. After ChatGPT's early-March transition to GPT-5.3 Instant, French consultancy Resoneo found that the average answer cited 19 unique domains before and 15 after, with unique URLs per answer falling from 24 to 19. The data came from Meteoria, which tracked 400 prompts daily over 14 weeks for 27,000 comparable responses. Roughly one in five cited sites dropped out of the average answer, and nobody's content changed to cause it.

A different set of winners. SISTRIX tracks German-language ChatGPT answers daily. When the model behind its sample switched from GPT-5 mini to GPT-5.5 on May 22 to 23, the domain distribution of citations shifted by 47% within 48 hours, against normal day-to-day drift of 1 to 2%. That comparison covered 800,000 responses, four days either side of the change, from a wider set of 3.8 million responses and more than 100 million citations. SISTRIX called it a ChatGPT core update, and the label fits: almost half of all citations went to different domains afterward.

Brand sites lost ground to forums. Writesonic compared GPT-5.3 Instant with its replacement, GPT-5.5 Instant, on May 6. First-party brand websites made up 13.4% of citations on the old model and 6.0% on the new one, and Reddit went from the ninth most cited domain to the first. This is a small study (42 prompts that ran cleanly and about 235 citations), so treat the exact ratio with caution. The direction matches what the larger datasets show: a model change re-ranks sources, and the re-ranking is not neutral between site types.

Citations can disappear outright. Google released Gemini 3.8 Flash into AI Mode on September 2, 2026, less than three weeks after Gemini 3.7 Flash. Within a day, practitioners posted screenshots of the new model answering without links or citations, particularly for top-of-funnel queries. Google's Robby Stein acknowledged on X that it "isn't working as intended", and a fix landed the next morning. For roughly a day, a page could be retrieved, used and still earn zero visible citations on that model.

#Why this breaks the usual diagnosis

A drop in AI citations usually triggers a content review: freshness, structure, schema, competitor pages. When the cause is a model swap, that review finds nothing, or worse, finds something unrelated and gets credit when citations drift back.

Three properties of model swaps make them easy to misread:

  1. They are rarely announced to site owners. OpenAI and Google announce new models to users, framed around reasoning and factuality. Neither publishes how a new model changes retrieval or source selection.
  2. They are not uniform across users. In Google AI Mode, newer Flash models have appeared as options in a model picker for paid Google AI Pro and Ultra subscribers rather than as the free default, per Search Engine Roundtable's coverage of the 3.8 Flash launch. Two people running the same query can get answers from different models.
  3. They are not uniform across prompts. Writesonic found that 8 of 50 prompts auto-escalated to GPT-5.5 Thinking even with the auto-switch toggle off, and described the routing as content-based rather than random. Within a single tracked prompt set, some answers come from one model and some from another.

The practical result is that "ChatGPT" and "AI Mode" are not single engines. They are a changing set of models behind one brand name, and your visibility is a property of each model separately.

#How to tell a model swap from a content problem

The signature of a model swap is breadth and timing. A content problem tends to hit specific pages or topics, and it builds over days or weeks as crawls and re-indexing catch up. A model swap hits many unrelated prompts on the same engine at the same time, and it often hits your competitors too.

Run these checks before touching any page:

  • Did the drop start on a single day across unrelated prompts? A step change within 24 to 48 hours, across topics that share nothing but the engine, points at the engine.
  • Did other engines move? If ChatGPT citations fell and Perplexity and AI Overviews held steady for the same prompts, the cause is probably inside ChatGPT.
  • Did the number of cited sources per answer change? The Resoneo data shows a model can shrink the citation list for everyone. If your share of the remaining slots held steady, you did not lose relevance. The pie got smaller.
  • Did the mix of source types change? A swing from brand sites to forums, or from aggregators to local publishers, is a selection change, not a ranking change for your page.
  • Is there a public model announcement in the window? Check the engine's release notes and the trade press for the dates in question. The SISTRIX team confirmed its date by asking the system which model it was running, and the answer changed on the day the citations moved.

If all of these point at the engine, the right response is usually to wait a week, keep measuring, and avoid rewriting pages to chase one model's preferences. Models are now replaced within weeks. Optimizing for the quirks of one is optimizing for something that may be gone by the next sprint.

#What to log so you can tell next time

Most of the diagnosis above is impossible without the right data captured at run time. For each tracked answer, store:

  • The model, when the surface exposes it. Some API responses and some interfaces report the model version. Record it every time it is available, and record "unknown" when it is not, rather than leaving the field blank.
  • The full list of cited URLs, not only whether you appeared. Citation count per answer is the quickest way to spot a list that shrank for everyone.
  • The source type of each citation. Brand site, forum, publisher, marketplace, reference. Shifts between types are often the first visible symptom of a model change.
  • The run timestamp, to the hour. Weekly snapshots blur a step change into a gradual slide.
  • A short annotation log. Note known model launches against your charts, the way SEO teams annotate core updates.

Keep engines apart in reporting. A blended "AI visibility" number across ChatGPT, AI Mode and Perplexity averages a model swap on one engine into a mild dip, which is exactly the kind of change that sends teams looking for a cause in their own content.

#The working assumption

Treat every AI engine as a rolling series of models, each with its own citation habits. Content quality still decides whether you are eligible to be cited. The model decides how many slots there are and which kinds of sources fill them, and that part changes on the vendor's schedule, not yours.

Related field notes

Share or discuss

Field notes in your inbox

New posts, no spam. Roughly monthly. Unsubscribe with one click.