AI visibility is per-language, not global
Crawlmind Engineering··5 min read
AI visibility is per-language: the language of the query selects the pool of sources an answer engine draws from, before anything about your individual page is weighed.
For a company selling in one language, that costs nothing. For everyone else it breaks the most common way GEO gets reported. A single citation share averaged across markets hides the case that actually matters, which is strong presence in English alongside structural absence in German or Swedish, with no page-level defect to explain it.
#The candidate set is filtered by language first
Temso AI analyzed 7,058,891 individual source citations drawn from 350,000 responses across four models, 12 countries, 47 verticals and six non-English languages. Every model cited local-language sources most of the time, but the margins were not close. Google AI Overview kept 85.4% of its citations in the prompt language, Copilot 76.7%, ChatGPT 70.2%, and Grok 51.7%, a spread of about 34 percentage points between the strictest and the loosest model.
Read it from the other side. English sources took 8.7% of Google AI Overview citations but 39.6% of Grok's. Same question, same market, and the set of pages eligible to be cited is close to a different corpus depending on which assistant the buyer happened to open.
The language pairs vary too. Italian prompts returned 89.9% local-language sources on Google AI Overview and 54.0% on Grok, while Dutch prompts returned 81.2% and 38.3%. Dutch is the weakest case in that study, which matches the intuition that engines fall back to English hardest where the local corpus is thinnest.
Vertical matters as much as language. Local-language citation rates ran from 77% in K-12 education down to 36% in hotels and hospitality. Travel is the giveaway: the audience is international, the available content pool is English-heavy, and the engine follows the pool.
#Language also changes which kinds of sources fill the pool
It is not only the language of the cited domain that shifts. The composition shifts with it.
Profound looked at 3.25 billion citations across seven models and 14 countries in March 2026, filtering prompts to native-language text only. Social sources behaved very differently by engine: Google AI Overviews cited social at 15.3%, Perplexity at 11.3%, ChatGPT at 9.1%, and Claude at 3.99%.
Language then moved those rates in opposite directions on different engines. Spanish queries drew about 1.4x the English social rate in Google AI Overviews (22.5% in Mexico, 20.0% in Spain), while ChatGPT went the other way at roughly half its English rate (4.70% in Mexico, 4.24% in Spain).
The mechanism shows up in the platform mix. Reddit accounts for 51% to 76% of ChatGPT's social citations in every country measured, and Reddit is not a language-diverse property, so ChatGPT's social channel thins out as soon as the prompt stops being English. On the other side, TikTok climbed to 16% of social citations for Spanish queries in Google AI Overviews, and YouTube fell from 38% on English prompts to 26% on Arabic ones.
The practical read is that the third-party surfaces worth investing in are market-specific. A Reddit program that earns ChatGPT citations in the United States does very little for the same brand in Spain.
#Machine translation is not the shortcut it looks like
The obvious cheap fix is to translate. Engines are closing that path, and at very different speeds.
Peec AI examined 64.77 million Reddit citations across 20 countries between March 1 and June 10, 2026. Reddit serves machine-translated versions of threads under a ?tl= URL parameter, which makes translated pages countable from the outside. 3.7 million of those citations, or 5.71% of the total, pointed at translated URLs.
Engine behavior split sharply. Google AI Overview served translated Reddit URLs in 8% to 11% of its Reddit citations and Google AI Mode in 10% to 14%, while Gemini stayed just under 1% and Perplexity never cited a single translated URL across the period. In non-English markets the Google figures went much higher: Google AI Overview reached 52% translated Reddit citations in Germany and over 70% in Sweden and Norway, and Google AI Mode hit 72% in Spain in early June.
ChatGPT reversed course inside six weeks. Its translated Reddit share fell from 6.14% in April to 0.66% in May and 0.30% by early June. In Germany the same drop ran 18.06% to 1.54%, and in Sweden 20.51% to 1.62%.
Whatever the intent behind it, the direction is one way. Translated duplicates are treated as lower-grade supply, and an engine that still serves them can stop inside a single release cycle. Visibility built on that surface can disappear with no warning and no change on your side.
Google's own documentation points the same way for your site. Localized pages are treated as duplicates only when the main content is left untranslated, and Google does not use hreflang or the HTML lang attribute to detect a page's language, it uses algorithms. Translation markup declares a relationship between pages. It does not declare quality, and it does not make a thin machine translation read as a native source.
#What this changes in practice
Report per language, never in aggregate. A blended citation share across markets is a composition statistic. It moves when your market mix moves, and it conceals the one market where you are absent.
Pick the tracked engines per market. If Grok pulls English sources into a Dutch answer far more often than Google AI Overview does (Temso), the two are measuring different things and should not be collapsed into one score.
Build the local corpus rather than mirroring it. A localized page that keeps the structure of the English original but was genuinely written for the market gives the engine a native-language source with its own definitions, examples and entity names. A translation-parameter duplicate gives it a copy that may be filtered out next month.
Budget third-party presence per market as well. Whichever community, review site or video platform carries weight in that language is the one worth earning, and the Profound platform split shows the answer changes market by market.
The underlying point is smaller than it sounds. AI answer engines do not maintain one global list of trusted sources. They assemble a candidate set per query, and language is the first filter applied. Optimizing a page that was never in the candidate set changes nothing.
Related field notes
September 1, 2026 · 5 min
Your AI bot policy now needs three answers
Cloudflare split AI traffic into Search, Agent, and Training, with new defaults landing September 15. One block rule no longer covers it.
September 1, 2026 · 5 min
What Google and Bing's AI reports measure
Google Search Console and Bing Webmaster Tools now ship first-party AI visibility data. Two metrics, two denominators, two blind spots.
September 1, 2026 · 5 min
Source diversity fails before accuracy does
A WWW '26 study shows synthetic content can take over 80% of top-10 retrieval while answer accuracy holds steady. Diversity collapses first, quietly.
Share or discuss
New posts, no spam. Roughly monthly. Unsubscribe with one click.