Where Does AI Get Brand Information? 85.7% Is Third-Party

Where Does AI Get Brand Information? 85.7% Is Third-Party

A study that classified 167,551 citations AI attached to brand answers across 12 European markets and 13 languages

Where does AI get brand information? Of 167,551 citations for 128 European brands, 85.7% were third-party sites. The top source shifts by language.

By · TRAIL Labs Research
GEOAEOAI CitationsBrand ReputationSource AnalysisMeasurement

When AI answered questions about a brand, 85.7% of the URL citations it attached pointed to sites that brand does not own. That is the conclusion of a study that classified, one by one, 167,551 citations that search-grounded AI models attached to answers about 128 brands across 12 European markets and 13 languages [1]. Instead of looking at what the answers said, the study looked one step earlier, at what the model read before writing. Owned sites made up just 14.3%, and Wikipedia was the top source in nearly every language, but at a 4–6% share.

We wrote this post from the full paper on arXiv (v1), checking the numbers in the text against the paper's Tables 1–6. Every number here comes from the paper's data, and we did not reproduce it. The author is affiliated with a company that builds AI brand intelligence tools, and the paper discloses that the three datasets analyzed are that company's research data [1]. In return, the data and analysis code are published on Zenodo under CC BY 4.0, so anyone can regenerate the numbers. One Polish figure appears in two different versions inside the paper, so for that part we carried over only the ratio. We also run an AI search diagnosis tool, so we have marked separately where the last section connects to our product.

Why look at citation sources instead of AI answers

A search-grounded model retrieves web pages first and writes its answer from what it retrieved, so the sources it reads set the limits of what it can say [1]. Most AI brand visibility monitoring looks at the answer text: was the brand recommended, what was the tone, did a competitor appear alongside it. This study asks the question one step earlier. Where does AI get its information about our brand, and can we influence that source?

The question is practical because the answer leads to opposite strategies. If AI mostly reads the brand's own site, you can improve the answer by writing a better site. If it mostly reads third parties, you have to earn mentions on sites you don't control. To know which, you have to count sources.

It also matters that citations do not accurately mirror their sources. An audit of verifiability in generative search engines found that only 51.5% of generated sentences were fully supported by their citations, and only 74.5% of citations actually supported the sentence they were attached to [3]. That is why this study draws the line up front: it measures the citation mix and does not claim it equals real reputation or how people see a brand [1]. What it measures stops at "what AI cited."

Data and method: three datasets, 167,551 URL citations

The study merged three citation datasets collected by a measurement company [1]. Each one was built by asking search-grounded LLMs about brands and recording the sources the models cited.

DatasetRecord unitSize
Nordic-Baltic (NB)brand × language × model × prompt × citation150,093 rows
Poland (PL)brand × prompt × model × iteration4,151 records, 35,880 citations
Central and Eastern Europe (CEE)attributed source mention4,001 rows (157 with a URL)

Table 1. Makeup of the three merged datasets. URL citations that can be analyzed by domain total 167,551: 131,514 from NB, 35,880 from PL, and 157 from CEE [1].

The data cover 128 brands, 13 languages, and 12 home markets: the Czech Republic, Denmark, Estonia, Finland, Germany, Latvia, Lithuania, Norway, Poland, Slovakia, Sweden, and Hungary [1]. The models in the Nordic-Baltic data were Perplexity Sonar Pro, Gemini 3.1 Pro, and GPT-5.4. The author states up front that the merged data is not one uniform table. Only NB and PL carry URLs that resolve to real domains; CEE attributed 96.1% of its sources only by keyword, such as "tier1 news." So every domain-level figure was computed on the NB and PL URL data.

The classification is simple [1]. Each cited URL was reduced to its registrable domain, and Wikipedia's language subdomains were merged into wikipedia.org. A citation counted as owned if the brand's name token appeared in the cited domain, for example wise.com for Wise. The data also carry source-type labels such as company website, Wikipedia, industry report, tier-1 news, social, review platform, government, and other web, so the author reports both the domain-based and the label-based figures and flags where the labels mislead.

Brand sites are only 14.3% of the sources AI cites in brand answers

Of 131,514 URL citations in the Nordic-Baltic data, 85.7% pointed to sites the brand does not own and 14.3% to owned domains [1]. Counting all 150,093 rows, including answers given from the model's internal knowledge with no URL, the split is 14.4% owned, 76.1% third-party, and 9.5% implicit. However you count it, AI read third parties about a brand four to six times as often as it read the brand's own site.

Card showing, side by side, the share of sources AI cited about a brand that were sites the brand does not own and the share that were owned

Figure 1. A bar splitting third-party sites (85.7%) from owned sites (14.3%), plus the study's scale: 167,551 URL citations, 128 brands, 12 European markets, 13 languages. A card we drew from the values in the paper's Table 2 [1].

SplitOwnedThird-party
All URL citations14.3%85.7%
By company website label16.4%83.6%
B2B brands13.1%86.9%
B2C brands15.7%84.3%

Table 2. Owned versus third-party citation share, based on 131,514 Nordic-Baltic URL citations [1].

The gap between B2B and B2C was small. B2C brands drew slightly more owned citations, but the direction was the same. Owned citations don't disappear; they concentrate in a few brands [1]. Among brands with at least 1,500 citations, the most self-cited were Tatra Banka at 34.4%, Statkraft at 33.9%, ESET at 33.2%, and Wise at 32.4%. Even the best case is a third. At the other end, several brands such as Slovak Telekom, PKN Orlen, and Kiwi.com drew 0.0% owned citations. In describing those brands, AI never once used the brand's own site as a source.

This direction matches earlier work. A study of generative engine optimization reported that AI search systematically favors authoritative third-party media over brand-owned content [2]. This paper confirms the same tendency at the level of individual brand citations, with multi-country data.

Citation sources are concentrated in a small set of domains

The web AI reads about brands is wide, but most citations come from a small head [1]. Of 20,815 registrable domains, 3,778, about 18.2%, supplied 80% of all citations. By host, 547 (2.3%) supplied half of all citations, and the remaining 20% of citations spread across more than 17,000 tail domains.

The author fit a Zipf law to the top 1,000 domain ranks with a log-log regression [1].

is the number of citations received by the domain ranked . An exponent of 0.86 means that each time the rank doubles, a little more than half of the citations remain, and an of 0.983 means this simple law explains the top-1,000 distribution almost exactly. The tail is real but thin.

By source-type label, "other web" dominates at 77.0%, followed by owned 16.4%, Wikipedia 3.9%, industry report 1.0%, tier-1 news 0.7%, social 0.4%, review platform 0.3%, and government 0.1% [1]. The author warns against reading this table at face value, because the label classifier dropped most real news and review domains into other web. The honest summary is that the open third-party web is the dominant source, ahead of any curated category. Across the whole dataset, the most-cited domains were Wikipedia (5,800 citations), YouTube (2,826), and Statista (1,310), followed by brands' own sites and social platforms such as Reddit.

The top cited source changes with the language

Wikipedia was the most-cited domain in 11 of 12 languages, but each market also had local sources that could beat it [1]. Wikipedia's share ranged from 3.71% in Polish to 5.70% in Finnish. Saying "Wikipedia dominates" means it ranks first, not that it supplies most citations.

Card showing the most-cited domain and its share for each of 12 languages as bars, with only Lithuanian highlighted where vz.lt leads

Figure 2. Top cited domain by language. wikipedia.org leads 11 languages at 3.71–5.70%, and only in Lithuanian does vz.lt lead (4.38%). The box below notes the Polish recruitment-portal case and that Korean is not in the study. A card we drew from the values in the paper's Table 4 [1].

In the one exception, Lithuanian, the national business daily Verslo žinios (vz.lt) edged out Wikipedia at 4.38% [1]. The author finds the exception instructive. An English-only or Wikipedia-centric view of AI sourcing would miss exactly this kind of local business outlet, and outlets like it appear in smaller-language markets where local press fills space the encyclopedia does not.

Poland is even sharper [1]. Across 35,880 citations for 46 Polish national brands, the most-cited single domain was YouTube (6.4%), followed by the price-comparison and marketplace sites ceneo.pl and allegro.pl. Just below them, four HR and careers portals (pl.indeed.com, livecareer.pl, interviewme.pl, randstad.pl) together were cited about twice as often as Polish Wikipedia. In a market with strong employer-review sites, AI reads those sites to describe a brand. How a brand looks on a careers portal can make it into AI answers more often than its encyclopedia entry.

Interestingly, the mix of source types barely changed across languages. The chi-square test was significant (χ² = 1906.4, p < 0.001), but the effect size, Cramér's V, was a tiny 0.036 [1]. At this sample size significance comes cheap, and the small V is the honest summary. What changes by language is not the source-type mix but which domain leads in that market. By language family, Uralic (Estonian, Finnish) and Baltic (Lithuanian, Latvian) answers leaned a bit more on owned sites and Wikipedia (15.5% and 15.0% owned), and Slavic (Polish, Czech, Slovak) leaned on them least (12.5% owned). The author reads this as AI falling back on brand sites and the encyclopedia in languages with a thinner native web.

23,027 Gemini citations arrived as Google redirect addresses

The most useful part of this paper for measurement practice is not a result but one paragraph in the method section [1]. In the Nordic-Baltic data, 23,027 rows, 17.5% of URL rows, held a redirect host, vertexaisearch.cloud.google.com, instead of the real source domain. All of them were Gemini citations, virtually all of Gemini's 23,032. The citation links returned through Grounding with Google Search in the Gemini API were Google's relay addresses rather than the original URLs. The real domain sat in each citation's title field, and the author resolved all 23,027 from the title.

The author calls this step load-bearing [1]. Without it, the redirect host becomes the top domain in every language, and Wikipedia loses first place in every language. In the Polish data, Gemini citations also arrived as google.com redirect addresses and were resolved the same way.

Card comparing Gemini's owned-citation rate when the redirect is read as the source with the rate after resolving the real domain from the title field, plus a bar chart of owned-citation share by model

Figure 3. The Gemini redirect trap. The owned-citation rate of 0.0% when the redirect is read as is versus 5.8% after resolution, plus owned-citation share by model (Perplexity Sonar Pro 16.8%, GPT-5.4 12.9%, Gemini 3.1 Pro after resolution 5.8%). A card we drew from the values in the paper's Section 2.4 and Table 6 [1].

ModelCitationsDomainsOwned (by domain)Owned (by label)Wikipedia
Perplexity Sonar Pro90,27615,99516.8%20.7%4.9%
Gemini 3.1 Pro23,0326,5685.8%0.0%2.6%
GPT-5.418,2063,28412.9%16.2%4.3%

Table 3. Citation behavior by model, Nordic-Baltic data after redirect resolution [1].

Gemini's owned rate reads 0.0% by label because the source-type classifier ran on the redirect address instead of the real domain [1]. Counted again on the resolved domain, it is 5.8%. In the Polish data, Gemini also read 0.0% owned because of the redirect (against Perplexity 18.7% and OpenAI 11.7%). The author's conclusion is clear: any comparison of owned-citation rates across models must resolve the redirect first, or it will report a model's collection artifact as a finding about brands.

This trap is not unique to this paper. An AI recommendation audit that enumerated 4,776 restaurants in Bali also reports that Gemini's API returned grounding links only as relay addresses, so its author recovered the real domains from citation titles [4]. Two different studies independently doing the same step is a signal that any measurement handling Gemini citations needs it. The differences between models were large in their own right. Perplexity cited the most and most widely, accounting for 68.6% of Nordic-Baltic citations across 15,995 domains, while GPT-5.4 leaned on a narrow set of 3,284 domains.

Limitations of this study

The biggest limitation is geography. All 12 markets are in Europe, and Korea and the rest of Asia are not in the data [1]. There is no basis for carrying these shares over to Korean-language answers.

The author also states the limits of the data structure [1]. The three datasets were collected under related but not identical protocols, so the merged result is an NB and PL URL backbone with a CEE cross-link. Owned detection relies on a heuristic, whether the brand name appears in the domain, so it misses owned domains without the brand name and can count third-party domains that happen to contain it. Citations attach at the answer level, so when a model lists several brands and cites one source, that source attaches to every brand listed.

There are limits of time and version, too. The data reflect specific model versions and collection windows, and search-grounded models change how they retrieve frequently [1]. Domain shares and per-model figures vary with time and version.

One figure is inconsistent within the paper. The Polish recruitment-portal and Wikipedia citation counts appear as one pair of values in the abstract, results, and conclusion, and as a different pair in the limitations section [1]. Both pairs give a ratio of about two, so in this post we carried over only the "about twice" ratio. Finally, there is a conflict of interest: the data were collected by a measurement company that sells tools adjacent to this question. The author states the company had no separate influence on design or interpretation and published the data and code, but we have not reproduced the results.

How to use this in AEO and GEO practice

The author draws four practical conclusions directly from the data [1]. First, earn third-party mentions instead of relying on your own site. A clear company page is necessary, but it is not where AI mostly reads a brand. Second, win the head of the distribution. Since 80% of citations come from about 18% of domains, mentions on head domains are worth more than scattered mentions in the tail. Third, Wikipedia is a common source across almost every language, while local outlets differ by market. Fourth, audit by market and by model. An audit in one language with one model shows only a slice of where AI reads a brand.

These conclusions point the same way as posts we have written before. The off-site source map in Off-Site AEO: Why AI Won't Cite Your Flawless Site, the third-party-heavy sourcing across five AI engines in Which Brands Do AI Search Engines Mention?, and our argument in Who Owns AEO? that brand consensus has to reach beyond the SEO team all get one more layer of support from this study's multi-country data.

From TRAIL's side, this paper made us check our own code. TRAIL Search's diagnosis sorts the citation URLs an engine attaches into owned, wiki (including Namuwiki), YouTube, reviews and community (including Naver blogs and cafes), media, and listicles. After reading this paper, we checked whether the Gemini redirect paragraph applied to our pipeline, and it did. Our pipeline was also reading the vertexaisearch.cloud.google.com address in Gemini citations as the domain, which meant owned sources could never register in Gemini citations. We have now fixed it to resolve the real domain from the citation title field, and deployed the fix. The most practical lesson from this paper is that a single choice in how the collection path is handled can produce a number that looks like a finding about brands.

And Korea is not on this map. We have not yet seen public numbers on whether Korean-language answers lean more on Wikipedia or Namuwiki, or how often Naver blogs and cafes serve as sources. Whether Korea has a source that leads only there, like Lithuania's business daily or Poland's recruitment portals, is something only direct measurement can answer. It is the question we want to measure next.

Frequently asked questions

How often does AI cite a brand's own website when answering about that brand?

In this study of 128 European brands, only 14.3% of URL citations pointed to the brand's own domain. The other 85.7% pointed to sites the brand does not own. Even the most self-cited brands reached about 34%, and several brands reached 0%.

Which sites are cited most in AI answers?

In this study's Nordic-Baltic data, Wikipedia was the most-cited domain in 11 of 12 languages. Its share was only 3.71–5.70%, though, so leading does not mean dominating. In Lithuanian, the business daily vz.lt led, and in the Polish brand data, YouTube led.

Why do Gemini citations need separate handling?

In this study, 23,027 Gemini citations arrived as Google redirect addresses instead of the real sources. The real domain sat in the citation title field. Left unresolved, the redirect becomes the top domain in every language, and Gemini's owned-citation rate reads 0.0% by label. Resolved, it is 5.8%.

Can these results be applied directly to the Korean market?

No. All 12 markets are in Europe, and there is no Korean-language data. This study can't tell you whether Korean answers lean on Wikipedia or Namuwiki, or how often Naver blogs and cafes serve as sources. Use the direction as a reference, but measure the numbers yourself.

References

  1. [1]Dmitrij Żatuchin, "How Large Language Models Source Brand Reputation Across Languages and Markets", arXiv 2026
  2. [2]Chen, Wang, Chen & Koudas, "Generative Engine Optimization: How to Dominate AI Search", arXiv 2025
  3. [3]Liu, Zhang & Liang, "Evaluating Verifiability in Generative Search Engines", Findings of EMNLP 2023
  4. [4]Vladimir Pitenin, "Invisible to the Machine: Auditing AI Restaurant, Café, and Bar Recommendation Against a Complete Market Census", arXiv 2026

Summary

  • This study classified 167,551 URL citations that AI attached to answers about 128 brands across 12 European markets and 13 languages, by domain and source type.
  • 85.7% of citations pointed to sites the brand does not own, and 14.3% to owned sites. Even the most self-cited brands reached only about a third.
  • 80% of citations came from about 18% of domains, and the domain-rank distribution fit a Zipf law with exponent 0.86.
  • Wikipedia led 11 of 12 languages but at a 4–6% share; in Lithuanian the business daily vz.lt led, and for Polish brands YouTube led.
  • 23,027 Gemini citations arrived as Google redirect addresses, and only resolving the real domain from the title field corrected the owned-citation rate (0.0% to 5.8%) and the top domain.

Check this topic against your own brand

TRAIL Search measures how ChatGPT and Perplexity answer your customers' questions and finds and fixes where the brand is missing, turning this post into a diagnosis. Start with 10 questions, no card required.

More posts

Keep reading on this topic

This post is part of the Papers category, which collects all 15 posts on the topic. See all posts in Papers