Does AI Search Cite AI-Written Content? About 16% of Sources

Does AI Search Cite AI-Written Content? About 16% of Sources

An AIES 2026 audit of ChatGPT, Copilot, Gemini and Perplexity citations, read next to Google's spam policy

Does AI search cite AI-written content? About 16% of sources four AI answer engines cited were AI-generated. How engines differ and what Google's policy says.

By · TRAIL Labs Research
GEOAEOAI CitationsAI-Generated ContentGoogle Spam UpdateMeasurement

About 16% of the sources that four AI search engines cited in their answers were classified as AI-written. A Northwestern University team sent 712 real user queries to ChatGPT (with web search), Copilot, Gemini and Perplexity, then ran the text of every cited page through an AI-generated text detector [1]. The engines differ a lot: Copilot was at 27.8% and ChatGPT at 7.3% [1]. Over the same period, Google started its fourth spam update of the year on September 24 and finished around October 8 [4].

We read the paper on arXiv (v1), including the appendix tables, and checked every number against the table values. All figures are from the paper's experiments, and we did not reproduce them [1]. Facts on Google's side come only from Google Search Central's spam policies, its generative AI content guidance, and the Search Status Dashboard [2][3][4]. We did not use personal observation posts about how the update hit particular sites. The paper states two values differently in two places, and there we followed the tables and the results section. We also build AI search diagnosis tools, so anything that connects to our products sits in the last section.

Does AI search cite content written by AI?

It does. AI-generated sources showed up in all four engines, and overall 3,056 of the 19,154 classified sources, about 16%, were labeled AI-generated [1]. The authors read this as a lower bound on the true share. On a test of articles from sites reported as AI content farms, all assumed AI-written, the detector labeled about 31% human [1].

The question matters because users cannot tell sources apart. A link under an answer looks the same whether it points to a government site or an automatically generated niche site. The authors do not assume all AI-generated content is low quality, but they flag the risk that sources able to carry over hallucinations and bias get mixed into answers with the same weight as authoritative ones [1].

For marketers the question also runs the other way. If AI-written content does not stop a page from being cited, should you mass-produce content with AI, and would Google Search penalize that? This paper answers the first question. Google's official documents answer the second.

A card showing the share of sources cited by four AI answer engines that were classified as AI-generated, a distribution bar, and counts of queries, engines and URLs

Figure 1. How the cited sources were classified.

How did the study count AI-generated sources?

The team typed real user queries into each engine's interface, scraped the pages each answer cited, and classified them with a detector [1]. From Search Arena, a dataset of real generative search sessions, they took 175 politics and 257 health queries, and from a climate Q&A dataset, 280 environment queries [1]. All are in English and limited to US-related questions.

They automated each engine's web interface instead of using the API. Earlier work found that API and interface answers differ, so this choice captures the sources users actually see [1]. Sending 712 queries to four engines produced 2,848 responses, and 91.2% of them carried citations [1]. The cited sources added up to 26,266 unique URLs on 7,675 domains.

They compared two detectors, Pangram and GPTZero. On a test of 200 human-written and 200 AI-written texts, both were perfect [1]. On a second test of 105 recent articles from seven domains that journalists had reported as AI content sites, Pangram missed 31.4% and GPTZero about 30.4% by labeling them human [1]. To avoid inflating the share, the authors picked Pangram, the one that misses slightly more.

Text extraction worked for 19,154 pages (72.9%). The rest were PDFs, videos or images and fell outside the analysis [1]. Of Pangram's four categories, the ambiguous Possibly AI was dropped, and only Highly Likely AI and Likely AI counted as AI-generated [1].

Pangram labelSourcesShare
Unlikely AI (classified human)15,81582.5%
Highly Likely AI2,91615.2%
Likely AI1400.7%
Possibly AI (excluded)2831.4%

Table 1. Classification of the 19,154 cited sources with extracted text. AI-generated is Highly Likely plus Likely, 3,056 sources [1].

How much of what ChatGPT or Perplexity cites is AI-generated content?

ChatGPT at 7.3% and Perplexity at 9.4% were on the low end of the four, and Copilot was highest at 27.8% [1]. Gemini was at 14.7%. The highest and lowest engines are nearly 4x apart.

EngineAvg. citations per responseUnique sourcesAI-generated share
Copilot8.446,00727.8%
Gemini7.265,16814.7%
Perplexity7.925,6369.4%
ChatGPT14.6810,4537.3%

Table 2. Citation volume and AI-generated share by engine. Volume is from the paper's Table 1 and shares from the results section [1].

ChatGPT cited the most sources per response, 14.7 on average, and still had the lowest AI-generated share [1]. Citing more sources does not by itself bring in more AI-generated ones. Across the 12 topic and engine cells, Copilot cited the most AI-generated sources in every topic [1].

By topic, 16.3% of sources cited for health queries were AI-generated, 13.3% for the environment and 11.1% for politics [1]. In raw counts the environment led with 1,633. Health queries also had the highest rate of answers with no citations at all, 11.6% [1].

A card showing bars for the AI-generated source share by engine and AI-generated shares for Wikipedia, PMC, Reddit and Times of India

Figure 2. AI-generated share by engine and by domain.

Where do AI-generated sources come from?

Mostly from outside the big, frequently cited sites. 97.1% of AI-generated sources came from outside the 25 most-cited domains [1]. Those 25 domains account for 23.8% of all citations, and the rest spreads across a long tail [1]. 59.1% of cited domains were cited only once, and 16.5% only twice [1].

Inside the top 25, the engines' anchors are clear. Wikipedia alone takes 15.2% of top-25 citations, and government sites together take 33.0% [1]. Social media such as Reddit and YouTube take 8.8%, academic publishers and ResearchGate 14.6%, and news outlets barely appear [1]. The Gini index of domain concentration was 0.68 overall, and by engine ChatGPT 0.648, Gemini 0.595, Perplexity 0.563 and Copilot 0.492 [1]. Copilot, the engine with the highest AI-generated share, is also the one that spreads its sources most widely. With only four engines, it is too early to read that as a relationship.

Big sites run low. Of 1,180 Wikipedia pages cited, 15 (1.3%) were AI-generated, and of 1,078 pages from the medical paper archive PMC, 44 (4.1%) [1]. Reddit was at 61 of 520 pages (12%) and the Indian daily Times of India at 35 of 78 (45%) [1].

At the other end, several domains had every cited page classified as AI-generated. Many were niche sites on topics such as climate and health, and five subdomains of the same directory site appeared in the table together [1]. The authors opened 200 pages classified as AI-generated by hand, and not one disclosed AI use [1].

This shape matches what we saw in the study on where AI gets brand information. Engines lean on a few big domains while drawing widely on a long tail cited once or twice. The authors frame the difficulty of vetting that long tail as a design trade-off between source diversity and quality checks [1].

Does Google's spam update penalize AI-written content?

Google's official documents look at whether content was mass-produced without value more than whether AI wrote it. The scaled content abuse section of the spam policy defines generating many pages primarily to manipulate rankings, and its first example is using generative AI tools or similar tools to generate many pages without adding value for users [2]. The same section treats unoriginal content with little to no value as a problem no matter how it's created [2].

The Google Search Status Dashboard shows four spam updates this year: March 24, June 24, August 18, and the September update that started on September 24 and rolled out over about 13 days and 16 hours [4]. The policy page was updated on August 28 [2].

The generative AI content guidance adds one more condition. It says it is critical to manually fact-check and review all AI-generated content for accuracy and trustworthiness before publishing, and that this review also applies to title elements, meta descriptions, structured data and image alt text [3]. That guidance was updated on October 1 [3].

A card showing a timeline of the four 2026 Google spam update start dates and quotes from the spam policy and the generative AI guidance

Figure 3. Google's 2026 spam update schedule and policy wording.

After an update, several observation posts claim that AI content is being dropped from the index at scale. The patterns they describe may be real, but ranking screenshots and individual cases cannot separate causes. What the public record supports is the policy wording above.

Is it okay to publish AI-generated content on a blog?

By Google's standard you can, on the condition of value and human fact-checking [3]. Read together, the two sources point the same way. AI answer engines cite a substantial number of sources classified as AI-generated [1], and Google Search treats valueless mass production as spam regardless of the tool [2]. So asking whether a page can be verified fits both sides better than asking whether AI wrote it.

A verifiable page shows it through a few habits. Any paragraph with numbers or statistics links to the original source in the same paragraph. A person checked the facts before publishing, and the page says what was checked. The site does not stamp out look-alike pages that only swap the topic. The sourcing and author details covered in our E-E-A-T content checklist point the same way.

The study also shows a gap. None of the 200 pages classified as AI-generated disclosed AI use [1]. Google's guidance suggests giving readers context about how content was created [3]. One paragraph that states what a person checked gives both readers and search engines something to judge by.

Limitations

This study does not test what makes content get cited. It counts the makeup of cited sources, so it cannot tell you whether AI-written content helps or hurts your chances of being cited [1].

There is no web-wide baseline for the share of AI-generated pages, so the paper alone cannot say whether 16% is high or low [1].

Classification rests on a single detector, Pangram. Another detector could give a different share, and the authors advise reading the result as a lower bound [1]. "AI-generated" is not a quality judgment either. The authors did not separately evaluate accuracy or claims.

The scope is narrow. It covers English queries on US-related topics, single-turn questions, and three topics: politics, health and the environment [1]. The 27.1% of sources whose text could not be extracted (PDFs, videos, images) fall outside the analysis. Korean queries and Naver AI Briefing were not measured.

The paper states two values in two ways. The data section text says 257 politics and 175 health queries, while the final sample and Table 2 say 175 politics and 257 health. The last paragraph of the limitations section says "approximately one in five," while the results section and the abstract say about 16% [1]. We followed the tables and the results section.

At TRAIL

Our products sit on the same standard. When TRAIL Search drafts an article, it counts whether each paragraph with numbers has a source link and rewrites any paragraph that lacks one. New posts on our own blog have to pass the same check before they can be committed.

On the other hand, our products have no feature that detects whether content was written by AI. Our check looks at verifiability, and Google's policy also judges value and human fact-checking [2][3].

Measurement work remains. If the AI-generated share differs by nearly 4x across engines [1], the kinds of sources that get competitors cited may also differ by engine. Our diagnosis splits cited sources into owned, wiki, YouTube, reviews and communities, media and listicles, but does not yet separate suspected AI-generated sources as their own category. No public number yet shows this share for Korean-language answers. We cover the gap between what engines read and what they cite in the difference between retrieval and citation.

Frequently asked questions

Do AI search engines cite content written by AI?

Yes. In a study that sent 712 real user queries to ChatGPT (with web search), Copilot, Gemini and Perplexity, about 16% of the 19,154 cited sources whose text could be classified were labeled AI-generated. The detector misses some AI text, so the authors read this as a lower bound.

Which AI search engine cites the most AI-generated sources?

In this study Copilot was highest at 27.8%, followed by Gemini at 14.7%, Perplexity at 9.4% and ChatGPT at 7.3%. Copilot ranked first in all three topics: politics, health and the environment. The study covers English queries on US-related topics only.

Does Google's spam update penalize AI-written content?

Google's official documents judge value more than the tool used. The scaled content abuse policy lists using generative AI tools to generate many pages without adding value for users as an example, and treats unoriginal content with little to no value as a problem no matter how it's created. Google's AI content guidance asks for human fact-checking before publishing.

Is it okay to publish AI-generated content on a blog?

By Google's standard you can, with conditions. The content has to add value for users, and a person has to check it for accuracy before it goes live. In this study, none of 200 pages classified as AI-generated disclosed AI use. Sourcing every number and not mass-producing look-alike pages fits both standards.

References

  1. [1]Mowafak Allaham & Nicholas Diakopoulos, "Synthetic Sources?: Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources", AIES 2026
  2. [2]Google Search Central, "Spam policies for Google web search" (updated 2026-08-28)
  3. [3]Google Search Central, "Google Search's guidance on using generative AI content on your website" (updated 2026-10-01)
  4. [4]Google Search Status Dashboard, Ranking incident history

Summary

  • An AIES 2026 study sent 712 real queries to ChatGPT, Copilot, Gemini and Perplexity and ran the text of every cited source through an AI detector.
  • About 16% of the 19,154 classified sources were labeled AI-generated, a lower bound because the detector labeled 31.4% of articles from reported AI content sites as human.
  • By engine, Copilot was at 27.8%, Gemini 14.7%, Perplexity 9.4% and ChatGPT 7.3%, a gap of nearly 4x.
  • 97.1% of AI-generated sources came from outside the 25 most-cited domains, and none of 200 pages classified as AI-generated disclosed AI use.
  • Google shipped four spam updates this year, and its official policy judges value and human fact-checking more than the tool used.

Check this topic against your own brand

TRAIL Search measures how ChatGPT and Perplexity answer your customers' questions and finds and fixes where the brand is missing, turning this post into a diagnosis. Start with 10 questions, no card required.

More posts

Keep reading on this topic

This post is part of the Papers category, which collects all 17 posts on the topic. See all posts in Papers