[FREE TOOL] Common Crawl, LLM Training Data, and the Domain Authority Question
If you've been following the AI industry lately, you've probably noticed the growing tension between AI companies and content creators.
Lawsuits are piling up. Publishers are blocking crawlers. And in November 2025, The Atlantic published an investigation that put a relatively unknown nonprofit at the center of the controversy.
That nonprofit is Common Crawl.
For most of its 17-year existence, Common Crawl operated quietly in the background, archiving the web and making that data freely available to researchers. But since GPT-3 launched in 2020, it has become the backbone of most large language models you use today.
This article does three things:
Summarizes the current Common Crawl controversy with verified facts
Examines what we actually know about LLM citation patterns
Introduces a free tool I built to explore Common Crawl’s WebGraph data
Let’s dig in.
The Common Crawl Controversy (2023-2025)
The Atlantic Investigation
On November 4, 2025, journalist Alex Reisner published an investigation that sent shockwaves through the industry. His reporting revealed that Common Crawl has been supplying paywalled content from major publishers to AI companies, despite publicly claiming it only scrapes “freely available content.”
The key findings:
The paywall workaround. Many news sites show the full article text briefly before JavaScript checks your subscription status and hides the content. Common Crawl’s scraper never executes that JavaScript code. Result: it captures the full articles anyway.
By Reisner’s estimate, CC’s archives contain millions of articles from The New York Times, The Wall Street Journal, The Economist, The New Yorker, Harper’s, The Atlantic, and others.
The takedown problem. Multiple publishers have requested content removal. The New York Times sent a removal notice in July 2023. The Danish Rights Alliance initiated requests in July 2024. But Reisner’s research found that none of these requests appear to have been fulfilled, content files in the archives haven’t been modified since 2016.
The misleading search tool. CC’s public search function returns “no captures” for domains like nytimes.com, even though articles from those domains exist in the archives. Reisner discovered over 1,000 domains producing these incorrect results, most belonging to publishers who requested removal.
The quotes. Rich Skrenta, CC’s Executive Director, made some memorable statements:
“You shouldn’t have put your content on the internet if you didn’t want it to be on the internet.”
He also told Reisner that The Atlantic “is not a crucial part of the internet” and that “whatever you’re saying, other people are saying too, on other sites.”
Common Crawl’s response came the same day. They denied lying to publishers and stated: “Our web crawler, known as CCBot, collects data from publicly accessible web pages. We do not go ‘behind paywalls,’ do not log in to any websites, and do not employ any method designed to evade access restrictions.”
Skrenta also explained that the archive’s file format is “meant to be immutable”. You can’t delete anything from it.
One data point that’s hard to dispute: CCBot has become the most widely blocked crawler by the top 1,000 websites, surpassing even OpenAI’s GPTBot.
The Mozilla Report (February 2024)
Before The Atlantic investigation, Mozilla Foundation published the most comprehensive analysis of Common Crawl’s role in AI development. Their report, “Training Data for the Price of a Sandwich,” revealed just how foundational CC has become.
The numbers:
64% of 47 LLMs analyzed (2019-2023) used at least one filtered version of Common Crawl
GPT-3: Over 80% of tokens came from filtered CC data
Archive size: 9.5+ petabytes containing billions of web pages
Academic impact: 10,000+ papers have cited Common Crawl
But here’s what caught my attention: Mozilla documented how Common Crawl decides what to crawl.
CC uses Harmonic Centrality to prioritize domains. This metric measures how “close” a domain is to all other domains in the web’s link graph. Higher scores mean the domain is crawled more frequently.
The implication? Domains with lower centrality scores—including what Mozilla called “digitally marginalized communities”—are less likely to be included.
As CC’s main crawl engineer told Mozilla:
“Often it is claimed that Common Crawl contains the entire web, but that’s absolutely not true. Based on what I know about how many URLs exist, it’s very, very small.”
Mozilla also found that popular filtered versions of CC (like Google’s C4 and EleutherAI’s Pile-CC) use “simplistic automated filtering” that leaves problematic content untouched while sometimes removing non-toxic content from marginalized communities.
The Washington Post Analysis (April 2023)
Before the Mozilla report, The Washington Post analyzed Google’s C4 dataset—a filtered version of Common Crawl used to train models like T5.
What they found inside:
15 million websites in the dataset
Top sources include patents.google.com, nytimes.com (#4), theguardian.com (#7)
Also present: RT.com (#65), Breitbart (#159), and vdare.com (#993)
Personal blogs, voter registration databases, and copyrighted content
As the Post noted: “Experts say many companies do not document the contents of their training data.”
The Financial Connections
In 2023, after 15 years of near-exclusive funding from founder Gil Elbaz’s family trust, Common Crawl started receiving donations from AI companies:
OpenAI: $250,000
Anthropic: $250,000
NVIDIA: Listed as “collaborator”
Amazon Web Services: Sponsors data hosting
Make of that what you will.
What We Know About LLM Citation Patterns
Let’s shift from controversy to data. Multiple studies have analyzed which domains LLMs actually cite when generating responses.
Semrush Study (June 2025)
Analysis of 150,000+ citations across 5,000 keywords:
DomainCitation FrequencyReddit40.1%Wikipedia26.3%Google23%YouTube23%
Reddit’s dominance isn’t entirely organic—it likely reflects Google’s $60 million API licensing deal with Reddit in early 2024.
Profound Study (August 2024 - June 2025)
The largest analysis I’ve seen: 680 million citations tracked across ChatGPT, Google AI Overviews, and Perplexity.
Key findings by platform:
ChatGPT:
Wikipedia: 7.8% of total citations
Within ChatGPT’s top 10 sources, Wikipedia accounts for 47.9%
Heavy preference for authoritative knowledge bases
Perplexity:
Reddit: 6.6% of total citations
Strong emphasis on community-driven content and real-time sources
Google AI Overviews:
Reddit: 2.2% of citations
More balanced mix: Reddit (21%), YouTube (18.8%), Quora (14.3%) among top 10
Across all platforms:
.com domains: 80.41% of all citations
.org domains: 11.29%
Search Atlas Study (August-September 2025)
Analysis of 5.17 million citations across 907,003 unique domains:
Commercial domains dominate
Academic and government sources remain underrepresented
Their conclusion: “LLM citations reflect the structure of the public web rather than institutional authority”
What Actually Drives Citations?
Based on the research, here’s what we know:
Confirmed factors:
Content freshness and recency (40-60% of cited sources change monthly)
Semantic relevance to query
Structured data and formatting
Cross-platform presence and brand mentions
Platform-specific retrieval preferences
Possible contributing factors:
Frequency of domain appearance in training data
Authority signals embedded in parametric knowledge
Brand search volume and entity recognition
The recency factor is significant. LLMs aren’t static—they’re constantly updating which sources they cite based on real-time retrieval.
The Underexplored Question
Here’s what most coverage of Common Crawl misses: CC doesn’t just archive web pages. It also publishes WebGraph data.
Every month, CC releases domain authority metrics for 94-163 million domains:
Harmonic Centrality (HC): How “close” a domain is to all other domains in the link graph
PageRank: Authority based on quality and quantity of incoming links
This is the same data CC uses internally to decide what to crawl more frequently. And it’s publicly available.
The Open Questions
1. Training data composition
If CC prioritizes crawling high-HC domains, those domains appear more frequently in training data. Does this create a baseline familiarity in LLMs with certain sources?
2. Correlation vs. causation
The domains that rank highest in CC’s WebGraph (Facebook, Google, YouTube, Wikipedia) are also among the most-cited by LLMs. Is this because:
They’re genuinely authoritative (likely primary factor)
They were overrepresented in training data (possible contributor)
They perform well on real-time retrieval (confirmed factor)
All of the above
3. The long tail problem
Mozilla noted CC’s crawling underrepresents “digitally marginalized communities.” If your domain is in CC’s long tail (rank >1M among 607M domains), does that correlate with citation difficulty—regardless of content quality?
4. Authority threshold
Is there a minimum domain authority level below which LLMs rarely cite a source, even with excellent content and freshness?
I don’t have definitive answers to these questions. But I built a tool to help explore them.
Introducing CC Rank Checker
I created a free tool that makes Common Crawl’s WebGraph data accessible:
Check your domain’s CC authority:
What you can do:
Check any domain’s HC Rank and PageRank across 607 million domains
View rank history across 5 time periods (2023-2025)
Track authority changes over time
Compare up to 10 domains simultaneously
Browse the Top 1000 domains
What the Data Shows
Here are the top 15 domains from the latest WebGraph (October-November-December 2025):
Interesting observations:
Wikipedia is #14 in Harmonic Centrality but only #37 in PageRank—yet it’s ChatGPT’s most-cited source at 7.8% of total citations
CDN and infrastructure domains (gstatic, cloudflare, jsdelivr) rank extremely high because they’re embedded across millions of sites
Social platforms dominate the top 10
How to Use This Data
Benchmark your domain against competitors
Track authority trends over time (are you gaining or losing ground?)
Investigate correlations between CC rank and citation frequency in your niche
Identify authority gaps that content optimization alone may not solve
What This Means for AEO
I want to be clear: I’m not claiming CC authority is the primary driver of LLM citations. The research shows content freshness, relevance, and platform-specific factors matter significantly.
But domain authority likely plays some role in the equation. Here’s my balanced take:
Confirmed factors (focus here first):
Content quality and relevance
Freshness and recency
Structured formatting
Cross-platform presence
Brand search volume
Possible contributing factors (worth monitoring):
Historical training data presence
WebGraph-derived authority signals
Entity recognition and associations
Practical takeaways:
Don’t ignore authority. While content and freshness matter most, domain-level signals probably contribute to the overall picture.
Track multiple metrics. CC Rank is one data point—not a silver bullet, but useful for benchmarking.
Understand platform differences. Wikipedia dominates ChatGPT. Reddit dominates Perplexity and Google AI Overviews. Optimize accordingly.
The long tail question. If your domain ranks >1M in CC, it’s worth investigating whether this correlates with citation challenges in your space.
More research needed. The relationship between CC authority and LLM citations deserves rigorous empirical study. I’m planning follow-up analysis.
Conclusion
The Common Crawl debate is about more than copyright and paywalls. It’s about understanding the foundation of AI’s knowledge—and potentially, its tendencies toward certain sources.
What we know:
Most major LLMs were trained on CC data (64% of models, 80%+ of GPT-3 tokens)
CC prioritizes high-authority domains via Harmonic Centrality
These same domains tend to be cited most frequently by LLMs
What we don’t know:
How much training data composition directly influences citations vs. real-time signals
Whether CC authority metrics have independent predictive value
How these factors interact with confirmed signals like freshness
The CC Rank Checker is my small contribution to making this data accessible. The bigger questions require more research, more data, and more transparency from AI companies about their training data composition.
Check your domain’s CC authority: https://webgraph.metehan.ai
References
Investigations & Reports
Reisner, A. (November 4, 2025). “Common Crawl Is Doing the AI Industry’s Dirty Work.” The Atlantic.
Baack, S. & Mozilla Insights (February 6, 2024). “Training Data for the Price of a Sandwich.” Mozilla Foundation.
Schaul, K., Chen, S.Y., & Tiku, N. (April 19, 2023). “Inside the secret list of websites that make AI like ChatGPT sound smart.” Washington Post.
Citation Studies
Semrush (June 2025). LLM Citation Analysis — 150,000+ citations
Profound (August 2024-June 2025). AI Platform Citation Patterns — 680 million citations
Search Atlas (August-September 2025). Industry Patterns in LLM Responses — 5.17 million citations
Academic Papers
Brown, T. et al. (2020). “Language Models are Few-Shot Learners.” arXiv:2005.14165




🔥🔥🔥 good job Metehan
This is such a great tool with valuable insights. Metehan, you are the golden nugget of the AEO/SEO deep dives and experimentation. Glad I followed you early on.