Connect with us

AI

European Publishers Feed a Crawl That Sends No Readers

European publishers take four times the AI scrapes of North American sites as Common Crawl’s European share hits 30.04 percent.

Published

on

European news sites took four times as many AI scrapes as North American ones in the first half of 2026, TollBit found, and sent almost no readers back. The open crawl that feeds model training has spent years loading up on European domains, and the labs that need those languages still treat a robots.txt file as a preference.

TollBit, which sells bot tracking and paid access to publishers, read named AI bots from 40 scraping vendors across 3,906 publisher sites. The European slice of that panel is paying the server bill for a multilingual data race it did not start.

TollBit’s Sample Puts Europe at Four Times the Load

Median AI scrapes per site were four times higher in Europe than in North America. European publishers got one human referral visit from AI apps for every 179 AI bot visits, a return more than three times worse than on North American sites. That is the fourfold scrape gap on European sites, measured on the same vendor panel.

In the first half of 2026, just 0.05% of external referrals to European sites came from an AI app, against 0.16% in North America. TollBit also counted about one AI bot scrape for every 33 human visits on the European side. The median European site’s robots.txt rules against scraping were ignored nearly three times more often than on North American sites.

EUROPE VERSUS NORTH AMERICA, H1 2026

Metric Europe North America
Median AI scrapes per site Four times higher Baseline
AI app share of external referrals 0.05% 0.16%
Human referral per AI bot visits 1 per 179 More than 3× better
Sites that disallow Claude-User 9% 26%
Sites that disallow Perplexity-User 13% 26%

Olivia Joslin, TollBit’s cofounder, said the split may come from Europe’s many languages, as models try to learn them and pull copy from more sources. The company’s own docs say it mainly tracks large, self-identified bots such as ChatGPT, Perplexity and Claude, not the long tail of scrapers that pretend to be browsers.

A Crawl That Tilted Toward Europe

The language theory has a supply-chain paper trail. Common Crawl, the nonprofit archive that many model builders still filter for pre-training text, has shifted hard toward European country-code domains over 17 years of monthly snapshots.

EUROPE’S SHARE OF COMMON CRAWL

  1. 2009: European mapped TLDs are 4.75% of pages;.com and.net are 80.38%.
  2. 2016: Europe’s share is 12.77% as the crawl widens past generic domains.
  3. 2020: Europe reaches 28.61% of the archive.
  4. 2026: Europe is 30.04%;.com and.net together are 44.84%, and Germany’s.de alone is 4.38% of pages.

Common Crawl’s own tables put European domains at 30.04% of the 2026 crawl, more than six times the 4.75% share in 2009. Grzegorz Piechota, researcher-in-residence at INMA, had already flagged the same tilt. He argued that consumer demand for non-English information is steering AI firms and their data suppliers, and he pointed at research drives from Google, OpenAI and Microsoft aimed at multilingual users, especially in Europe.

Consumer demand for non-English language information is influencing the strategies of AI companies and their suppliers of data.

Grzegorz Piechota, researcher-in-residence, INMA

Anthropic’s September 2025 Economic Index put the United States at 21.6% of Claude usage, with India at 7.2% and Brazil at 3.7%. Piechota read those figures as evidence that English-speaking countries make up less than a third of Claude users, so most usage, and by extension much of the scraping that supports it, is tied to non-English copy. That is a reading of one lab’s chat product, not a census of every crawler. It still matches the direction of the archive that feeds so many of them.

Why the Other Dashboards Disagree

Cloudflare Radar does not show European publishers taking more total bot hits than North American ones. Lai Yi Ohlsen, who runs the dashboard, said the absolute number of bot requests is consistently higher for North America. On Cloudflare’s media and publishing customers, though, European sites have a higher share of all requests coming from bots. North America has more bot activity in bulk. Europe has a heavier bot mix on the sites themselves.

Jérôme Segura, vice president of threat research at DataDome, said his firm’s logs do not show the same steady regional gap. Variance from publisher to publisher is huge, he said, and it likely tracks how prominent a site is and how much of the scraping comes from bots that name themselves. On some individual publishers, regardless of region, he saw AI traffic as high as roughly one AI visit for every 10 human visits.

WHERE THE COUNTS DISAGREE

  • TollBit: Median scrapes per European site are four times the North American median, on a panel of 3,906 publishers that includes 456 European titles from the Telegraph to Ringier to Styria.
  • Cloudflare Radar: North America still leads in raw bot request volume; European media sites show a higher bot share of their own traffic.
  • DataDome: No stable Europe-versus-North-America gap; a single large title can sit at one AI visit per 10 human visits once you leave regional averages behind.

Piechota also warned that TollBit’s customer mix could be doing some of the work. A panel weighted toward European local and sports titles will look different from a network weighted toward huge English-language destinations. The Common Crawl shift does not depend on that panel. The scrape-to-referral collapse on TollBit’s European sites still does.

The Disallow Lines Europe Never Wrote

TollBit’s first-half 2026 report also scored a narrower class of traffic: page fetchers, the agents that load a URL in real time because someone asked a chat app a question. About 15% of those named fetchers on European sites reached URLs the sites had marked as disallowed. ChatGPT-User, Bytespider and Youbot each hit disallowed pages on nearly half of the European sites that had listed them. ChatGPT-User reached the most sites.

The other half of that file is empty. Only 9% of European websites disallow Claude-User, Anthropic’s fetching agent, compared with 26% in North America. Perplexity-User sits at 13% in Europe and 26% in North America. Most of the newest fetchers sit in single-digit disallow rates across Europe. ChatGPT-User is the outlier that sites actually name, which is why it also racks up the most recorded bypasses. An agent that never appears in robots.txt cannot violate a rule nobody wrote.

A low bypass count on a new bot is not proof of good manners. It can just mean the publisher never added the line. Europe is writing fewer of those lines than North America, then watching the ones it did write get skipped more often.

227 Scrapes for Every Referred Reader

The scrape load on European sites was nearly 20% higher in June 2026 than in January. Local news and sports titles posted the highest quarter-over-quarter growth. The scrape-to-referral ratio on the European side moved from 150 to 1 in the first quarter of 2026 to 227 to 1 in the second. That quarterly series is a different cut from the 179 to 1 H1 referral ratio, and both point the same way: more taking, less sending back.

Local newsrooms are a useful target if you are starving a model of German match reports, Polish city-hall copy or French club coverage. They publish all day, they sit on country-code domains the archive now weights more heavily, and they rarely have the bot-management staff of a global English title. Sports desks add another lure, structured stats and timely recaps that assistants like to quote.

The return path is the part that does not scale with that hunger. AI apps accounted for 0.05% of external referrals to European sites in the first half of 2026. The bots keep visiting. The readers mostly do not.

OpenAI Treats a Chat Click as Permission

Fetchers are not the same software as training crawlers. GPTBot collects copy to train models. OAI-SearchBot is the agent OpenAI tells webmasters to use if they want to appear in ChatGPT search. ChatGPT-User is the live fetch that runs when a person asks about a page. OpenAI’s crawler overview says that because those actions are initiated by a user, robots.txt rules may not apply. ChatGPT-User is not used to decide whether a site shows up in Search, and the company tells publishers to manage that opt-out with OAI-SearchBot instead.

Perplexity takes a similar line on Perplexity-User. Anthropic says all three of its bots respect the file. TollBit still counts any request to a disallowed URL as a bypass, whatever the operator’s docs claim. Three large assistant makers now hold two different positions on whether the same text file governs the same live fetch.

A publisher that blocks both OAI-SearchBot and ChatGPT-User gives up the search listing OpenAI says it will honor, and keeps a fetch control OpenAI says it may ignore. The file still matters for training crawlers that choose to read it. It is a weak lock on the agent that shows up because a reader pasted a URL into a chat box.

September 15 Changes the Default on Ad Pages

Cloudflare spent 2025 and 2026 trying to move that fight off the text file and onto the network. On July 1, 2026, Jin-Hee Lee and Bryan Becker described a split of automated traffic into three jobs, plus new defaults for Training and Agent crawlers that take effect for new domains on September 15, 2026.

THE THREE CRAWLER JOBS

  • Search: Indexing a site so a later query can point back, the one class Cloudflare still wants allowed by default because it can send people.
  • Agent: A live fetch or browser agent acting for a person right now, including chat fetch bots such as ChatGPT-User.
  • Training: A crawl whose job is to absorb the page into a model.

On September 15, 2026, new domains joining Cloudflare will block Training and Agent crawlers by default on pages that display ads. Search stays allowed. Multi-purpose crawlers that mix Search with Training will be judged on all of their jobs, so a block on Training can also catch Googlebot, Applebot and BingBot unless the site owner opts out. Existing customers can freeze their settings before that date. The Agent bucket is the same class TollBit measured walking through disallow lines in Europe.

Cloudflare is also testing a robots.txt content-use signal (immediate, reference or full) on top of older Content Signals language that already separates search from training. Bots that reproduce pages in full cannot hold Verified status on its network. That is still a preference until the edge rule fires.

European publishers sit at the loud end of this mismatch. The archive that supplies training text now draws more than six times the European-domain share it did in 2009. TollBit’s European panel is taking four times the scrapes of its North American one and sending back one referred reader per 179 bot visits. On September 15, 2026, new Cloudflare domains will block Training and Agent crawlers on ad pages by default. The file that was supposed to sort those visits remains optional once a person clicks.

Harry is the editor of Oton Technology, an independent site he owns and edits, covering the part of technology that people actually have to act on. After ten years in journalism, first reporting and then editing, he works from primary material by habit: the advisory rather than the write up of it, the filing rather than the press release, the changelog rather than the launch video. Every figure in an article carries its source and its date, and where a number comes from a vendor or an analyst model rather than a count, he says so plainly instead of letting it stand as established fact. What he leaves out is anything he could not verify himself, which on a beat full of unnamed supply chain claims removes a great deal. That standard applies across all the sections the site publishes for an international audience, from artificial intelligence and security to phones, computers, gaming, crypto and the software businesses depend on. He corrects errors in the open and labels them, because a site that hides its mistakes is asking readers to trust the rest on nothing.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending