Connect with us

AI

European Sites Face Four Times the AI Scrapes With Worse Returns

TollBit data shows European sites take four times the median AI scrapes of North American peers and far fewer referrals as multi-language model demand rises.

Published

on

European publishers face median AI scrapes four times higher than North American sites and receive just one human referral from AI apps for every 179 bot visits, according to TollBit’s latest analysis of thousands of publishers. The scrape-to-referral gap widened from 150:1 in Q1 2026 to 227:1 in Q2, and local news and sports saw the fastest growth.

The pattern sits inside a larger multi-language push by AI labs. European content is feeding models at rising rates while the traffic coming back stays thin.

The TollBit Numbers Side by Side

TollBit identified AI bots from 40 scraping vendors across 3,906 publishers, including 456 European titles ranging from The Telegraph to Ringier and Styria. The TollBit State of the Bots Q1 and Q2 2026 report supplies the regional split.

Metric European sites North American sites
Median AI scrapes per site 4× higher Baseline
AI bot visits per human referral from AI apps 179:1 Roughly one-third that rate
Share of external referrals from AI apps (H1 2026) 0.05% 0.16%
AI bot scrapes per human visit (approx.) 1 per 33 Lower load
Robots.txt “do not scrape” ignored (median) Nearly 3× more often Baseline

AI scrapes on European sites rose nearly 20% from January to June 2026. Across TollBit’s network the firm logged 987 billion-plus website visits analyzed, 22 billion-plus AI bot scrapes detected, and 1.9 billion-plus scrapes that bypassed or ignored robots.txt instructions.

Local News and Sports Absorb the Climb

Quarter-over-quarter growth hit local news and sports hardest. Those categories already operate on thinner margins and smaller technical teams than national or lifestyle brands.

  • Local news sites saw some of the strongest scrape increases as models seek current, place-specific information.
  • Sports properties supplied timely scores, lineups and match reports that retrieval bots pull in real time.
  • National and lifestyle titles still face high absolute volumes but slower percentage jumps in the TollBit sample.

On the ground, operators report that even Cloudflare and robots.txt settings do not stop the load. Aggressive crawlers and third-party scrapers keep coming, and the bandwidth bill rises whether the bot self-identifies or not.

Language Demand Pulls More Crawlers Across the Continent

TollBit cofounder Olivia Joslin pointed to Europe’s language diversity. Large language models need training and retrieval data in many tongues, so crawlers fan out across more domains than the relatively English-heavy North American web.

Grzegorz Piechota, researcher-in-residence at INMA, reached a similar conclusion from usage data. He noted that Claude users in the United States accounted for 21.6% of total Claude usage in Anthropic figures from last September, while English-speaking countries together made up less than a third of global Claude users. Most activity, he argued, therefore ties to non-English content.

Consumer demand for non-English language information is influencing the strategies of AI companies and their suppliers of data.

Piechota said Google, OpenAI and Microsoft have all launched multi-language research and product efforts aimed especially at Europe. Suppliers responded. European country domains made up only 4.75% of pages captured by the Common Crawl open web crawl repository in 2009. By 2026 that share reached almost 29.98% in his analysis.

“The Big Tech and AI industry seems to be catching up with demand for high-quality information in multiple languages,” Piechota said. The Anthropic Economic Index geographic usage cuts show adoption remains concentrated in higher-income countries overall, yet the training and retrieval side still needs volume from many languages to serve those users.

Other Datasets Draw a Different Picture

Not every tracker sees the same European overweight.

Lai Yi Ohlsen, head of Cloudflare’s internet tracking dashboard Cloudflare Radar, said absolute bot request volumes stay consistently higher for North American sites. European publishers do show a higher percentage of their total site requests coming from bots in Cloudflare’s media and publishing customer data. In short, Europe carries a heavier bot share of its own traffic, while North America still sees more total bot activity.

Jérôme Segura, VP of threat research at DataDome, reported no consistent regional gap matching TollBit’s. Variance publisher by publisher is “massive,” he said, driven by prominence, size and how many bots self-identify. Some individual sites, regardless of region, already see AI traffic as high as roughly one AI visit for every ten human visits.

What we know

  • TollBit’s publisher sample shows clear Europe-over-North-America gaps on median scrapes, referrals and robots.txt compliance.
  • Cloudflare absolute volumes favor North America; bot share of traffic is higher in Europe for media customers.
  • DataDome sees large site-level swings that swamp regional averages.

What remains unconfirmed

  • Whether language mix, sample composition or both explain TollBit’s Europe premium.
  • How much of the extra European load comes from self-declared AI bots versus stealth third-party scrapers.

Piechota himself flagged that TollBit’s customer mix could influence the regional picture. The nearly 4,000-publisher set is not a random census of the open web.

Robots.txt Gets Ignored More Often on This Side of the Atlantic

The median European site’s instructions not to scrape were ignored nearly three times more often than the North American median. That gap matters because robots.txt remains the main voluntary signal publishers still control.

1.9 billion-plus bot scrapes on TollBit’s network bypassed or ignored robots.txt in the measured period.

Third-party scrapers advertise evasion and paywall penetration tools, then resell access, according to earlier TollBit work on the “leaky pipes” of the scraper ecosystem.

Mixed-use crawlers further blur intent, making it hard for a publisher to allow discovery while blocking training or retrieval.

A Cloudflare report on the agentic Internet market notes that more than one-third of crawler activity on its network still comes from mixed-use bots. Live Cloudflare Radar live bot traffic share dashboards continue to show bots as a large and rising fraction of HTML requests worldwide. Enforcement tools help, yet the underlying crawl demand keeps rising.

Turning Agent Traffic Into a Paid Gate

Publishers are not waiting for perfect data agreement. TollBit and similar platforms let sites monitor bots, set terms and charge for access. Arc XP, the Washington Post’s publishing platform, integrated TollBit so customers can redirect scrapers to a paywall. Nearly 20% of TollBit’s network has already earned money that way, ranging from hundreds to tens of thousands of dollars a month in earlier disclosures.

Scale figures from TollBit’s H1 2026 summary include more than 2.6 billion AI bots directed to its bot paywall. Some operators also experiment with selling visibility itself. One approach now pitched to advertisers treats AI citation and appearance metrics the way publishers once sold Comscore numbers; that is the territory covered when publishers selling AI visibility metrics to advertisers reframe the relationship.

The economics still look one-sided for many European titles. A 179:1 scrape-to-referral ratio leaves little organic discovery to replace the lost pageviews. Local newsrooms and sports desks that saw the sharpest scrape growth have the least spare capacity to negotiate licenses or run bot infrastructure. Absolute North American volumes remain high, yet the per-site median load and the referral return look comparatively better there in TollBit’s cut.

Every tracker agrees the scrapes are real. They disagree on the size of the Europe premium and on the precise mix of language demand versus sample bias. What is no longer in dispute is the direction: more multi-language content is entering the training and retrieval pipelines, robots.txt is a weak gate, and the referral bargain that once funded the open web continues to fray fastest where the new languages live.

Logan Pierce is a writer and web publisher with over seven years of experience covering consumer technology. He has published work on independent tech blogs and freelance bylines covering Android devices, privacy focused software, and budget gadgets. Logan founded Oton Technology to publish clear, no nonsense tech news and reviews based on real hands on testing. He has personally tested and reviewed dozens of mid range and budget Android phones, written extensively about app privacy, and built and managed multiple WordPress publications over the past decade. Logan holds a bachelor's degree in English and studied digital marketing at a certificate level.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending