Skip to content
RU

How bot traffic changed by 2026: one honest trend and five irreconcilable measurements

TL;DR. Between December 2025 and August 2026 the human share of HTML page requests fell from 47% to 38.1%.

Between December 2025 and August 2026 the human share of HTML page requests fell from 47% to 38.1%. This is one of the few comparisons in the field built on the same metric from one source rather than stitched from different reports.

Below: that trend with its limits, six vendors offering five different "leading" AI crawlers, and the metric that explains the tension around the subject — how many pages each platform takes per visitor it sends back.

A trend that can be measured honestly

Most "trends" in this area are stitched together from different vendors with different methodologies — which makes them not trends at all. But one comparison is sound: Cloudflare published the same metric twice — the breakdown of HTML page requests.

Who requests the page2 Dec 202520–26 Aug 2026Change
Human47%38.1%−8.9 pp
Non-AI bots44%52.8%+8.8 pp
AI bots4.2%5.7%+1.5 pp

In nine months humans stopped being the majority of those who open pages. The gain went mostly to ordinary bots, not AI ones: the AI-crawler share rose by just a point and a half.

Boundaries: this is Cloudflare’s network, not the internet, and HTML requests only. Counted across all HTTP requests (images and scripts included), bots come to 35.5% for the same week — see our breakdown of denominators.

Six vendors name five different leaders

There is no single answer to "which AI crawler is the most active". Measurements over overlapping periods:

VendorPeriodLeader
FastlyJanuary 2026Google Other — 36%
AkamaiJul–Dec 2025OpenAI — 37%
HUMAN Securitycalendar 2025OpenAI — about 69%
Radware2025 holiday peakOpenAI — 65%
DataDomeFebruary 2026Meta-ExternalAgent — about 25%
DataDomeQ2 2026Meta, both agents — a majority

These cannot be reduced to one number, and the reasons are clear: different customer verticals, different classification schemes (Fastly deliberately excludes ordinary search crawlers from its "AI" class; Akamai aggregates by vendor-owner rather than User-Agent string), and different units — rule triggers, requests, or fetches per site.

A separate caution about vendors’ own trends. Fastly’s Q3 2025 edition reported bots at 29%; its January 2026 edition reports 49%. A twenty-point jump between consecutive editions of one report speaks more of a classification change than a real surge, and the source offers no explanation.

What crawlers come for, and what they give back

Cloudflare classifies declared crawl purpose. For 20–26 August 2026: training 39%, training and search combined 37.4%, search 17%, user action 4.7%, undeclared 1.9%.

The practical meaning is clearer from another metric — how many pages a crawler takes per visitor it sends back (same week):

PlatformPages per visitor
Perplexity816.8
Anthropic707
OpenAI498.1
Microsoft35.7
Yandex25.5
Google5.6
DuckDuckGo1.7

The distance between 5.6 and 816.8 is essentially the distance between a search engine and a training crawler: the first takes a page in order to send a reader after it, the second so that no reader is needed.

Caveat: referral attribution depends on the Referer header, which AI clients send inconsistently. The ratios hold as orders of magnitude, not to the decimal.

A robots.txt disallow is not always honoured

Per TollBit’s H1 2026 data — the logs of 3,906 publishers — about 15% of AI crawler fetches went past a disallow directive. For individual agents the share is higher: ChatGPT-User at 54% of its own fetches, Bytespider 48%, PerplexityBot 42%.

An important caveat from the source itself: only self-identifying agents are counted, so the real volume exceeds the measured one.

From which follows a practical conclusion worth stating plainly: robots.txt is a request, not an enforcement mechanism. If content genuinely must be closed, that happens at the server — by User-Agent and address ranges, not by a line in a text file. The file is still needed: it records your position, and most crawlers honour it.

To check that your file parses the way you intended, use the robots.txt checker.

What cannot be used in this area

Vercel’s AI-crawler data is dated December 2024. Its figures (GPTBot 569M, Claude 370M, PerplexityBot 24.4M against Googlebot 4.5B) are the most-reprinted in this space, they are roughly twenty months old, and Vercel has published no newer study.

Imperva’s 2026 reports carry 2025 data. Published April 2026, measurement period calendar 2025: bots 53% (against 51% in 2024), bad bots 40%, benign automation 13%. Presenting this as a 2026 measurement is incorrect.

Some figures are gated or lack methodology. Kasada’s 2026 report is behind a form, and its public claims state neither period nor population. Barracuda has published nothing newer than April 2025, and there the sample is three web applications.

The mark of a reliable figure here is simple: it names the measurement period, the size and composition of the population, and the unit of count. If any of the three is missing, the figure is better left where you found it.

Frequently Asked Questions

Is bot traffic bad?

No. Googlebot, monitoring, RSS feeds — legitimate. The problem is the 32 % malicious share.

How do you tell them apart?

User-Agent + reverse DNS + behavioral fingerprinting (JA3/JA4). UA alone isn't enough (spoofable).

Should I block AI scrapers?

Depends: publishers — block (protect monetization). SaaS docs — allow (AI sends traffic back via links).

Try the live tool that powered this guide

Free plan — 10 monitors, checks every 5 min, no card required. Upgrade for 1-minute interval and multi-region monitoring.