Between December 2025 and August 2026 the human share of HTML page requests fell from 47% to 38.1%. This is one of the few comparisons in the field built on the same metric from one source rather than stitched from different reports.
Below: that trend with its limits, six vendors offering five different "leading" AI crawlers, and the metric that explains the tension around the subject — how many pages each platform takes per visitor it sends back.
Most "trends" in this area are stitched together from different vendors with different methodologies — which makes them not trends at all. But one comparison is sound: Cloudflare published the same metric twice — the breakdown of HTML page requests.
| Who requests the page | 2 Dec 2025 | 20–26 Aug 2026 | Change |
|---|---|---|---|
| Human | 47% | 38.1% | −8.9 pp |
| Non-AI bots | 44% | 52.8% | +8.8 pp |
| AI bots | 4.2% | 5.7% | +1.5 pp |
In nine months humans stopped being the majority of those who open pages. The gain went mostly to ordinary bots, not AI ones: the AI-crawler share rose by just a point and a half.
Boundaries: this is Cloudflare’s network, not the internet, and HTML requests only. Counted across all HTTP requests (images and scripts included), bots come to 35.5% for the same week — see our breakdown of denominators.
There is no single answer to "which AI crawler is the most active". Measurements over overlapping periods:
| Vendor | Period | Leader |
|---|---|---|
| Fastly | January 2026 | Google Other — 36% |
| Akamai | Jul–Dec 2025 | OpenAI — 37% |
| HUMAN Security | calendar 2025 | OpenAI — about 69% |
| Radware | 2025 holiday peak | OpenAI — 65% |
| DataDome | February 2026 | Meta-ExternalAgent — about 25% |
| DataDome | Q2 2026 | Meta, both agents — a majority |
These cannot be reduced to one number, and the reasons are clear: different customer verticals, different classification schemes (Fastly deliberately excludes ordinary search crawlers from its "AI" class; Akamai aggregates by vendor-owner rather than User-Agent string), and different units — rule triggers, requests, or fetches per site.
A separate caution about vendors’ own trends. Fastly’s Q3 2025 edition reported bots at 29%; its January 2026 edition reports 49%. A twenty-point jump between consecutive editions of one report speaks more of a classification change than a real surge, and the source offers no explanation.
Cloudflare classifies declared crawl purpose. For 20–26 August 2026: training 39%, training and search combined 37.4%, search 17%, user action 4.7%, undeclared 1.9%.
The practical meaning is clearer from another metric — how many pages a crawler takes per visitor it sends back (same week):
| Platform | Pages per visitor |
|---|---|
| Perplexity | 816.8 |
| Anthropic | 707 |
| OpenAI | 498.1 |
| Microsoft | 35.7 |
| Yandex | 25.5 |
| 5.6 | |
| DuckDuckGo | 1.7 |
The distance between 5.6 and 816.8 is essentially the distance between a search engine and a training crawler: the first takes a page in order to send a reader after it, the second so that no reader is needed.
Caveat: referral attribution depends on the Referer header, which AI clients send inconsistently. The ratios hold as orders of magnitude, not to the decimal.
Per TollBit’s H1 2026 data — the logs of 3,906 publishers — about 15% of AI crawler fetches went past a disallow directive. For individual agents the share is higher: ChatGPT-User at 54% of its own fetches, Bytespider 48%, PerplexityBot 42%.
An important caveat from the source itself: only self-identifying agents are counted, so the real volume exceeds the measured one.
From which follows a practical conclusion worth stating plainly: robots.txt is a request, not an enforcement mechanism. If content genuinely must be closed, that happens at the server — by User-Agent and address ranges, not by a line in a text file. The file is still needed: it records your position, and most crawlers honour it.
To check that your file parses the way you intended, use the robots.txt checker.
Vercel’s AI-crawler data is dated December 2024. Its figures (GPTBot 569M, Claude 370M, PerplexityBot 24.4M against Googlebot 4.5B) are the most-reprinted in this space, they are roughly twenty months old, and Vercel has published no newer study.
Imperva’s 2026 reports carry 2025 data. Published April 2026, measurement period calendar 2025: bots 53% (against 51% in 2024), bad bots 40%, benign automation 13%. Presenting this as a 2026 measurement is incorrect.
Some figures are gated or lack methodology. Kasada’s 2026 report is behind a form, and its public claims state neither period nor population. Barracuda has published nothing newer than April 2025, and there the sample is three web applications.
The mark of a reliable figure here is simple: it names the measurement period, the size and composition of the population, and the unit of count. If any of the three is missing, the figure is better left where you found it.
No. Googlebot, monitoring, RSS feeds — legitimate. The problem is the 32 % malicious share.
User-Agent + reverse DNS + behavioral fingerprinting (JA3/JA4). UA alone isn't enough (spoofable).
Depends: publishers — block (protect monetization). SaaS docs — allow (AI sends traffic back via links).
Free plan — 10 monitors, checks every 5 min, no card required. Upgrade for 1-minute interval and multi-region monitoring.