Every machine that
named itself, and
what it came for.
14,923 automated requests to this site between 18 August 2026 and 1 September 2026. 83.0% of them were collecting training data. 133 were a page being fetched because a person had just asked a question and was waiting for it.
Only agents that declare themselves in a user-agent string are counted. Anything arriving dressed as a browser is invisible to this instrument and absent from every number below. This is one small site over a fortnight, not an industry — and nothing here identifies a human reader, because nothing about one is collected.
In short
Automated requests over 15 days, from 20 agents run by 15 operators
Collecting training data, which returns no citation, no link and no reader
Requests — 0.9% — made because a person had just asked a question
A proxy in front of every document request reads the user-agent header, matches it against 31 declared clients, records the agent, its operator, the purpose and the address, and passes the request on. That is the whole instrument. The summary it produces is served as data at /crawler-census.json.
- Counted
- Requests for documents from clients that publish a name for themselves and use it. One crawl of one page is one row.
- Not counted
- Anything presenting a browser user-agent string — whether that is a person or an agent choosing not to say so. There is no estimate of it here, because we have no way to make one worth reading.
- Not collected
- No address, no cookie, no identifier of a human reader. The table has no column for one, which is a stronger statement than a policy is.
- The window
- 18 August 2026 to 1 September 2026 — 15 days, of which 15 carry traffic. Days are UTC days: the site is served from the Gulf and read from everywhere, so the server's clock is the only boundary a reader elsewhere can reproduce.
- Excluded
- Our own probes. Testing this instrument means sending it a spoofed user-agent, and a census that counts its own calibration as a finding is the failure this page exists to avoid.
- Recomputed
- Hourly, in the database rather than here. This copy was computed at 2026-09-01T14:49:25Z.
This is the finding, and it is not the one the field talks about. Training crawls took 12,380 of 14,923 requests, and they return nothing: no citation, no link, no reader. The category that contains a person waiting for an answer is 0.9% of the total.
Corpus collection. No citation, no link, and nobody at the other end. 6 agents, 6 operators, 323 addresses, active on 15 of 15 days.
Building an index. A page missing from these is missing from the answers assembled out of them later. 8 agents, 8 operators, 317 addresses, active on 15 of 15 days.
Third-party crawlers building link and keyword databases that are sold to somebody else. 2 agents, 2 operators, 317 addresses, active on 12 of 15 days.
One address fetched because somebody is waiting for the answer. The only category with a human in it. 3 agents, 3 operators, 21 addresses, active on 15 of 15 days.
Preservation crawls — the Internet Archive and its kin. Not once in this window.
Requests that named a known agent and then asked for something no agent asks for. Counted apart, never as the name they gave. 1 agent, 1 operator, 6 addresses, active on 1 of 15 days.
The zero is worth stopping on. Nothing on this host is blocked from any crawler — the robots.txt says so agent by agent — and not one archive crawler arrived in 15 days. A page nobody archives has exactly one copy, and its publisher owns it.
The smallest category is the interesting one, because every row in it is a question somebody asked a minute earlier. 133 requests landed on 21 addresses. One explainer accounts for 26.3% of them by itself.
/ · first 18 August, last 1 September
/blog/uae-civil-transactions-law-2026-what-it-changes-for-websites-and-service-contracts · first 18 August, last 31 August
first 19 August, last 31 August
/blog/cognitive-prosthetics-architecture-the-evolution-of-proxy-apps-and-the-future-of-digital-interaction · first 18 August, last 31 August
first 21 August, last 24 August
/glossary/asset-register · first 18 August, last 23 August
The remaining 15 addresses were fetched 15 times between them, none of them more than once.
We sell this work, so the reading has to be stated narrowly — starting with the row at the top, which is the front page. It took 57 of the 133 fetches, 42.9% of the category, and it is where a question about who we are lands whatever the answer turns out to be. What the rest supports is that on this site, in this window, the addresses fetched behind it were the ones that explain a subject in checkable terms — a legal explainer, a compliance checklist, a glossary entry. Every page under /services and /products took 5 requests between them. It supports nothing about how many people saw the answer, and nothing about anybody else’s site.
ChatGPT-User, OpenAI’s user-triggered fetcher, made 130 of the 133 requests with somebody waiting behind them. The same operator’s training crawler, GPTBot, did not arrive once.
Training crawls. 252 addresses, 12 of 15 days, first 21 August. Busiest day 24 August: 5,139.
Search and answer-engine indexers. 124 addresses, 15 of 15 days, first 18 August. Busiest day 20 August: 97.
Search and answer-engine indexers. 312 addresses, 13 of 15 days, first 18 August. Busiest day 21 August: 290.
Training crawls. 280 addresses, 13 of 15 days, first 18 August. Busiest day 27 August: 94.
Training crawls. 232 addresses, 15 of 15 days, first 18 August. Busiest day 19 August: 68.
SEO tooling. 317 addresses, 10 of 15 days, first 18 August. Busiest day 20 August: 218.
Search and answer-engine indexers. 106 addresses, 6 of 15 days, first 18 August. Busiest day 18 August: 220.
Training crawls. 62 addresses, 15 of 15 days, first 18 August. Busiest day 18 August: 65.
Search and answer-engine indexers. 59 addresses, 15 of 15 days, first 18 August. Busiest day 27 August: 54.
Search and answer-engine indexers. 42 addresses, 14 of 15 days, first 18 August. Busiest day 18 August: 28.
A person asked, a moment ago. 19 addresses, 15 of 15 days, first 18 August. Busiest day 25 August: 14.
Search and answer-engine indexers. 58 addresses, 15 of 15 days, first 18 August. Busiest day 24 August: 16.
SEO tooling. 36 addresses, 4 of 15 days, first 22 August. Busiest day 1 September: 25.
Training crawls. 1 address, 9 of 15 days, first 19 August. Busiest day 20 August: 2.
Declared identity contradicted by the request. 6 addresses, 1 of 15 days, first 21 August. Busiest day 21 August: 6.
A person asked, a moment ago. 2 addresses, 2 of 15 days, first 24 August. Busiest day 24 August: 1.
Search and answer-engine indexers. 2 addresses, 2 of 15 days, first 20 August. Busiest day 20 August: 1.
Training crawls. 1 address, 1 of 15 days, first 18 August. Busiest day 18 August: 2.
A person asked, a moment ago. 1 address, 1 of 15 days, first 20 August. Busiest day 20 August: 1.
Search and answer-engine indexers. 1 address, 1 of 15 days, first 18 August. Busiest day 18 August: 1.
11 of the 31 agents this instrument knows did not appear at all.
Nothing here is blocked, so an absence is a decision taken at the other end — or a site this small never reaching their queue. Both readings are available, the log cannot separate them, and neither can we.
- MistralAI-User
- Mistral — user-triggered fetcher.
- Meta-ExternalFetcher
- Meta — user-triggered fetcher.
- Claude-SearchBot
- Anthropic — search or answer-engine indexer.
- Applebot-Extended
- Apple — training crawler.
- GPTBot
- OpenAI — training crawler.
- Google-CloudVertexBot
- Google — training crawler.
- CCBot
- Common Crawl — training crawler.
- cohere-ai
- Cohere — training crawler.
- ia_archiver
- Internet Archive — archive crawler.
- archive.org_bot
- Internet Archive — archive crawler.
- DataForSeoBot
- DataForSEO — SEO tool.
The tallest column is 24 August 2026: 5,326 requests, of which 5,155 were training crawls. 5,139 of them came from Meta-ExternalAgent alone, on the busiest day it has had here. That is one crawler arriving, not a level of load. Averaging it across the window and calling the result traffic would be the easiest lie on this page to tell.
Training crawlsEverything else
One column per UTC day, 18 August to 1 September. The figures under the columns are days of the month. The last column is the UTC day this copy was computed on and was still running at 2026-09-01T14:49:25Z — a short column there is an unfinished day, not a fall.
Where a total is really one afternoon.
Meta · training crawler
11,258 requests across 12 days, of which 5,139 — 45.6% — arrived on 24 August 2026.
Perplexity · search or answer-engine indexer
521 requests across 13 days, of which 290 — 55.7% — arrived on 21 August 2026.
Ahrefs · SEO tool
335 requests across 10 days, of which 218 — 65.1% — arrived on 20 August 2026.
Yandex · search or answer-engine indexer
248 requests across 6 days, of which 220 — 88.7% — arrived on 18 August 2026.
This is the cheapest question in the field to answer from a log and the one most often answered from opinion. robots.txt was fetched 482 times by 13 agents. llms.txt was fetched 1 time by 1 agent.
13 agents, most recently 1 September.
1 agent, most recently 1 September.
1 agent, most recently 1 September.
1 agent, most recently 28 August.
We publish an llms.txt anyway, and the file that generates it says why in the same words: no assistant vendor has documented fetching it, so its retrieval value today is close to zero, and it ships as a cheap option on a convention that may yet be adopted. The census is what keeps that sentence honest instead of hopeful.
Markdown representations of these pages were fetched 0 times by 0 agents. Every page that has one now advertises it in a Link: rel="alternate" header, and that was added recently — so the zero is a starting line rather than a verdict. If it is still zero in a month, the honest move is to delete the mechanism rather than to explain it.
6 requests in this window named a known agent and then asked for things no agent asks for. They were filed as somebody waiting for an answer until we looked at what they had requested.
- Perplexity-User
- asked for /.aws/config on 21 August
- Perplexity-User
- asked for /.aws/credentials on 21 August
- Perplexity-User
- asked for /.git-credentials on 21 August
- Perplexity-User
- asked for /.git/config on 21 August
- Perplexity-User
- asked for /.git/HEAD on 21 August
- Perplexity-User
- asked for /wp-json on 21 August
The claimed name is kept verbatim, because the pair is the finding — and the honest reading is that the name is unverifiable, not that the company named did this. They are counted in a row of their own and never inside any of the categories above; they stay inside the window total, which is what a total is for. After the correction, not one request from that operator’s user-triggered fetcher survives in this window: those were all of them.
The rule that catches them is deliberately narrow. A requested address can only downgrade an agent that already declared itself; it never enrols a new one. A scanner arriving as a browser is still not recorded at all, and widening this into “log anything that looks like a probe” would turn an instrument that watches declared crawlers into one that watches everybody.
We sell search and AI visibility, so this page is an argument we benefit from. The limits belong beside it rather than under it.
- It supports
- That most of what is called "AI traffic" is not an audience: 83.0% of it here returns no citation, no link and no reader. That behind the front page, the fetches with a person waiting landed on explanatory pages rather than on the pages that sell — every /services and /products address together took 5 of 133. And that publishing a file is not the same as being read — 482 fetches against 1.
- It does not support
- Any claim about how many people saw an answer, clicked, or bought anything. A fetch is a fetch. Nothing here connects to a conversion, and a supplier selling you a number that does should be asked to show this layer first.
- It is one site
- 15 days, 14,923 requests, one small B2B site in the UAE. A different site in a different category would produce different proportions, and a fortnight is a fortnight.
- It misses the quiet ones
- Every figure counts only clients that declare themselves. An agent presenting a browser user-agent appears in no number on this page, and the honest position is that we do not know how many of those there are.
A proxy, one table and one aggregate query. We built it for our own site before selling it to anybody, which is the order we prefer. If the question is what to publish so that an assistant has something worth citing, that is the work on the search and GEO page.
The same figures as data: /crawler-census.json