- Python 100%
The path breakdown was populated only inside the IP-range branch, so it silently covered one of the two verification methods. Vendors confirmed by PTR or ASN - YouBot, Amazonbot, Meta - were counted in their own table but contributed nothing to the only output that shows what crawlers actually fetch. On the sample instance that hid roughly 40% of the verified requests from the path list. Collect paths per address during the log pass and merge them once the address is confirmed, whichever method confirms it. Private-source hits are excluded, since those addresses are never judged either way. Move the section to the end of the report and retitle it: it now spans both tables above, so placing it under the IP-range table implied a narrower scope than it has. Assisted-by: Claude:opus-5 |
||
|---|---|---|
| crawler-alert.env.example | ||
| crawler-alert.py | ||
| crawler-alert.service | ||
| crawler-alert.timer | ||
| crawler-weekly.service | ||
| crawler-weekly.timer | ||
| LICENSE | ||
| README.md | ||
| traefik-accesslog.yml | ||
crawler-alert
Tells you which AI crawlers are really visiting your site, and which are lying about it.
Reads a Traefik access log, mails a threshold alert daily and a summary
weekly. Every bot is checked against its vendor's published IP ranges,
so a forged User-Agent does not get counted as the real thing.
Python 3.9+, standard library only. No database, no daemon, no agent.
Why user-agent alone is not enough
A User-Agent header is a string the client chooses. Anyone can send
ClaudeBot/1.0. On the Forgejo instance this was written for, measured
over 2.27M requests between November 2025 and August 2026:
| Bot | Claimed | Verified | Forged |
|---|---|---|---|
| Claude-User | 5,000 | 0 | 4,992 |
| ChatGPT-User | 2,670 | 13 | 2,657 |
| PerplexityBot | 2,038 | 0 | 846 |
| ClaudeBot | 4,301 | 1,070 | 852 |
| Google-Extended | 879 | 0 | 747 |
| GPTBot | 1,864 | 766 | 1,098 |
About 83% of AI-labelled traffic was not from the vendor it named.
The forgeries shared infrastructure: roughly 100 Google Cloud hosts
rotating through six vendor identities, sometimes within the same hour.
Their requests went to /.env, /terraform.tfstate,
/serviceAccountKey.json, /.env.prod.bak — credential harvesting
wearing a crawler's name.
The genuine crawlers behaved completely differently: 2,682 requests,
zero robots.txt violations, no repository source fetched, ~125 MB
total. Mostly they re-read robots.txt and sitemap.xml.
That gap is the point. Counting by user-agent tells you AI crawlers are hammering your server. Counting by verified IP tells you they are not, and that something else is — which is a different problem with a different fix.
How verification works
Most major vendors publish their crawler IP ranges as JSON. A request counts as legitimate only when its source address falls inside the range list belonging to the vendor its user-agent names.
| Vendor | Endpoint |
|---|---|
| OpenAI | openai.com/{chatgpt-user,gptbot,searchbot}.json |
| Anthropic | claude.com/crawling/bots.json |
developers.google.com/static/crawling/ipranges/*.json |
|
| Perplexity | perplexity.ai/perplexitybot.json |
| Apple | search.developer.apple.com/applebot.json |
Meta, ByteDance, Amazon and You.com publish no range file. Those are checked by forward-confirmed reverse DNS instead: the address's PTR record must end in a vendor domain, and that hostname must resolve back to the same address. The forward step is what makes it evidence — a PTR alone is written by whoever controls the address block. Where even a PTR is absent, a single-tenant ASN match is used as a fallback (see below).
| Bot | rDNS suffix |
|---|---|
| Amazonbot | .crawl.amazonbot.amazon |
| meta-externalagent, FacebookBot | .facebook.com, .fbsv.net |
| YouBot | .search.you.com |
| Bytespider | .bytedance.com, .byteoversea.com |
| PetalBot | .petalsearch.com, .aspiegel.com |
| DuckAssistBot | .duckduckgo.com |
Diffbot, Cohere and Common Crawl publish neither ranges nor a documented rDNS convention; they appear in the user-agent tally only.
ASN, and why it is used for exactly one vendor
An ASN match is usually worthless as verification: genuine YouBot
(68.67.112.227) and genuine Amazonbot (100.24.167.60) both live in
AS14618, Amazon — the same ASN as every EC2 instance on the
internet, including any spoofer renting one. Matching there would prove
the traffic came from a cloud, not from the vendor.
It is only evidence when the ASN is single-tenant. Meta is the one
case that qualifies here: AS32934 is Meta's own network and sells no
compute, Meta publishes no IP range file (their documentation says to
email webmasters@meta.com), and their IPv6 crawler space carries no
PTR records — so without ASN there is nothing to check them against at
all. The allowlist is BOT_ASN in the script; adding a cloud ASN to it
would defeat the purpose. Disable with CRAWLER_ASN=0.
The four outcomes
| Column | Meaning |
|---|---|
PTR ok |
Forward-confirmed reverse DNS |
ASN ok |
No PTR, but address is in the vendor's single-tenant ASN |
FORGED |
A PTR exists and points elsewhere — the only column that is evidence of a lie |
no-PTR |
No PTR and no usable ASN — unknown, not accused |
The split matters. An earlier version had one "failed" bucket, which put "lying about who you are" and "your vendor doesn't publish DNS records" in the same number — and reported every genuine Meta request as though it were a forgery.
Ranges are fetched at run time and cached. If a vendor endpoint is unreachable the script falls back to the cache and says so in the mail, so an outage degrades the report instead of breaking it.
Requirements
- Traefik writing an access log with the User-Agent header kept —
it is dropped by default. See
traefik-accesslog.yml; without it nothing here works. - Python 3.9+
- An SMTP account, or a Forgejo/Gitea container whose
app.ini[mailer]section the script can borrow credentials from.
Install
sudo install -m755 crawler-alert.py /usr/local/bin/
sudo cp crawler-alert.env.example /etc/crawler-alert.env
sudo editor /etc/crawler-alert.env # set CRAWLER_RECIPIENT
sudo cp crawler-*.service crawler-*.timer /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now crawler-alert.timer crawler-weekly.timer
Check it against your own log before enabling the timers:
CRAWLER_LOG='/var/log/traefik/access.log*' crawler-alert.py --weekly --dry-run
--dry-run prints the mail instead of sending it and needs no SMTP
config, so it is safe to run repeatedly while you tune things.
Configuration
All settings are environment variables, read from
/etc/crawler-alert.env by the systemd units. Only CRAWLER_RECIPIENT
has no default.
| Variable | Default | Meaning |
|---|---|---|
CRAWLER_RECIPIENT |
(required) | Where reports are mailed |
CRAWLER_SITE |
this site |
Name shown in subject/header |
CRAWLER_LOG |
/var/log/traefik/access.log* |
Log glob |
CRAWLER_ROUTER |
(empty) | Traefik router to count, e.g. forgejo@docker. Empty counts everything |
CRAWLER_THRESHOLD |
100000 |
Daily alert fires above this |
CRAWLER_CACHE |
/var/cache/crawler-alert/ranges.json |
Cached ranges |
CRAWLER_MAIL_CONTAINER |
forgejo |
Container to read SMTP from |
CRAWLER_MAIL_INI |
/data/gitea/conf/app.ini |
Path inside it |
CRAWLER_SMTP_* |
(unset) | Configure SMTP directly instead |
Set CRAWLER_THRESHOLD near your normal peak. Far above it, the alert
never fires and is not a tripwire — run --dry-run for a few days and
pick a number a busy day would actually reach.
Modes
crawler-alert.py # daily: mail only if over threshold
crawler-alert.py --weekly # 7-day summary, always mails
crawler-alert.py --dry-run # print instead of mailing
Only --weekly fetches ranges; the daily check needs no network.
Sample output
==============================================================
LEGITIMATE CRAWLERS (source IP inside vendor's published ranges)
==============================================================
bot verified spoofed via CF IPs
ChatGPT-User 0 561 0 0
ClaudeBot 222 161 0 14
OAI-SearchBot 23 209 0 14
GPTBot 34 196 0 2
PerplexityBot 0 196 0 0
1909 requests claimed an AI-bot identity; 323 verified (16.9%).
1586 failed IP verification (spoofed). 0 arrived via Cloudflare
and cannot be verified by IP - not counted either way.
--------------------------------------------------------------
VERIFIED BY DNS / ASN (vendors publishing no IP ranges)
--------------------------------------------------------------
bot PTR ok ASN ok FORGED no-PTR CF priv IPs
meta-externalagent 0 1 14 0 0 21636 1
Amazonbot 47 0 1361 6 264 0 45
YouBot 90 0 0 0 0 0 22
Bytespider 0 0 11 0 2 0 0
Reading that: YouBot is entirely genuine. Amazonbot is mostly not —
1,361 requests carry a PTR pointing somewhere other than Amazon.
Meta's single ASN ok is a real request from AS32934; its 14 FORGED
resolve to googleusercontent.com, i.e. GCP hosts wearing Meta's
name. The 21,636 under priv are the IPv6 blind spot described in
Limitations, not a judgement.
Top paths, verified crawlers:
117 /robots.txt
73 /sitemap.xml
13 /explore/repos
Limitations
Cloudflare-proxied requests cannot be verified. If traffic reaches
Traefik through Cloudflare, the logged address is Cloudflare's edge and
not the client. Those hits go in a separate via CF column rather than
being guessed at. traefik-accesslog.yml shows how to recover real
client IPs with forwardedHeaders.trustedIPs.
Some vendor lists are stale. Perplexity's has 8 prefixes and was last updated in February 2025; Apple's dates to 2023. Genuine traffic from those two may fail verification. A 0% verified rate for a small vendor is weaker evidence than it looks.
Private source addresses cannot be judged at all. If the logged
client IP is RFC1918 (172.18.0.1, a Docker bridge gateway, say), the
real address was replaced before it was written. Those hits go in the
priv column, neither verified nor spoofed.
The most common cause is worth calling out, because it is silent and
it is easy to have without knowing. If your host publishes a port on
:: but the Docker network has EnableIPv6=false, every IPv6 visitor
is logged as the bridge gateway. The kernel cannot DNAT v6 into a
v4-only bridge, so docker-proxy accepts the connection in userland
and opens a fresh IPv4 connection to your reverse proxy — the original
address is not forwarded, it is gone.
On the sample instance this hid 21,624 requests, the entire IPv6 population, behind one address. Check with:
docker network inspect <net> --format '{{.EnableIPv6}}'
dig +short AAAA your.site
If that returns false and an AAAA record exists, you have this
problem. Fixing it means recreating the network with --ipv6 and a
subnet from your provider's prefix, which requires restarting every
container attached to it.
rDNS is rate-limited by design. Each unique address costs two DNS
lookups, so only the busiest CRAWLER_RDNS_MAX (default 400) addresses
per bot are resolved; the rest are reported as unresolved. Raise it if
your DNS resolver is fast and local.
IPv6 is verified only where vendors publish v6 prefixes, which most do; addresses outside those are treated as unverified.
This is a reporting tool, not an enforcement one. It tells you what
happened. Blocking is a separate decision, and robots.txt only works
on the bots that already respect it — which, per the numbers above, are
exactly the ones not causing the load.
License
0BSD — see LICENSE. Do what you like with it.