Your server’s access log is the only place that shows which AI bots actually fetched your pages. Each vendor’s crawler announces itself with a published token in the user-agent string: GPTBot, OAI-SearchBot and ChatGPT-User for OpenAI, ClaudeBot, Claude-SearchBot and Claude-User for Anthropic, PerplexityBot and Perplexity-User for Perplexity, Googlebot for Google’s AI Overviews and AI Mode, and meta-externalagent for Meta. Search the log for those tokens, confirm the requesting IP sits in the vendor’s published range, and read the status code on each request. A 200 means the bot got the page. A 403 or 429 means something in your stack turned it away, whatever robots.txt says.
Before you open a log, check what your rules ask for. Cituna’s free checker tests eight AI user agents against your robots.txt, names the line that decides each one, and shows how much of your homepage a crawler can read without JavaScript. The log then tells you what the bots really did.
The AI user-agent tokens, by vendor
Every major AI vendor now runs more than one bot, and they do different jobs. The difference matters, because blocking the training crawler and blocking the search crawler have opposite effects on whether you appear in answers. The table below is taken from each vendor’s own documentation (OpenAI, Anthropic, Perplexity, Google, Meta), read on October 6, 2026.
| Vendor | Token in the log | What it does | robots.txt | Published IP list |
|---|---|---|---|---|
| OpenAI | GPTBot | Collects pages for training OpenAI’s foundation models | Respected | openai.com/gptbot.json |
| OpenAI | OAI-SearchBot | Finds pages to show and link in ChatGPT search | Respected | openai.com/searchbot.json |
| OpenAI | ChatGPT-User | Fetches a page when a ChatGPT user or a custom GPT asks for it | May not apply, per OpenAI | openai.com/chatgpt-user.json |
| Anthropic | ClaudeBot | Collects pages that may contribute to training Anthropic’s models | Respected, including Crawl-delay | claude.com/crawling/bots.json (one list for Anthropic) |
| Anthropic | Claude-SearchBot | Analyses pages to improve Claude’s search results | Respected | claude.com/crawling/bots.json |
| Anthropic | Claude-User | Fetches a page when a Claude user asks a question | Respected, per Anthropic | claude.com/crawling/bots.json |
| Perplexity | PerplexityBot | Finds pages to show and link in Perplexity answers; Perplexity says it is not used for training | Respected | perplexity.com/perplexitybot.json |
| Perplexity | Perplexity-User | Visits a page during a user’s question | Generally ignored, per Perplexity | perplexity.com/perplexity-user.json |
| Googlebot | Crawls for Google Search, which AI Overviews and AI Mode draw on | Respected | googlebot.json on developers.google.com | |
| Meta | meta-externalagent | Crawls for training foundation AI models and indexing content for Meta’s products | Respected | None on Meta’s crawler page |
| Meta | meta-externalfetcher | Fetches individual links at a user’s request, including agent tasks | May bypass, per Meta | None on Meta’s crawler page |
Read on each vendor’s documentation on October 6, 2026. Every IP file listed returned a JSON list of prefixes that day. Anthropic’s help page says a crawler whose source IP is on its list is coming from Anthropic.
Two entries in that table catch people out.
Google-Extended will never appear in your log. Google’s crawler documentation says it is a robots.txt token only and has no request user agent of its own: the fetching is done by Google’s usual crawlers. So a log with no “Google-Extended” in it tells you nothing. Google’s AI Overviews and AI Mode use pages Googlebot crawled for Search, so for those surfaces, Googlebot is the bot to watch.
The user-triggered fetchers are the ones that look like your audience. ChatGPT-User, Claude-User, Perplexity-User and meta-externalfetcher arrive because a person asked an assistant something and the assistant opened your page to answer. A rising count on these is closer to demand than a rising count on GPTBot or ClaudeBot, which crawl on their own schedule. It is still not a visit: the person may never click through. Human clicks from those answers show up as referrals in analytics, which is a separate measurement covered in measuring AI agent referrals across seven engines.
OpenAI also documents OAI-AdsBot, which only checks landing pages submitted as ChatGPT ads. Unless you advertise there, you will not see it.
Step 1: find the log that saw the request
Bots hit the first server in your path. If a CDN or a web application firewall sits in front of your origin, a request it blocked or served from cache never reaches your origin log. That is the most common reason a site owner says “I see no GPTBot at all” while the bot is in fact being turned away at the edge.
- Plain server (nginx, Apache, Caddy): the access log on the origin is complete. On many Linux installs it lives at /var/log/nginx/access.log or /var/log/apache2/access.log, rotated daily with older days compressed.
- Behind a CDN or WAF: read the edge’s request logs or its firewall events as well, filtered on user agent. The origin log only shows what got through.
- Managed hosting or serverless: export the request logs from the host’s dashboard or log drain. Make sure the export includes the user agent, path, status code and client IP, because some default views drop the user agent.
Pull at least seven days. Most AI crawlers do not visit every day, and a single day can look empty by chance.
About the commands below
Step 2: count requests per AI bot
On a combined-format log, one command gives a count per token:
grep -oE '(GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User|meta-externalagent|meta-externalfetcher|Googlebot)/' access.log \
| sort | uniq -c | sort -rnMatching the token with its trailing slash, and keeping the match case-sensitive, stops the count from picking up the lowercase links inside the user-agent strings themselves (OpenAI’s strings end in a URL such as openai.com/gptbot). For rotated files, run the same pipe over zcat -f access.log* instead of a single file.
The result is your first answer: which AI bots came, and roughly how hard each one crawled. If a bot you allowed is missing after a week, go back to step 1 and check the edge (CDN or WAF logs) before you assume the vendor is ignoring you.
Step 3: see what each bot asked for and what it got
The count says who came. The path and status code say whether the visit did you any good. In combined format the path is field 7 and the status is field 9:
# status codes returned to one bot
awk '/OAI-SearchBot\// {print $9}' access.log | sort | uniq -c | sort -rn
# the pages that bot requested most
awk '/OAI-SearchBot\// {print $7}' access.log | sort | uniq -c | sort -rn | head -20Swap the token to repeat it per bot. Read the status codes like this:
| Status | What happened | What to do |
|---|---|---|
| 200 | The bot received the page | Nothing. This is what you want on every page you expect to be cited |
| 304 | The bot asked whether the page had changed and was told it had not | Nothing. Healthy |
| 301 or 308 | A redirect | One hop is fine. If a bot keeps requesting old URLs that redirect, your internal links or sitemap still point at them |
| 403 | Refused | robots.txt does not produce a 403; a WAF rule, a bot-management setting or a server rule does. If you allowed the bot in robots.txt, your rules and your infrastructure disagree |
| 429 | Rate-limited | The bot got some pages and was told to slow down on others. Check which paths were cut off |
| 404 | The page does not exist | Find where the bot got the URL: an old sitemap, a broken internal link, or a link elsewhere on the web |
| 5xx | Your server failed | Fix the error. A bot that keeps getting errors may come back less often |
Then read the paths. A training crawler spending its visits on tag archives, filtered URLs and pagination, while your pricing and product pages get nothing, is a crawl-budget problem you can fix with internal links and a cleaner sitemap. The page on testing AI crawler access across seven engines covers the server-side failure patterns in more depth.
Step 4: confirm the bot is genuine
A user agent is a line of text, and anything can send it. Scrapers routinely borrow GPTBot’s or Googlebot’s name. Before a number goes into a report, check the IP address.
OpenAI, Anthropic, Perplexity and Google publish the IP ranges their bots use, as JSON files; Meta’s crawler page lists none. OpenAI and Perplexity publish one file per bot. Anthropic publishes a single list and says a crawler whose source IP is on it is coming from Anthropic. Google publishes googlebot.json and also supports a reverse DNS check: a genuine Googlebot IP resolves to a host ending in googlebot.com or google.com, and that host resolves back to the same IP. A short script does the IP check:
curl -s https://openai.com/gptbot.json | python3 -c '
import json, sys, ipaddress
nets = [ipaddress.ip_network(p.get("ipv4Prefix") or p.get("ipv6Prefix"))
for p in json.load(sys.stdin)["prefixes"]]
for line in open("access.log"):
if "GPTBot/" in line:
ip = line.split()[0]
ok = any(ipaddress.ip_address(ip) in n for n in nets)
print(ip, "genuine" if ok else "NOT in published range")
' | sort | uniq -cSwap the JSON file and the token for each bot. The files change, so fetch them fresh each time rather than saving a copy. If your server sits behind a CDN, the first field in the log may be the CDN’s address rather than the bot’s; log the real client IP from the forwarded header before you run this.
Requests that claim an AI bot’s name but come from outside its published range are not that vendor’s crawler. Exclude them from your counts, and treat them as ordinary scraper traffic when you decide what to block.
Step 5: make it a weekly record
One log read answers “which AI bots crawl my site today”. The useful question is whether that changes after you change something: a robots.txt edit, a WAF rule, a site migration, a new section. Keep a small weekly table per bot with four numbers: requests, share returning 200, share returning 403 or 429, and the number of distinct pages fetched. A drop in 200s after a deploy is a regression you can catch in days rather than discovering months later that an engine stopped citing you.
Keep the decision about which bots to allow separate from the measurement. The log tells you what happened; whether GPTBot should be allowed at all is a policy choice, discussed in how to allow AI crawlers in your robots.txt safely.
What a crawl does not tell you
A bot fetching your page is a prerequisite for being cited, not evidence of it. GPTBot can read every page you have and ChatGPT can still recommend a rival when a buyer asks. The log ends where the answer begins: it cannot show which questions the engines answered with your page, or whether they named you.
That is the gap Cituna fills. It asks ChatGPT, Perplexity, Gemini, Claude, Grok, Google AI Overviews and Google AI Mode your buyers’ questions every day, records which brands and pages each answer names and cites, and writes the fix for each answer you are missing from. Read the log for access, and the answers for outcome.
Drafted with AI assistance from our own research and Search Console data, and reviewed by Rahul A before publishing. Rules and prices change; check the linked official source before you act.