Skip to main content
AI Visibility10 min read

How to See Which AI Bots Crawl Your Site in Server Logs

Find GPTBot, ClaudeBot, PerplexityBot and the other AI crawlers in your server logs, check each one is genuine against its published IP range, and read what every status code means.

Published

Your server’s access log is the only place that shows which AI bots actually fetched your pages. Each vendor’s crawler announces itself with a published token in the user-agent string: GPTBot, OAI-SearchBot and ChatGPT-User for OpenAI, ClaudeBot, Claude-SearchBot and Claude-User for Anthropic, PerplexityBot and Perplexity-User for Perplexity, Googlebot for Google’s AI Overviews and AI Mode, and meta-externalagent for Meta. Search the log for those tokens, confirm the requesting IP sits in the vendor’s published range, and read the status code on each request. A 200 means the bot got the page. A 403 or 429 means something in your stack turned it away, whatever robots.txt says.

Before you open a log, check what your rules ask for. Cituna’s free checker tests eight AI user agents against your robots.txt, names the line that decides each one, and shows how much of your homepage a crawler can read without JavaScript. The log then tells you what the bots really did.

The AI user-agent tokens, by vendor

Every major AI vendor now runs more than one bot, and they do different jobs. The difference matters, because blocking the training crawler and blocking the search crawler have opposite effects on whether you appear in answers. The table below is taken from each vendor’s own documentation (OpenAI, Anthropic, Perplexity, Google, Meta), read on October 6, 2026.

AI crawler tokens you can find in a server log
VendorToken in the logWhat it doesrobots.txtPublished IP list
OpenAIGPTBotCollects pages for training OpenAI’s foundation modelsRespectedopenai.com/gptbot.json
OpenAIOAI-SearchBotFinds pages to show and link in ChatGPT searchRespectedopenai.com/searchbot.json
OpenAIChatGPT-UserFetches a page when a ChatGPT user or a custom GPT asks for itMay not apply, per OpenAIopenai.com/chatgpt-user.json
AnthropicClaudeBotCollects pages that may contribute to training Anthropic’s modelsRespected, including Crawl-delayclaude.com/crawling/bots.json (one list for Anthropic)
AnthropicClaude-SearchBotAnalyses pages to improve Claude’s search resultsRespectedclaude.com/crawling/bots.json
AnthropicClaude-UserFetches a page when a Claude user asks a questionRespected, per Anthropicclaude.com/crawling/bots.json
PerplexityPerplexityBotFinds pages to show and link in Perplexity answers; Perplexity says it is not used for trainingRespectedperplexity.com/perplexitybot.json
PerplexityPerplexity-UserVisits a page during a user’s questionGenerally ignored, per Perplexityperplexity.com/perplexity-user.json
GoogleGooglebotCrawls for Google Search, which AI Overviews and AI Mode draw onRespectedgooglebot.json on developers.google.com
Metameta-externalagentCrawls for training foundation AI models and indexing content for Meta’s productsRespectedNone on Meta’s crawler page
Metameta-externalfetcherFetches individual links at a user’s request, including agent tasksMay bypass, per MetaNone on Meta’s crawler page

Read on each vendor’s documentation on October 6, 2026. Every IP file listed returned a JSON list of prefixes that day. Anthropic’s help page says a crawler whose source IP is on its list is coming from Anthropic.

Two entries in that table catch people out.

Google-Extended will never appear in your log. Google’s crawler documentation says it is a robots.txt token only and has no request user agent of its own: the fetching is done by Google’s usual crawlers. So a log with no “Google-Extended” in it tells you nothing. Google’s AI Overviews and AI Mode use pages Googlebot crawled for Search, so for those surfaces, Googlebot is the bot to watch.

The user-triggered fetchers are the ones that look like your audience. ChatGPT-User, Claude-User, Perplexity-User and meta-externalfetcher arrive because a person asked an assistant something and the assistant opened your page to answer. A rising count on these is closer to demand than a rising count on GPTBot or ClaudeBot, which crawl on their own schedule. It is still not a visit: the person may never click through. Human clicks from those answers show up as referrals in analytics, which is a separate measurement covered in measuring AI agent referrals across seven engines.

OpenAI also documents OAI-AdsBot, which only checks landing pages submitted as ChatGPT ads. Unless you advertise there, you will not see it.

Step 1: find the log that saw the request

Bots hit the first server in your path. If a CDN or a web application firewall sits in front of your origin, a request it blocked or served from cache never reaches your origin log. That is the most common reason a site owner says “I see no GPTBot at all” while the bot is in fact being turned away at the edge.

  • Plain server (nginx, Apache, Caddy): the access log on the origin is complete. On many Linux installs it lives at /var/log/nginx/access.log or /var/log/apache2/access.log, rotated daily with older days compressed.
  • Behind a CDN or WAF: read the edge’s request logs or its firewall events as well, filtered on user agent. The origin log only shows what got through.
  • Managed hosting or serverless: export the request logs from the host’s dashboard or log drain. Make sure the export includes the user agent, path, status code and client IP, because some default views drop the user agent.

Pull at least seven days. Most AI crawlers do not visit every day, and a single day can look empty by chance.

About the commands below

They are examples for a combined-format log, the default for nginx and Apache. We ran each one on a short sample log in that format, not on our own traffic, so adapt the field numbers and file names to your server before you trust the output.

Step 2: count requests per AI bot

On a combined-format log, one command gives a count per token:

grep -oE '(GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User|meta-externalagent|meta-externalfetcher|Googlebot)/' access.log \
  | sort | uniq -c | sort -rn

Matching the token with its trailing slash, and keeping the match case-sensitive, stops the count from picking up the lowercase links inside the user-agent strings themselves (OpenAI’s strings end in a URL such as openai.com/gptbot). For rotated files, run the same pipe over zcat -f access.log* instead of a single file.

The result is your first answer: which AI bots came, and roughly how hard each one crawled. If a bot you allowed is missing after a week, go back to step 1 and check the edge (CDN or WAF logs) before you assume the vendor is ignoring you.

Step 3: see what each bot asked for and what it got

The count says who came. The path and status code say whether the visit did you any good. In combined format the path is field 7 and the status is field 9:

# status codes returned to one bot
awk '/OAI-SearchBot\// {print $9}' access.log | sort | uniq -c | sort -rn

# the pages that bot requested most
awk '/OAI-SearchBot\// {print $7}' access.log | sort | uniq -c | sort -rn | head -20

Swap the token to repeat it per bot. Read the status codes like this:

What each status code means for an AI crawler
StatusWhat happenedWhat to do
200The bot received the pageNothing. This is what you want on every page you expect to be cited
304The bot asked whether the page had changed and was told it had notNothing. Healthy
301 or 308A redirectOne hop is fine. If a bot keeps requesting old URLs that redirect, your internal links or sitemap still point at them
403Refusedrobots.txt does not produce a 403; a WAF rule, a bot-management setting or a server rule does. If you allowed the bot in robots.txt, your rules and your infrastructure disagree
429Rate-limitedThe bot got some pages and was told to slow down on others. Check which paths were cut off
404The page does not existFind where the bot got the URL: an old sitemap, a broken internal link, or a link elsewhere on the web
5xxYour server failedFix the error. A bot that keeps getting errors may come back less often

Then read the paths. A training crawler spending its visits on tag archives, filtered URLs and pagination, while your pricing and product pages get nothing, is a crawl-budget problem you can fix with internal links and a cleaner sitemap. The page on testing AI crawler access across seven engines covers the server-side failure patterns in more depth.

Step 4: confirm the bot is genuine

A user agent is a line of text, and anything can send it. Scrapers routinely borrow GPTBot’s or Googlebot’s name. Before a number goes into a report, check the IP address.

OpenAI, Anthropic, Perplexity and Google publish the IP ranges their bots use, as JSON files; Meta’s crawler page lists none. OpenAI and Perplexity publish one file per bot. Anthropic publishes a single list and says a crawler whose source IP is on it is coming from Anthropic. Google publishes googlebot.json and also supports a reverse DNS check: a genuine Googlebot IP resolves to a host ending in googlebot.com or google.com, and that host resolves back to the same IP. A short script does the IP check:

curl -s https://openai.com/gptbot.json | python3 -c '
import json, sys, ipaddress
nets = [ipaddress.ip_network(p.get("ipv4Prefix") or p.get("ipv6Prefix"))
        for p in json.load(sys.stdin)["prefixes"]]
for line in open("access.log"):
    if "GPTBot/" in line:
        ip = line.split()[0]
        ok = any(ipaddress.ip_address(ip) in n for n in nets)
        print(ip, "genuine" if ok else "NOT in published range")
' | sort | uniq -c

Swap the JSON file and the token for each bot. The files change, so fetch them fresh each time rather than saving a copy. If your server sits behind a CDN, the first field in the log may be the CDN’s address rather than the bot’s; log the real client IP from the forwarded header before you run this.

Requests that claim an AI bot’s name but come from outside its published range are not that vendor’s crawler. Exclude them from your counts, and treat them as ordinary scraper traffic when you decide what to block.

Step 5: make it a weekly record

One log read answers “which AI bots crawl my site today”. The useful question is whether that changes after you change something: a robots.txt edit, a WAF rule, a site migration, a new section. Keep a small weekly table per bot with four numbers: requests, share returning 200, share returning 403 or 429, and the number of distinct pages fetched. A drop in 200s after a deploy is a regression you can catch in days rather than discovering months later that an engine stopped citing you.

Keep the decision about which bots to allow separate from the measurement. The log tells you what happened; whether GPTBot should be allowed at all is a policy choice, discussed in how to allow AI crawlers in your robots.txt safely.

What a crawl does not tell you

A bot fetching your page is a prerequisite for being cited, not evidence of it. GPTBot can read every page you have and ChatGPT can still recommend a rival when a buyer asks. The log ends where the answer begins: it cannot show which questions the engines answered with your page, or whether they named you.

That is the gap Cituna fills. It asks ChatGPT, Perplexity, Gemini, Claude, Grok, Google AI Overviews and Google AI Mode your buyers’ questions every day, records which brands and pages each answer names and cites, and writes the fix for each answer you are missing from. Read the log for access, and the answers for outcome.

Drafted with AI assistance from our own research and Search Console data, and reviewed by Rahul A before publishing. Rules and prices change; check the linked official source before you act.

Frequently asked questions

Does blocking GPTBot remove my site from ChatGPT?

Not on its own. OpenAI documents GPTBot as its training crawler and OAI-SearchBot as the one that finds pages for ChatGPT search, and controls them separately in robots.txt. Blocking GPTBot while allowing OAI-SearchBot keeps you eligible for ChatGPT search results. ChatGPT-User, which fetches pages a user asks for, is described by OpenAI as one that robots.txt rules may not apply to.

Why do I see ChatGPT-User or Perplexity-User but never GPTBot?

The user-triggered fetchers only arrive when someone asks an assistant about a page or a topic that leads to it, while the crawlers run on their own schedule and may not have reached you yet. If GPTBot is allowed in robots.txt and still absent after a few weeks, check your CDN or firewall logs for 403s on that token before anything else.

Can I see AI bot visits in Google Analytics?

No. GA4 measures visits from browsers that run its JavaScript, and AI crawlers do not run it, so bot fetches never reach your analytics. What GA4 can show is the people who click a link in an AI answer and arrive with a referrer such as chatgpt.com or perplexity.ai. Bot visits live only in server, CDN or WAF logs.

Is a sudden spike from one AI bot a problem?

Usually it is a fresh crawl after a sitemap change or a new section, and it settles. It becomes a problem when it pushes up 5xx errors or 429s for real visitors. Anthropic documents that ClaudeBot respects a Crawl-delay line in robots.txt, which slows it without blocking it. Check first that the traffic comes from the vendor’s published IP range, because a spike under a borrowed name is a scraper, not the vendor.

Turn this checklist into fixes for your site

Use Cituna to review your technical SEO, AEO and GEO issues alongside measured AI answers. Prioritize the fixes and content gaps that need your attention.

Start free trial

3-day free trial · Card required, cancel anytime · Plans from $39 a month

Check crawler readiness free