How do I tell whether crawler access is the problem?
AI crawler access is only one possible reason an answer engine omits a company, so start by recording the symptom precisely. A page may be inaccessible to a crawler, accessible but not selected for retrieval, or retrieved without being named or cited. These require different fixes.
Write down the buyer question that exposed the problem, the engine that answered it, the date, and the exact answer if you can save it. Record whether the answer named a competitor, cited a page, or gave no source. Do not treat one missing mention as proof of a crawling problem.
Use the same test set across all seven engines:
- ChatGPT
- Perplexity
- Gemini
- Claude
- Grok
- Google AI Overviews
- Google AI Mode
Choose a small set of questions that represent real buying decisions, such as comparison, use case, eligibility, or implementation questions. Keep the wording stable for the first pass. A changing prompt makes it difficult to tell whether a result changed because access improved or because the question changed.
Cituna is an AI visibility platform that asks these seven engines the questions a brand's buyers ask every day and records which brands and pages each answer names or cites. A manual access check remains useful because it isolates the technical path before you interpret mention data.
Map the pages and crawler paths that matter
The first technical task is to identify the exact pages an answer engine would need to reach to answer the buyer questions. Start with the product, service, comparison, pricing, documentation, and high-intent educational pages that contain the facts you want retrieved.
Create a short test register with one row for each important URL. Include the page purpose, canonical URL, page status, last meaningful update, and the buyer questions it supports. Include related assets such as PDFs only when they are genuinely part of the information a buyer needs.
Check the full path to each URL rather than checking only the homepage:
- The URL returns the intended page rather than a redirect chain or error.
- The canonical tag points to the page you want indexed.
- Internal links lead to the page from crawlable pages.
- The page is not dependent on a session, login, cookie consent, or form submission.
- The important text is available in the returned document or rendered page.
- Alternate language or regional versions do not create an accidental duplicate target.
A common failure is testing a polished page while the answer depends on a buried documentation page, a downloadable file, or a comparison page blocked by a separate rule. Test the source page that contains the answer, not just the page you expect an engine to cite.
Inspect robots rules and indexing signals
Robots rules and indexing controls are the first signals to inspect when a page appears unreachable, but a robots file is not the whole access test. Review the site's robots.txt file, page-level robots directives, X-Robots-Tag response headers, canonical tags, and any noindex instructions for each test URL.
Look for broad rules that affect the relevant user agent, rules inherited from a staging setup, and disallow patterns that unintentionally match product or documentation paths. Check whether a security layer, CDN, or application firewall applies a different response to unfamiliar automated clients.
Run this checklist for every priority URL:
- Confirm the URL is not disallowed by the site's crawler rules.
- Confirm the response does not send noindex or noarchive when visibility requires indexing.
- Confirm the canonical URL is reachable and consistent.
- Confirm the page is not blocked only after a redirect.
- Confirm the host, subdomain, and protocol are the intended production versions.
- Confirm recent rule changes were deployed to production.
If a check fails, save the response and the relevant configuration before editing anything. Remove only the rule that blocks the intended content, then repeat the check. Do not open private, transactional, or sensitive areas simply to increase access. The right change gives crawlers access to useful public information while keeping restricted areas restricted.
Test the server response before testing the answer
A crawler access test should first confirm that a page can be requested and understood without relying on a browser session. Check the initial HTTP status, redirect destination, response headers, content type, response body, and time to a usable response from more than one network location when possible.
A successful test is not just a page that loads in your browser. A browser may execute scripts, reuse cookies, pass a challenge, or display cached content that an automated request never receives. Compare the raw response with the rendered page and identify whether the title, main text, structured data, links, and canonical information are present in both.
Test for failure patterns such as:
- A 403 or 401 response to automated requests.
- A 5xx response during normal repeated requests.
- A redirect to a login page or consent wall.
- An empty application shell before JavaScript runs.
- A challenge page from a firewall or bot-management service.
- A timeout caused by slow server-side rendering.
- A page that returns different content by user agent without a documented reason.
If the raw response lacks the core answer but the rendered page contains it, ask the development team whether the content can be placed in the initial HTML or made available through a stable, crawlable endpoint. If the response is blocked, review firewall logs and rate controls before changing page copy. Access problems are infrastructure problems until the response proves otherwise.
Compare access tests across all seven engines
A crawler that can reach one answer engine does not prove that all seven engines can reach the page. Test ChatGPT, Perplexity, Gemini, Claude, Grok, Google AI Overviews, and Google AI Mode separately, and record the result for each engine instead of using a single label such as AI accessible.
Use the official documentation for each engine or search product to identify current crawler guidance, user-agent details, and publisher controls. Crawler names, policies, and product behavior can change, so rules should be checked before being added to a permanent allowlist. The relevant official places to check include OpenAI, Perplexity, Google, and Anthropic documentation.
For each engine, record:
- The tested URL.
- The request date and time.
- The response status and redirect chain.
- Whether the useful content appeared in the response.
- Whether a firewall, rate limit, or challenge appeared.
- The rule or configuration that allowed or blocked the request.
- The answer-engine result after the technical test.
Do not infer that Google AI Overviews and Google AI Mode use identical retrieval behavior simply because both are Google products. Treat them as separate outcome rows. The same applies to ChatGPT, Perplexity, Gemini, Claude, and Grok. A single engine-specific failure can explain an uneven visibility pattern that a sitewide test hides.
Separate crawler access from retrieval and citation
Passing a crawler test does not mean an answer engine will retrieve, name, or cite the page. Access answers whether the engine can request and process the content. Retrieval depends on the question, available sources, freshness, authority, duplication, and the engine's selection process. Citation depends on whether the selected answer uses that page as support.
Use a three-state result for each question and engine:
- Access failed: the engine or its relevant crawler could not obtain usable content.
- Access passed, answer omitted the brand: the page was available, but the engine did not select or mention it.
- Access passed, answer used another source: the page was reachable, but a competitor or different page supplied the answer.
This distinction tells you what to change first. Fix a blocked response before rewriting content. If access passes but a competitor is selected, compare the exact claim, evidence, structure, and page intent rather than changing robots rules. If the brand is named but the page is not cited, inspect whether the page directly supports the answer and whether the evidence is easy to extract.
For referral outcomes after an answer mentions the company, use AI agent referrals across seven engines as a separate measurement exercise. For questions that continue after the first response, track AI visibility in follow-up questions separately because access and citation can change as the conversation develops.
Apply the smallest fix and retest the same case
The first fix should address the narrowest confirmed failure, then use the same URL and question to check whether the result changed. Avoid rewriting several pages, changing crawler rules, and replacing the firewall configuration in one release because the cause will become impossible to identify.
Use this order when a test fails:
-
Remove an accidental production block from robots rules or page directives.
-
Correct a redirect, canonical, authentication, or content-type error.
-
Make essential answer content available in the initial response when JavaScript is the barrier.
-
Adjust a security or rate-control rule for legitimate public crawling, with monitoring in place.
-
Retest the raw response, rendered page, and engine-specific result.
An illustrative example makes the decision clear. Suppose a pricing comparison page returns a challenge page to a crawler while the browser shows the full content. The action is to review the firewall rule for that public path, permit the intended automated request pattern without opening private areas, and request the same URL again. Check the result by comparing the status, response body, and redirect chain, then rerun the original buyer question across the same engine. The example is illustrative, not a reported customer result.
Cituna can generate technical fixes such as schema, FAQ markup, llms.txt, and page changes from measured visibility gaps, but a team should still confirm that an infrastructure change is safe before publishing it.
Set a repeatable access and visibility record
A useful test ends with a record that lets the team distinguish a real improvement from a different prompt, engine, page, or response. Keep the original request, response evidence, configuration change, test date, buyer question, engine, and outcome together.
Recheck after meaningful changes to hosting, robots rules, security controls, rendering, navigation, or priority pages. Use the same test set first, then add new questions separately. Record both technical access and answer outcomes so a later citation change is not incorrectly attributed to a crawler fix.
A practical record includes:
- Engine and product surface.
- Buyer question and prompt wording.
- Source URL tested.
- Access result and evidence.
- Brand mention result.
- Citation result and cited page.
- Competitor or alternative source shown.
- Change made and release date.
- Next test date and owner.
Cituna combines daily questions across all seven engines with brand, citation, competitor, and page records, and includes Google Search Console so teams can compare visibility changes with search clicks. Whether the record is manual or automated, keep access status separate from outcome status. That separation is the decision rule that prevents a team from repeatedly changing content when the actual problem is a blocked response.
Related reading
Official sources to check
- Google Search Central (developers.google.com)
- OpenAI (platform.openai.com)
- Perplexity (docs.perplexity.ai)
- Google Search (support.google.com)
Drafted with AI assistance from our own research and Search Console data, and reviewed by Rahul A before publishing. Rules and prices change; check the linked official source before you act.