What should an AI visibility audit measure first?
A useful AI visibility audit starts by defining exactly what one observation means and what decision the data must support. Decide whether you are checking brand mentions, citations, answer position, competitor presence, or a specific page being selected as a source. Do not combine those outcomes into one score before reviewing them separately.
Write the audit question in a form that can be tested, such as: “When buyers ask which accounting tools suit a ten-person consultancy, does our brand appear, and which page supports the answer?” The question gives you a prompt, an audience, an intent, and a result to inspect.
Use a short audit record with these fields:
- Prompt text and intent.
- Engine and model surface.
- Run date and location or language, if relevant.
- Brand mention, citation, and answer position.
- Competitors named instead.
- Pages or domains cited.
- Notes about ambiguity, refusal, or a changed answer.
The distinction between a mention and a citation matters. A response may name a company without linking to a supporting page, or cite a page without giving the brand a prominent position. The page that answers the question and the brand that receives the mention can be different, so record both.
Freeze the prompt set and run conditions
Comparable tracking data requires a stable prompt set and documented run conditions. Copy the exact prompt text into a controlled list, preserve punctuation and qualifiers, and record whether the prompt is asking for a recommendation, comparison, definition, or source-backed answer. Small wording changes can alter which entities and pages an engine retrieves.
Create prompt groups rather than treating every question as equivalent. A category question, a problem-solving question, a product comparison, and a branded question reveal different visibility gaps. Keep the groups separate when you calculate or discuss results.
Record the conditions that could change the answer:
- Logged-in or logged-out state, where applicable.
- Country, language, device, and search surface.
- Model or engine version when exposed.
- Conversation history and system instructions.
- Date, time, and number of attempts.
An audit fails comparability when last month’s prompts were broad and this month’s prompts contain a brand name. If the prompt set must change, create a new version and label the break. Do not describe the difference as a visibility improvement until the old and new sets have a controlled overlap.
Check engine coverage and collection completeness
A complete audit checks whether data was collected from all seven tracked surfaces: ChatGPT, Perplexity, Gemini, Claude, Grok, Google AI Overviews, and Google AI Mode. Missing runs can look like poor visibility, while duplicate or stale runs can make a brand appear more visible than it is.
For every engine, compare the expected run count with the recorded run count. Then check whether each result contains the fields your decision needs. A response without citations is not necessarily a failed collection, but a missing response body, missing engine label, or missing timestamp is a data-quality failure.
Use a simple completeness review:
- Count expected prompts by engine.
- Count completed responses by engine.
- Identify retries, timeouts, and duplicated responses.
- Confirm that citation URLs were captured exactly enough to inspect.
- Mark unavailable surfaces instead of converting them to zero visibility.
When one engine has a collection problem, quarantine that slice and continue auditing the complete engines. Never average an unavailable run into the result as if the brand was absent. A clean missing-data label is more useful than a precise-looking score built from incomplete observations.
Reconcile mentions, citations, and answer position
The core reconciliation step is to inspect the response text and source list together, then classify the brand’s actual role in the answer. A brand can be first in the narrative, listed later, mentioned only in a caveat, cited without being named, or absent while one of its pages is used as background.
Use consistent rules for each outcome. For example, count a direct brand name separately from a page citation, record the first meaningful position separately from the total number of mentions, and flag a citation that supports a competitor rather than your brand. Apply the same rules across ChatGPT, Perplexity, Gemini, Claude, Grok, Google AI Overviews, and Google AI Mode.
Inspect the raw answer before trusting an extracted field. Automated parsing may mistake a cited domain for a brand mention, treat a list position as an answer position, or miss a brand inside a table or generated summary. Keep the original response available for every disputed record.
The [AI Visibility Metrics] page can help define the separate measures before you reconcile them. The important audit result is not a single number. It is a traceable explanation of what the engine said, what it cited, and how your measurement system classified that result.
Separate source changes from answer changes
A change in an AI answer does not prove that your site caused the change. An engine may change its retrieval, model, index, citation policy, or surrounding search features while your pages remain identical. The audit should therefore compare the answer, cited sources, and tracked page changes as three related but separate records.
Create a change log beside the visibility data. Record publishing dates, title or template changes, redirects, schema edits, robots changes, major product updates, and changes to the prompt set. Then compare the timing of each event with the first run that changed.
Use this decision rule:
- If the answer changed and your cited page changed, inspect the page and its supporting evidence first.
- If the answer changed but your pages did not, treat the movement as an engine or retrieval event until repeated evidence says otherwise.
- If the cited page changed but the answer did not, do not claim the edit worked yet.
- If several engines change after the same page update, investigate the shared source before assuming a causal result.
This separation prevents a common reporting error: turning correlation into a content recommendation. It also tells you what to test next instead of encouraging broad edits based on one surprising response.
Test repeatability and investigate outliers
A visibility movement is actionable only when it survives a repeatability check appropriate to the engine and prompt. Generative answers can vary, so one response should be treated as an observation, not a trend or a verdict about the whole category.
Repeat the exact prompt under the same recorded conditions, then compare the raw answers rather than only the extracted labels. If the answer varies, record the range of outcomes and the reason for the classification. If the answer is stable but the citation changes, investigate source selection separately.
Illustrative example: a marketing lead runs the exact prompt “Which payroll tools work for a ten-person consultancy?” in ChatGPT and sees the brand named second with a product page cited. The lead repeats the prompt under the same conditions and sees the brand omitted, while a competitor’s comparison page is cited. The correct action is to label the prompt unstable, inspect both raw responses, and run the same check across the other six engines before rewriting the product page.
Flag, rather than hide, unusual records:
- A response that stops halfway through.
- A citation that returns an error or unrelated page.
- A sudden answer change after no site change.
- A prompt answered in a different language.
- A result that appears copied from an earlier response.
Outliers can reveal collection defects or meaningful retrieval behavior. They should trigger investigation, not automatic optimization.
Choose the tracking method and assign ownership
The right tracking method depends on whether the team needs occasional diagnosis, repeatable monitoring, or monitoring connected directly to fixes. Manual checks are useful for understanding a confusing answer, but they are slow to repeat and easy to record inconsistently. A spreadsheet can preserve prompts and notes, but someone still has to run every engine, capture sources, and maintain the change log.
Cituna is one option for teams that want the measurement and the work connected: it asks ChatGPT, Perplexity, Gemini, Claude, Grok, Google AI Overviews, and Google AI Mode the questions buyers ask, then records mentions, citations, positions, competitors, and pages appearing instead. It also generates fixes for identified gaps, including schema, FAQ markup, llms.txt, and page changes.
Choose manual review when the question is narrow and the team needs to understand one disputed answer. Choose a repeatable platform when the team needs daily collection across seven engines, consistent fields, and a path from a verified gap to an assigned content change. Whatever method you choose, name one owner for prompt governance, one owner for source review, and one owner for approving changes.
Review the current [AI visibility tracking pricing] page when comparing recurring monitoring with manual work. Treat price as only one part of the decision. Data completeness, raw-answer access, correction workflow, approval controls, and the time required to validate a change are equally important.
Turn verified gaps into the next controlled action
The final audit step is to select one verified gap, make one targeted change, and define how the next run will be judged. Do not rewrite several pages at once when the audit is meant to tell you what caused movement.
Start with the gap that has a clear buyer intent and a clear source problem. For example, if the same product-comparison prompt repeatedly names competitors and cites their comparison pages, check whether your site has a directly comparable page, clear claims, accessible supporting content, and suitable structured data. A structured data review should use a separate [Structured Data for AI Search] checklist rather than assuming markup alone will change an answer.
Use this handoff sequence:
-
Save the raw response and the extracted result.
-
State the gap in one sentence, including the prompt and engine.
-
Choose one page or content change that addresses the gap.
-
Record the publication date and the expected evidence of improvement.
-
Re-run the unchanged prompt set after a sensible observation period.
-
Compare the edited page, cited source, answer position, and competitor results separately.
A practical next step is to run Cituna’s free AI visibility scan before deciding how much tracking you need. The check concerns crawler readiness, not brand mentions or citations, so do not treat it as proof of visibility. Use its result to remove access uncertainty, then begin a controlled tracking audit with a defined prompt set.
Related reading
- How often should I check my AI visibility
- AI Content Visibility: An 8-Step Check
- Measure AI Visibility for Long-Tail Prompts
Official sources to check
- Google Search Central (developers.google.com)
- OpenAI platform documentation (platform.openai.com)
- Perplexity API documentation (docs.perplexity.ai)
- Anthropic (anthropic.com)
Drafted with AI assistance from our own research and Search Console data, and reviewed by Rahul A before publishing. Rules and prices change; check the linked official source before you act.