Step 1: What should the baseline capture?
A useful baseline records the exact question, engine, answer date, brand mention, citation, cited page, and relevant competitors. Without those fields, a later result can show that visibility changed without showing why.
Start with the questions buyers actually ask about your category, including comparison, recommendation, problem-solving, and product-specific questions. Record the wording exactly, including punctuation and any location or audience qualifiers. Treat ChatGPT, Perplexity, Gemini, Claude, Grok, Google AI Overviews, and Google AI Mode as separate observation points because they can produce different answers from the same question.
Capture the full answer where permitted, not just a yes-or-no visibility label. Note whether your brand is mentioned, whether a page is cited, which URL is cited, and whether the answer places you among recommendations, alternatives, or background sources. Save the date and any relevant search context. The baseline is not a report card. It is the reference point that lets you identify a change, reproduce a result, and decide which evidence deserves investigation.
For more context, read How Often Do Ai Answers Change.
Step 2: How do I freeze the prompt set?
Freeze a core set of prompts before measuring change, then add new questions in a separate discovery set. Changing the questions and the measurement period at the same time makes the trend impossible to interpret.
Keep the core set stable long enough to compare like with like. Include the same wording, language, market, product terms, and buyer context in every scheduled run. Store an owner and purpose for each prompt so someone can tell whether it represents a real buying question or an experimental query. Retire a prompt only when the underlying customer need no longer exists, and keep its historical results rather than deleting them.
A separate discovery set can test newly emerging concerns, competitor names, or questions found in sales conversations and search data. Do not mix discovery results into the core trend until the prompt has a clear reason to remain. Prompt drift is a commonly missed source of false improvement. A broader or more favorable question can raise apparent visibility even when the original buyer question still produces no mention or citation.
For more context, read How Often Should I Check Ai Visibility.
Step 3: Which parts of each answer should I record?
Record the answer text, brand status, citation status, cited URL, competitor names, and answer context for every engine run. A single visibility score cannot show whether a brand was recommended, mentioned in passing, or supported by a useful page.
Separate at least four outcomes: no brand mention, brand mention without a source, brand mention with a source, and a source citation without a clear brand mention. Also note the position and role of the brand in the answer. A company listed first in a recommendation differs materially from one named as an alternative near the end. Record cited competitor pages as well, because the replacement source often points to the content or evidence that the answer currently prefers.
Keep engine-specific details separate. Google AI Overviews and Google AI Mode appear within Google experiences, while ChatGPT, Perplexity, Gemini, Claude, and Grok have their own response formats and source behavior. Do not treat a missing visible link as proof that no source influenced an answer. Record what the reader can actually verify, and mark uncertain cases for review instead of forcing them into a confident category.
Step 4: How often should I measure to find a real trend?
Measure on a consistent schedule that matches how quickly you can act, and use repeated observations before calling a change a trend. Daily collection can reveal volatility, but one unusual answer should not trigger a content rewrite.
Keep the run conditions consistent wherever possible. Use the same prompt wording, market, language, logged-in state, and device context. Record outages, interface changes, major product updates, and unusual events that could affect answers. Engines can update their systems or retrieval sources between observations, so a result can change without any change to your website.
Review results at two levels. The first is the latest run, which helps identify urgent losses or new citations. The second is a rolling comparison against the baseline, which shows whether a change persists. Look for repeated movement across related prompts, not only a single question. A citation that disappears once may be normal answer variation. A citation that disappears across several relevant prompts and runs is a stronger reason to inspect the cited page, competing sources, and recent site changes.
Step 5: When does a citation loss deserve investigation?
Investigate a citation loss when it is repeated, concentrated in related questions, or accompanied by a change in the cited source or answer role. A single missing citation is usually weaker evidence than a consistent shift across the same topic.
First check measurement conditions. Confirm that the prompt, engine, language, market, and answer date were recorded correctly. Next check whether the engine answered a different interpretation of the question. Then compare the lost citation with the page now being cited instead. Look for differences in specificity, freshness, first-party evidence, clear definitions, and coverage of the question. Avoid assuming that the longest page or the page with the strongest traditional ranking is automatically the preferred source.
Also inspect whether your own page changed. A redirect, removal, template change, broken access path, altered title, or weaker explanation can affect how an answer uses it. If the loss appears only in one engine, begin with engine-specific behavior. If it appears across several engines for related prompts, investigate the underlying topic coverage and source signals before changing isolated wording.
Step 6: What should I connect to Google Search Console?
Connect citation observations to Google Search Console data so you can compare answer visibility with the search demand and page performance behind the same topics. The connection is useful for diagnosis, not proof that one channel caused the other.
Match tracked prompts to relevant query groups and landing pages. Check impressions, clicks, average position, and page changes for the period before and after a citation shift. A page can gain search impressions while losing AI citations, or gain citations while receiving little conventional search traffic. Those cases suggest different decisions and should not be collapsed into one visibility score.
Use the comparison to prioritize review. A cited page with growing search interest may deserve stronger internal links, clearer supporting evidence, or better maintenance. A page with weak search demand but repeated AI citation may still matter for a high-value buyer question. Keep the time windows aligned and note publication, update, indexing, and technical changes. Cituna joins AI answer observations to Google Search Console data and provides SEO, AEO, and GEO fixes, which gives teams one place to connect the observed answer change with possible site actions.
Step 7: How do I test whether a change worked?
Test one meaningful change against the same prompt set and wait for repeated observations before attributing a citation gain to the change. Editing several pages, prompts, and technical elements at once removes the evidence needed to learn.
Define the intended outcome before editing. The outcome might be more citations to a specific page, clearer brand inclusion in recommendations, or better coverage of a recurring buyer question. Save the previous answer and source record, then document the page, claim, structure, or internal link that changed. Keep unrelated prompt and site changes visible in the same log.
After publication, compare the same engines and prompts against the baseline. Check whether the gain is broad or limited to one wording, whether the cited URL changed, and whether competitors remain the preferred source. Do not treat a newly visible mention without a relevant citation as the same result as a supported recommendation. If the change is positive, look for repetition across related questions. If it is neutral, review whether the edited page actually addresses the question better than the source that engines continue to cite.
Step 8: Which tracking setup fits a small team?
A small team should choose the simplest setup that preserves prompt history, engine-level results, cited URLs, and a clear action log. Manual checks can work for a small research sample, while recurring business measurement needs scheduled collection and consistent records.
A spreadsheet is suitable for an initial baseline when one person can run the questions and inspect answers carefully. Its weakness is inconsistent collection and slow comparison as the prompt set grows. A script or API-based process can standardize storage and reporting, but it still needs a review method for ambiguous mentions, changing answer formats, and source interpretation. A specialist tracker reduces collection effort when the team needs recurring coverage across several engines.
Cituna tracks whether seven AI answer engines, ChatGPT, Perplexity, Gemini, Claude, Grok, Google AI Overviews, and Google AI Mode, mention and cite a brand for buyer questions every day, and shows which competitors and pages they cite instead. Every engine is included on every plan, and Cituna does not track Microsoft Copilot. Choose based on the decision you need to make, not the largest dashboard: preserve raw evidence, review meaningful changes, and connect each finding to one accountable site action.
Related reading
- Which AI Visibility Tool Includes Google Search Console?
- How To Improve Ai Search Visibility With Answer Pages
Sources consulted
- OpenAI (platform.openai.com)
- Perplexity (docs.perplexity.ai)
- Google Search Central (developers.google.com)
- Google Search (support.google.com)
Drafted with AI assistance from our own research and Search Console data, and reviewed by Rahul A before publishing. Rules and prices change; check the linked official source before you act.