Skip to main content
AI Visibility8 min read

How to Measure Brand Sentiment in AI Answers (and Fix It)

A repeatable method for AI brand sentiment: a fixed prompt set, repeat runs, a rubric, a net sentiment score per engine, and fixes at the pages engines cite.

Published

To measure AI brand sentiment, ask the same fixed set of buyer questions to each AI engine several times, label every answer that names your brand as positive, neutral or negative using a written rubric, and compute a net sentiment score for each engine separately: positive answers minus negative answers, divided by all answers that named you. Keep factual errors in a separate count, because a warm answer with a wrong price is a different problem from a cold one. Then fix the scores at their source, which is usually a page the engine cited.

Why AI brand sentiment needs its own method

Social listening tools count thousands of posts, so one odd mention barely moves the score. AI answers are the opposite. A tracked prompt produces one answer per engine per run, and your brand may appear in only a handful of them. Three things follow.

Samples are small. In our own scan of October 4, 2026, cituna.com was tracked on 34 prompts across seven engines, and the engines named Cituna in 10 answers. A sentiment score built on 10 answers moves a lot when one answer changes, so the method has to collect more runs before it reports a trend.

Answers change between runs. Across 25 scans of cituna.com between September 29 and October 4, 2026, the prompt “Scrunch AI alternative” named Cituna in 14 scans and not in the other 11. The tone of an answer swings the same way. One run is an anecdote. Our guide on how often AI answers change looks at that volatility in detail.

Engines disagree. Of those 10 named answers, 6 came from ChatGPT, 2 from Claude, 1 from Grok and 1 from Google AI Overviews. Blend them into one score and you mostly get a ChatGPT score. Each engine reads different sources, so each needs its own number.

The method, step by step

1. Fix the prompt set

Pick 15 to 30 questions and do not change them during the measurement period. Mix three kinds:

  • Brand questions: “Is [brand] any good?”, “What do people say about [brand]?”, “[brand] pricing”. These ask for an opinion, so tone shows up most clearly.
  • Category questions: “Best [category] for [use case]”. These show how you are described when you are one name among several.
  • Comparison questions: “[brand] vs [rival]”, “[rival] alternatives”. These show whether you are framed as the stronger or weaker option.

Write the list down with the date. If you change a prompt, start a new series rather than comparing across it.

2. Run every prompt on every engine, more than once

Run each prompt on each engine you care about, then repeat. Daily runs over one to two weeks give enough answers to see a direction. Store the full answer text, not only whether you were named: sentiment lives in the sentences around your name.

3. Label each answer that names you

Score the answer as a whole, as it would read to a buyer, using a fixed rubric:

A sentiment rubric for AI answers that name your brand
LabelWhat it looks like in the answerExample wording
PositiveRecommends you, or names a strength as a reason to choose you“a strong choice for small teams because...”
NeutralNames you with facts and no judgement, or lists you without comment“Other options include [brand]”
NegativeNames a weakness as a reason not to choose you, or recommends a rival over you for the reader’s stated need“more expensive than most and harder to set up”
Inaccurate (tracked separately)States a wrong fact about you, whatever the toneA retired plan, an old price, a missing feature

The example wording is ours, written to show the rubric, not quoted from an engine.

Two rules keep the labels honest. First, have two people label the same 20 answers and compare; where they disagree, tighten the rubric text. Second, never let an inaccurate answer count as positive just because it is friendly. A glowing answer with a wrong price still sends a buyer to the wrong number.

4. Score each engine separately

For each engine, use the net sentiment formula common in brand monitoring: positive answers minus negative answers, divided by all answers that named you, times 100. Thematic publishes it in this form. Some tools, such as YouScan, divide by positive plus negative only and leave neutral answers out, which gives a higher score; pick one and keep it. The score runs from -100 (every mention negative) to +100 (every mention positive).

Worked example

Illustrative numbers, not data from any engine: over two weeks, one engine names your brand in 40 answers. 22 are positive, 12 neutral and 6 negative. (22 - 6) / 40 x 100 = +40. With the positive-plus-negative version, the same answers score (22 - 6) / 28 x 100 = +57. Neither is wrong, but the two cannot be compared with each other.

Report three numbers next to it, because the score alone hides them:

What to report for each engine
Per engineWhat it tells you
Net sentimentThe balance of tone when you are named
Answers that named youHow many answers the score rests on. Below about 20, treat it as a direction, not a figure
Inaccurate answersHow many answers state a wrong fact, whatever their tone

Keep how often you are named apart from how you are described. An engine that names you twice, warmly, has a high score and almost no reach. The share of answers that name you is a separate measure, covered in AI share of voice.

5. Find what each negative answer is built on

For every negative or inaccurate answer, read the sources that answer cited. Engines take their framing from the pages they read: a review with a recurring complaint, a comparison post written before your last price change, a forum thread. Note the page and the sentence the engine echoed. After a week or two, the same few pages usually account for most of the negative answers.

6. Re-measure on the same prompts

After you act on a source, keep running the same prompt set. Compare the per-engine scores on equal-length windows before and after. Expect the change to show first on engines that search the web live for that prompt, and later, if at all, where the answer leans on what the model learned in training.

How to fix negative or inaccurate AI sentiment

Engines repeat what they read, so the fixes are upstream of the answer.

  • Correct the cited source. If a negative answer cites a review or comparison with outdated facts, ask its author to update it, with the evidence. This is usually the fastest fix for a specific complaint. The steps are set out in how to fix wrong AI information about your business.
  • Answer the complaint on your own site, plainly. If engines say you are expensive, a pricing page that states each plan and what it includes, in server-rendered text, gives them something accurate to quote. If they say a feature is missing that you have, a page that names it in its first paragraph does the same.
  • Make the facts agree everywhere. The same name, category, price and one-line description on your site, your review profiles and directory listings. Conflicting facts make engines hedge, and hedged answers read as lukewarm.
  • Earn coverage where the engines look. The sites each engine cites for your category are where new, accurate descriptions of you have the most effect. Track which sites those are per engine; they differ more than most teams expect.

Sentiment is one of several ways an answer frames a brand, alongside category, comparison set and accuracy. Our guide on how AI engines shape brand perception covers the other dimensions.

Doing this with Cituna

Cituna does the collection steps. It runs your tracked prompts on all seven engines every day, stores the full answer text each engine gave, with the brands it named and the sources it cited, and keeps that history so you can compare runs. Cituna does not assign sentiment scores: the labelling and the net sentiment arithmetic in steps 3 and 4 are done by you or a reviewer on that text, with the rubric above. For step 5, its Sources page shows the pages each engine cites for your prompts and whether your brand is named on them, which is where most fixes start. It also writes the fix for each answer you are missing from.

Plan size matters for this method. Starter checks 10 prompts daily and Pro 30, so a full 15 to 30 prompt set needs Pro. The free trial tracks only your top 3 prompts daily, which is enough to see how the answers read, not to run the full method.

Start with the answers themselves

AI brand sentiment is measured one engine at a time, on a fixed prompt set, over repeated runs, and fixed at the pages engines read. Cituna asks ChatGPT, Perplexity, Gemini, Claude, Grok, Google AI Overviews and Google AI Mode your buyers’ questions every day, keeps every answer’s text and sources, and writes the fix for each answer you are missing from.

Drafted with AI assistance from our own research and Search Console data, and reviewed by Rahul A before publishing. Rules and prices change; check the linked official source before you act.

Frequently asked questions

How do you measure brand sentiment in AI answers?

Ask each AI engine the same fixed set of buyer questions several times, label every answer that names your brand as positive, neutral or negative with a written rubric, and compute a net sentiment score for each engine separately: positive answers minus negative answers, divided by all answers that named you, times 100. Count factual errors separately, whatever their tone.

What is a good AI brand sentiment score?

There is no published benchmark for net sentiment in AI answers, and scores depend heavily on the prompt set: brand questions draw more opinion than category lists. Compare each engine against its own earlier windows on the same prompts, and treat a falling score or a rising count of inaccurate answers as the signal to act.

How many answers do I need before the score means anything?

More than a single run gives you. With fewer than about 20 named answers per engine, one changed answer moves the score by five points or more, so report it as a direction. Running daily for one to two weeks on 15 to 30 prompts usually gets past that.

Should I use a language model to label sentiment?

It can speed up the first pass, but check it against human labels on a sample before you trust it. Keep the same rubric for both, and have a person review every answer labelled negative or inaccurate, since those are the ones you will act on.

Does Cituna score sentiment?

No. Cituna does not assign sentiment scores. It asks seven AI engines your buyer questions every day and stores each full answer with the brands named and the sources cited, which is the raw material for the labelling and net sentiment arithmetic described here.

Find your next AI visibility fix with Cituna

Cituna asks ChatGPT, Perplexity, Gemini, Claude, Grok, Google AI Overviews and Google AI Mode your buyers' questions every day, writes the fix for every answer you are missing from, and publishes new articles to your site. Run all of it from Claude or any AI agent.

Start free trial

3-day free trial · Card required, cancel anytime · Plans from $39 a month

Check crawler readiness free