← All guides

Guides

How to benchmark yourself against competitors in AI answers

Published 7 Aug 2026

In short

Benchmark on two axes and keep them apart. First, page signals: what a rival's page carries that yours does not — schema, entity links, an author byline, depth. That is comparable in seconds and fully fixable. Second, outcomes: who actually gets named when a customer asks. That needs real questions put to the engines repeatedly, and it is the one that decides revenue.

Why a score alone tells you nothing

"Your homepage scores 41 out of 100" prompts exactly one question: is that bad? Without a reference point the number is unusable. Forty-one against a category where everyone sits at 38 is a different situation from forty-one against a rival at 79, and the work you should do next is different too.

A comparison converts a score into a list. "They have FAQ schema, an author byline and entity links; you have none of the three" is not a verdict, it is three tickets.

Pick the right competitors

The instinct is to benchmark against the market leader. That is usually the least informative comparison you can run: a large incumbent turns up across the whole category for reasons no markup of yours can reproduce.

  • Someone your size who is being cited for a question you want — the most actionable comparison, because whatever they did is reachable.
  • The specific pages assistants already cite for your questions, whoever published them. These are the pages you are actually competing with, and they are often not your commercial rivals at all.
  • Yourself, three months ago. The comparison nobody runs and the only one that measures whether your own work did anything.

That third one matters more than it sounds. Competitive gaps close slowly and unevenly; your own before-and-after is the cleanest read on whether a change helped.

The two axes, and why mixing them misleads

Page signals and outcomes move on different timescales and for different reasons. Keeping them in separate columns stops you drawing the wrong conclusion from a coincidence.

  • Page signals — schema types, entity links, author and date, structure, depth. Comparable from a single fetch, fully under your control, and they change the day you deploy.
  • Outcomes — who gets mentioned, how often, and where in the answer. Measure them with the same questions across engines on a schedule. When they move, compare both on-site and off-site evidence before attributing a cause.

A rival who beats you on page signals and on mentions is a clear brief. A rival who loses on page signals and still gets mentioned more is the more interesting finding: it is a signal to inspect the off-site evidence — the benchmark shows the gap, not what is producing it.

Making the outcome comparison honest

Most competitive claims about AI visibility fall apart on method. Three things decide whether yours holds up.

  • Ask the same questions of both of you. A comparison across different queries measures the queries, not the brands.
  • Ask more than once. Assistants are non-deterministic — the same question can name different brands on two consecutive runs, so a single check is an anecdote. Sampling repeatedly is the only way a rate means anything.
  • Compare like engines. Being named by Perplexity and being named in a Google AI Overview are different systems with different source habits; averaging them hides where you are actually losing.

Share of voice is the metric that comes out of doing this properly: your share of the mentions across a set of questions, rather than a raw count that grows whenever you add a query.

What to do with the gap

Work the page signals first, because they are cheap, quick and entirely yours. Most of them are a template change rather than a content project, and they clear the eligibility bar that everything else depends on.

Then check whether the gap actually moved, rather than assuming. A fix that a re-check cannot detect is a fix you cannot claim — and that discipline is worth more over a year than any single tactic on this page.

Frequently asked questions

How many competitors should I track?

Three to five is usually the useful range: one leader for context, one or two your own size, and the pages actually cited for your questions. Beyond that the list stops being read, which makes it worse than a shorter one.

My competitor scores lower but gets mentioned more. What does that mean?

Almost always that the question is being decided off-site — they are in the roundups, threads or reference entries that the assistant reads when choosing who to name. Page signals set eligibility; they rarely decide selection on a competitive question.

How often should I re-run a benchmark?

Page signals only need re-checking after you change something, or roughly monthly to catch a template regression. Outcomes need a regular schedule to be meaningful at all, because a single sample of a non-deterministic system tells you very little.

See where you stand today

Run your questions across seven AI visibility surfaces and get your baseline — small runs often finish in about a minute. Free plan, no credit card.

Start free →Try the free AI Traffic Checker

Keep reading

AI share of voice: the metric that shows who owns your category's answersGEO vs SEO: what changes, what doesn'tHow to track your brand's mentions across AI assistants