A visibility rate is a proportion, and proportions have intervals
Most AI visibility tools report one number per period: the share of tracked prompts whose answer names your brand. Statistically that is a proportion, k answers out of n prompts, and a proportion measured on a sample has a range of true values it is consistent with. This guide gives that range for the prompt counts people actually track, and shows how the count changes it.
- n is the number of prompts (one answer each, unless a section says otherwise); k is how many of those answers named you.
- The interval is the 95% Wilson score interval, without continuity correction. It is the method the Promvia dashboard uses for the rates it prints with an interval. Promvia is our product.
- "±X points" means half the width of the interval. Wilson intervals are not symmetric around the observed rate, so the bounds are given as well.
- The interval covers sampling noise: which prompts you happened to pick and which answer you happened to get. It cannot fix a prompt set that leans toward questions you already do well on.
The textbook formula (the Wald interval, p ± 1.96·√(p(1−p)/n)) breaks at small counts, and that is the reason for Wilson. For 0 mentions in 20 prompts it gives 0% to 0%, a range of nothing on exactly the sample that most needs a warning; Wilson gives 0% to 16.1%. For 1 in 20 it gives −4.6% to 14.6%, a negative rate; Wilson gives 0.9% to 23.6%.
The table: 5 to 220 prompts
The pool sizes below are the tracked-question pools of Promvia's plans on 7 October 2026 (Free 5, Starter 20, Pro 80, Agency 220) plus 40, the size of our public AI Source Index panel. Two rates for each: 20%, a typical rate for a brand that is present but not dominant, and 50%, where the interval is widest.
- 5 prompts: 20% (1 of 5) → 3.6% to 62.4%, ±29.4 points. 50% → 17.0% to 83.0%, ±33.0.
- 20 prompts: 20% (4 of 20) → 8.1% to 41.6%, ±16.8 points. 50% → 29.9% to 70.1%, ±20.1.
- 40 prompts: 20% (8 of 40) → 10.5% to 34.8%, ±12.1 points. 50% → 35.2% to 64.8%, ±14.8.
- 80 prompts: 20% (16 of 80) → 12.7% to 30.0%, ±8.7 points. 50% → 39.3% to 60.7%, ±10.7.
- 220 prompts: 20% (44 of 220) → 15.2% to 25.8%, ±5.3 points. 50% → 43.4% to 56.6%, ±6.6.
Quadrupling the prompts roughly halves the interval: 20 to 80 prompts takes ±16.8 to ±8.7 points at 20%. That square-root law is why the jump from 5 to 20 prompts buys so much and the jump from 80 to 220 buys less.
How many prompts for a given margin
Turned around: the smallest number of prompts, one answer each, whose 95% Wilson interval is no wider than the margin. 50% is the worst case, so plan on that number if you do not know your rate yet.
- ±20 points: 21 prompts at 50%, 14 at 20%.
- ±15 points: 39 prompts at 50%, 26 at 20%.
- ±10 points: 93 prompts at 50%, 60 at 20%.
- ±5 points: 381 prompts at 50%, 245 at 20%.
The textbook formula asks for slightly more at these margins: 97 for ±10 and 385 for ±5 at 50%. Those are the "about 100" and "almost 400" in Peec AI's post on the same question (quoted below). The two methods agree away from the edges; they part company at small counts and at rates near 0% or 100%.
A one- or two-question move is noise
The number people react to is the change from one week to the next. A change between two rates has its own interval, and it is wider than either rate's. We use Newcombe's method, which builds it from the two Wilson intervals and treats the weeks as independent samples. For the same prompts measured twice that is conservative: if most prompts give the same verdict both weeks, a paired comparison can be tighter.
- 20 prompts, 4 → 5 mentioning answers (20% → 25%): change +5 points, 95% interval −20.6 to +29.9.
- 20 prompts, 4 → 6 (20% → 30%): +10 points, interval −16.6 to +34.9.
- 20 prompts, 4 → 8 (20% → 40%, the rate doubled): +20 points, interval −8.2 to +44.5. Still consistent with no change.
- 80 prompts, 16 → 18: +2.5 points, interval −10.2 to +15.1. 16 → 24: +10 points, interval −3.4 to +23.0.
- 220 prompts, 44 → 50: +2.7 points, interval −4.9 to +10.4.
Starting from 20%, the smallest week-to-week gain whose interval clears zero is 4 → 10 on 20 prompts (+30 points), 8 → 17 on 40 (+22.5), 16 → 28 on 80 (+15) and 44 → 62 on 220 (+8.2). Below those, the honest reading of a move is "no detectable change", in either direction.
Repeated runs and pooled weeks narrow it, up to a point
AI engines do not give the same answer every time. In our AI Source Index, where each of the six engines answers each of 40 fixed questions once per weekly run, 48.0% of the domains an engine cited for a question were cited again for the same question the next week (3,335 of 6,947; 95% interval 46.8% to 49.2%). A second run of the same prompts is far from a copy of the first, so running prompts again, or pooling several weeks, adds information.
It does not add as much as new prompts: answers to the same prompt are correlated. With R runs of n prompts and a run-to-run correlation ρ, the runs are worth n·R / (1 + (R − 1)·ρ) independent answers. We have not measured ρ for brand mentions, so here is the bracket for 20 prompts at a 20% rate:
- 1 run: ±16.8 points, whatever ρ is.
- 4 runs (80 answers): ±8.7 if every answer were independent (the same as 80 prompts), ±13.5 at ρ = 0.5, ±16.8 if every repeat were a copy.
- 8 runs (160 answers): ±6.2 if independent, ±12.8 at ρ = 0.5, ±16.8 if copies.
At ρ = 0.5, going from 4 to 8 runs moves the margin from ±13.5 to ±12.8 points; going from 20 to 80 prompts moves it to ±8.7. Runs and prompts also answer different questions. More runs pin down how often these particular prompts name you. Only more prompts pin down how often prompts like these would.
Split branded and unbranded prompts
A prompt that names your brand ("is Acme any good?") gets an answer about you nearly every time, so it says little about whether buyers who do not know you will hear of you. Mixed into one rate, branded prompts inflate it while the interval stays about as wide.
- 20 prompts, 5 of them branded and all 5 naming you, plus 3 of the 15 unbranded: pooled rate 8 of 20, 40% (21.9% to 61.3%).
- The unbranded rate alone: 3 of 15, 20% (7.0% to 45.2%). Half the pooled figure, on fewer prompts.
- The branded rate alone: 5 of 5, 100% (56.6% to 100%). Five prompts cannot say much even at the ceiling.
Report the two separately, and size the unbranded set for the margin you need, because that is the rate you are trying to move. The Promvia dashboard keeps prompts that name your brand out of its share-of-voice figure for this reason.
What other tools say about prompt counts
Otterly's ranges shrink the way the square-root law predicts: scaling its 32-point range at 10 prompts by √(10/n) gives 14.3 points at 50 prompts and 10.1 at 100, against the 13 and 9 it measured. Its subsets came from a fixed pool of 320 prompts over a 30-day window, so the match is in shape rather than to the point. Peec's post is the only one of the four that gives an interval, and it uses the Wald approximation, which its own author notes does not hold near 0% or 100%, where the rates of smaller brands sit. SE Ranking's and Promptwatch's counts are starting points without an interval attached. The tables above put an interval on any of those counts.
What to do with this
- Print the interval next to every rate you report, and the counts behind it ("4 of 20", not just "20%").
- Decide the margin you need first, then size the prompt set: 39 prompts for ±15 points in the worst case, 93 for ±10.
- Do not report a week-to-week change of one or two answers on a set of 20 as a trend. Compare four-week windows, or the same prompts over several runs, before concluding.
- Keep branded and unbranded prompts apart, and size the unbranded set, which is the one that measures discovery.
- Add prompts before adding runs once a set is past a few runs: new prompts narrow the interval faster and widen what it describes.
How Promvia samples, counts and prints intervals is on the methodology page. For setting up a prompt set in the first place, see how to track AI mentions; for how much the cited sources themselves change from week to week, see how long an AI citation lasts.
Frequently asked questions
How many prompts should I track for AI visibility?
Enough for the margin you need. For a 95% interval of ±15 points, 39 prompts in the worst case (a 50% rate); for ±10 points, 93; for ±5 points, 381. At a rate near 20% the counts are 26, 60 and 245.
Are 10 or 20 prompts enough?
For a rough level, yes. For tracking change, not on their own. At 20 prompts a 20% rate has a 95% interval of 8.1% to 41.6%, and a move from 4 to 8 mentioning answers between two weeks is still consistent with no change.
Does running the same prompts more often help?
Partly. Repeat answers differ (in our AI Source Index, 48.0% of cited domains were cited again the next week), so extra runs add information, but answers to the same prompt are correlated. At a run-to-run correlation of 0.5, eight runs of 20 prompts give about ±12.8 points at a 20% rate; 80 prompts run once give ±8.7.
Which confidence interval should I use for a visibility rate?
The Wilson score interval. The textbook Wald formula gives 0% to 0% for 0 mentions in 20 prompts and a negative lower bound for 1 in 20; Wilson gives 0% to 16.1% and 0.9% to 23.6%.