Every number in visitedby.ai is a measurement, which means it has a method, a sample size and an error bar. This page is the method. It is written so you can check our figures rather than trust them, and it states the limits of what we measure as plainly as the strengths, including the places where we are still validating our own statistics.
01WHAT WE ASK, AND HOW OFTEN
We track a set of prompts you approve. Each prompt runs once per engine per day.
Asking the same question ten times in one afternoon measures that afternoon. The engines change from day to day, and different questions surface different brands, so the variation worth measuring sits across days and across prompts rather than inside a single sitting. Ten runs today and one run a day for ten days produce a similar number of answers, and only the second tells you what the engine did over ten days.
So a rate is pooled across every tracked prompt over a rolling window. Twenty-one days of ten prompts on two engines rests on about 420 answers.
One exception reduces sampling rather than increasing it. A prompt whose rate has sat at or below 10%, or at or above 90%, for fourteen days is sampled every third day. A question that has answered itself for a fortnight is not where the next finding is.
02WHICH ENGINES, AND THROUGH WHAT
We measure ChatGPT, Perplexity, Gemini and Google AI Overviews. The first three are queried through their official APIs; AI Overviews is read from the search results page, because Google decides whether an overview appears at all. “No overview shown” is a real result and we record it as one.
Web search is forced on, and an answer that arrives without it is thrown away. Search-on and search-off answers diverge sharply, so a run where the model chose not to search belongs to a different distribution; recording it would blend two populations into one rate. Our runners refuse such an answer and retry rather than store it.
03THE PROXY, STATED PLAINLY
Measuring through an API is not the same as using the consumer app. The consumer products wrap the same model families in their own system instructions, tool orchestration, personalisation and product logic. Our numbers are a proxy for what those apps answer, not a copy of it.
We could narrow that gap by driving the consumer interfaces with browser automation. We do not, because it violates their terms of service. We think the honest option is to measure a surface we are allowed to measure and tell you exactly what it is, which is what you are reading now rather than discovering later.
04THE MODEL IS PART OF THE MEASUREMENT
A visibility number belongs to a specific model, not to a brand name. Every run we store records the engine, the exact model version that answered, whether search was enabled, and the locale, so any answer in your history can be traced to the conditions that produced it.
This matters because engines change. If we change the model behind an engine, the series will step for reasons that have nothing to do with your visibility. When that happens we mark it on your charts rather than letting it read as a trend, and we report the two periods as two different measurements.
05HOW A RATE IS COMPUTED
Each answer is one observation: your brand was named, or it was not. A rate is the number of answers naming you divided by the number of answers, pooled over the window. An answer that names you three times still counts once. Otherwise a wordy answer would outvote a decisive one.
Every rate carries a 95% Wilson confidence interval. It is the honest width of the estimate, and it is wide when the sample is small.
06DISCOVERY AND BRANDED PROMPTS ARE NEVER POOLED
A prompt that contains your brand name guarantees you appear in the answer. Pooling those with the rest measures how many of your prompts contain your own name, and it always flatters.
So the two populations are reported separately and never combined into one figure. Your headline visibility comes from discovery prompts only: the questions that do not mention you. Branded prompts get their own block, shown as counts rather than a percentage, because a percentage there invites exactly the comparison that should never be made. On our own data the difference between these two readings has been more than thirty-fold.
07WHAT COUNTS AS A CHANGE
Two rates being different is not the same as something having changed. We compare two periods with a two-proportion z-test at 95% confidence, and say “no significant change” whenever the test does not clear, even when the two percentages look different on screen.
We deliberately do not use the more common shortcut of asking whether two confidence intervals overlap. That test sounds stricter and is: it sits at roughly p<0.005 rather than p<0.05, and it discards about a third of the sensitivity the measurement was paid for. Movements that clear 80% but not 95% are shown as emerging and labelled as not yet significant, rather than being hidden or promoted.
08WHERE OUR STATISTICS ARE STILL BEING VALIDATED
One open limitation, stated here because you would otherwise have no way of knowing it. Confidence intervals and significance tests of this kind assume independent observations. Ours are not fully independent: a window is the same prompts and engines measured on consecutive days, so answers cluster within a prompt, and each day’s answers are related to the day before through changes in the engines themselves.
The practical consequence is that our intervals may be somewhat narrower than the underlying uncertainty warrants. We are measuring this directly, by running the same portfolio at different sampling rates and comparing our published intervals against a method that models the structure of the data. We will publish the result here and widen the bands if that is what it shows. We would rather disclose an unfinished validation than present a settled-looking number we have not checked.
09WHAT WE REFUSE TO SHOW
A measurement that cannot support a claim should not be dressed as one. So:
- below five answers, a percentage is not shown. The counts are, because those are what happened;
- a trend line needs fourteen days of measurement before it is drawn at all, because three points and a partial day draw a collapse that is not there;
- an automated check that cannot see its inputs reports inconclusive rather than passing;
- we never report AI prompt volume. Nobody can measure that without a licensed conversation panel and we do not have one. Where we show demand, it is Google search volume for a named keyword, printed with that keyword beside it.
10HOW ANSWERS ARE READ
We store every answer verbatim. Your brand and your tracked competitors are found by exact text matching, not by a model’s opinion. A language model then reads the same answer to find the other companies named, the sources cited, sentiment and the specific aspects of your brand the answer discusses. Anything it reports that cannot be located in the original text is discarded.
Because the raw answers are kept, improvements to this reading are applied to history rather than only to new data. Each version is recorded alongside the rows it produced, so a re-read never overwrites what a previous version found.
11LIMITS WE WILL NOT PAPER OVER
- API answers are a proxy for consumer-app answers (clause 03);
- your prompt set is chosen by you, and a portfolio that does not reflect how buyers actually ask will produce a precise measurement of the wrong thing;
- engines change underneath everyone, and no amount of sampling on our side removes that;
- small prompt portfolios produce wide intervals. We show the width rather than hiding it.
If you find a number on this page that disagrees with a number in the product, the product is the measurement and this page is the description. Tell us, and we will fix whichever one is wrong.
