How we measure, and what we refuse to score
Every number in a report comes from this procedure. It is published so you can check it, repeat it, and compare our results with any other tool's.
1. A frozen set of buyer questions
From your site, your one-line category and your competitors we generate a fixed set of questions a buyer would type into an AI assistant before choosing — "best X for Y", "A vs B", "how do I…", "is A worth it". The set never mentions your brand. It is frozen per site and category, so a re-test asks exactly the same questions and a change means something.
Before the set is frozen, we check it. Every drafted question is embedded and compared to the ones already accepted; anything at or above 0.85 cosine similarity is a near-duplicate of a question we already have, so it is rejected and the model is asked for a replacement — up to two extra rounds — rather than counted as a second real question. We also check the mix of question types (best-for, versus, how-to, evaluate) against an intended ~45/20/20/15 split and swap out whichever type has run more than 10 points over, and, where enabled, we show the model a handful of real search-suggestion phrasings for your category so its wording leans on how buyers actually search rather than one house style. Your report states how many questions were generated, how many near-duplicates were removed, the intent mix it landed on and whether real-search phrasing was used — and once the set is frozen, none of this runs again for that site and category.
2. Repeated samples through live AI search
Each question is asked several times through the engines' APIs with web search on (OpenAI's Responses API with search forced; Perplexity's Sonar API). We record the answer, every URL it cited, and whether it used live search. These are API samples: the consumer apps add memory, location and personalisation, so results there differ. We say so on every report.
3. Three separate facts per answer
- Recommended — a model judge reads the answer and decides whether your brand is presented as an option to consider, and at what position. When the judge cannot decide, the sample is marked and excluded rather than counted as "no".
- Mentioned — your brand name appears at all, including as a caution.
- Your own page cited — a page on your domain was one of the answer's sources.
4. Intervals, not points — and the question is the unit
The experimental unit is the buyer question, not the sample. Repeating one question measures that question's model noise; it does not add independent evidence about your brand. So the interval we lead with is computed over your questions (a bootstrap over questions, n = questions), and the older sample-level Wilson interval is shown beside it only to make the difference visible.
The difference decides verdicts. Two of our own scans of one site, run on the same day, came out at 42% and 31%. On sample-level intervals — 34–51% and 23–40% — each rate falls outside the other's interval, and the second falls below the first's lower bound: that arithmetic calls it a real drop. On question-level intervals — 30–55% and 18–44% — they overlap and each contains the other's point estimate: no change. The second answer is the correct one, because nothing about the site changed between the two scans. On those scans three repeats of a question carried the information of roughly 1.1–1.4 independent samples, not 3.
Every report prints its own measured repeat-to-repeat noise — how many of its questions gave the same verdict on every repeat — instead of quoting a constant. Published estimates put the floor at 3–7 percentage points (arXiv:2603.08924); we cite that as context and report our own number as the fact.
Your site not changing does not mean nothing changed. Models, indexes and retrieval move on their own, so two readings can differ with no change to your pages at all — which is why a re-test's verdict comes from the primary comparison specified in the current method version, not from a rate moving.
The unit is still the question. A question's verdict on one scan is recommended on at least half its repeats (⌈r/2⌉ of them, so a majority decides); comparing two scans produces the 2×2 table of old verdict × new verdict, which we publish on every compare page. The primary comparison specified in the current method version is an exact two-sided McNemar test on matched question-level majority verdicts. It is applied to that table's two discordant cells — the questions that moved either direction — at α = 0.05; the mid-p correction is reported beside it for reference and never decides the verdict, because the documented failure mode we are guarding against is false alerts, not missed ones. The effect is the paired difference between the two scans, in points, reported with a bootstrap interval — never a p-value alone. A compare tests the overall rate and each engine separately, so the whole family of comparisons in that compare (and, for a Fix Pack re-test, every item in the pack) is corrected together with Benjamini–Hochberg; a result the correction no longer supports is downgraded, and the page, the email and the JSON behind it are rewritten to agree. Which individual questions flipped is published too, but it describes what moved — it is never itself the test. Benjamini–Hochberg controls the expected false-discovery proportion across the tested family; it does not control the probability of any false alert. And however the verdict comes out, it only means something next to the design's own minimum detectable change: the movement this many questions, asked this many times, has about 80% power to detect. Every report states its own number. This is not a cutoff: smaller changes can sometimes be detected, and larger changes are not guaranteed to be detected.
5. Page-verified sources
Every cited URL is fetched and classified (your site, a competitor's site, a review directory, a listicle, a community thread, video, news). We check on the page itself whether it names you or a tracked competitor — never inferred from the answer text.
6. Reachability gates from the engines' own rules
Per engine and per page, we test only what the engines document: OpenAI states that sites opted out of OAI-SearchBot are not shown in ChatGPT search answers and that its fetcher does not run JavaScript; Google requires a page to be indexable with a snippet for AI Overviews and AI Mode and says Googlebot, not Google-Extended, is the control; Perplexity documents PerplexityBot but has been observed fetching through blocks, so its robots result is reported as "declared only".
Page size is not one rule for every engine. The 4 MB ceiling we fail a page on is ChatGPT's alone: it is over 4 MB — observed limit of ChatGPT's fetcher (RESONEO, 2026), not a figure any engine publishes, so we never apply it to Google or anyone else. Google's own documentation says "Googlebot crawls the first 2MB of a supported file type, and the first 64MB of a PDF" (developers.google.com/search/docs/crawling-indexing/googlebot), so a page over 2 MB gets an informational line for Google — content past the first 2 MB may never be read — and never a failing gate.
We fetch each page twice, once with the crawler's user-agent and once as a browser, to catch firewalls that answer bots differently. A user-agent fetch shows what a crawler would receive from your server; it is not proof that the engine crawled the page — user agents can be spoofed, the engines also verify crawlers by IP range, and a WAF can treat the vendor's real addresses differently from ours.
Reachability is checked for more engines than we sample answers from. The table above tests crawler and robots-rule reachability for four engines — ChatGPT search, Google AI Overviews / AI Mode, Perplexity and Claude — because all four either document or observably enforce their own crawler rules. We do not have a supported way to sample recommendation answers from Google AI Overviews or Claude, so those two are checked for reachability only: no recommended, mentioned or cited rate, and no citation count, anywhere in a report is measured on them. Today we sample answers from OpenAI's web-search API and Perplexity's Sonar API only, and every report's header names exactly which engines its numbers came from.
7. Deliberately not scored
- Schema.org / JSON-LD — a matched difference-in-differences study on 1,885 pages found no lift (AI Overviews −4.6%); ChatGPT strips JSON-LD before the model sees a page.
- llms.txt — 97% of existing files are never fetched; Google says it ignores them.
- IndexNow — no Google support; no AI engine documents using it.
- FAQ / Q&A re-formatting — controlled tests show no gain, and Q&A format slightly lowers absorption.
We do not sell work on these, and we do not promise rankings, mentions or citations. What we deliver is measured honestly, and the fixes we generate are the ones your own scan pointed at, written from facts you gave us, published only when you approve.
The similarity bands that triage which of your pages answers which question are internal cut points borrowed from a vendor study run on an embedding model that was never published; ours is text-embedding-3-small over page text. They are not calibrated on our data. On the scans we have, within one site they separate the questions you are recommended for from the ones you are not with an AUC of 0.60–0.63, where 0.50 is no separation at all; pooled across sites the figure drops below 0.50, because every site sits on its own similarity scale — one site in our set has the lowest similarities and the highest recommendation rate. So the bands order the work inside one scan and mean nothing between scans. Every report prints the figure measured on its own scan, and we describe similarity as triage — never as a prediction.
Cited studies: Ahrefs (schema, llms.txt, 1.4M-prompt citation study, self-promotion experiment), Seer Interactive (recency), Peec AI (listicle rank effect), arXiv 2607.14035 (survey of 45 studies on getting cited in AI answers), 2605.25517, 2604.25707, 2603.08924; engine documentation from OpenAI, Google Search Central, Perplexity and Anthropic.