Blog

We Read Eleven GEO Vendors' Documentation. Not One Publishes a Margin of Error.

GEO AI visibility tools accuracy: an AI visibility score shown as a single number with no error bar, next to a list of cited sources changing between scans.

By Justin Schroder, CTO and co-owner, SEMoptimize

Nobody has shown that a GEO tool's visibility score predicts anything downstream, and across the eleven vendors whose own documentation we read in August 2026, not one attaches a margin of error to the number it sells. Olivier Martinez's July 2026 review of 45 studies (arXiv:2607.14035) graded the claim that "citation scores predict clicks, conversions, or revenue" at very low confidence. The IAB counted more than 20 companies selling these tools that same month, so eleven is a sample and not a census. SEMoptimize is a Denver paid search agency, we have never held an account with any AI visibility platform, and we ran no correlation study. None of that was going to answer the question, so we ran our own panel instead: the same prompts, the same engines, three weeks, every cited source logged. Getting your brand into AI answers has a name now, generative engine optimization, and a budget line. Whether that work can be measured is the part nobody has settled.

Not one of the eleven publishes a margin of error on the score it sells

Across eleven vendors' own documentation in August 2026, not one publishes a confidence interval, a margin of error, a standard deviation, or any other statement of uncertainty on its headline visibility metric. Every score we found ships as a bare point estimate, with no range and no sample size beside it. Unvalidated is the accurate word for that, not inaccurate, because nobody has tested it in either direction.

The closest anyone comes is a blog post. Semrush's Head of Organic and AI Visibility wrote in May 2026 that the platforms are nondeterministic within a single day, then gave the honest reporting format: "a share of voice that swings between 20% and 40% over a day is normal, so '30% plus or minus 10%' is the honest way to report it." That is correct. It also appears nowhere in Semrush's published product documentation.

Credit where it belongs. Conductor refuses to sell an AI search volume metric, because "there is currently no definitive or reliable source for AI search volume data." Scrunch logged a definition change in December 2025 that lowered its own customers' presence scores, and Otterly ships no blended composite at all. At the other end, seoClarity advertises a "statistically significant presence rate" with no sample size published, and BrightEdge publishes no methodology of any kind.

Does getting mentioned in an AI answer bring anyone to your site?

AI mentions are associated with real site visits, and there is now a quasi experiment with published confidence intervals behind that association. Profound, an AI visibility platform, joined an AI interaction stream to a browsing stream across a double opt in US panel of more than 2 million conversations, January to June 2026 ("The AI mention effect," July 2026). Visits to a brand's site in the seven days after an assistant introduced it ran above the forecasted baseline: in Google AI Overviews, 7.79 percent treated against 4.83 percent forecast, a lift of 2.96 points, 95 percent interval 2.76 to 3.17. More than 97 percent of those visits carried no UTM parameter, which is why your analytics has been quiet.

Profound sells the product this research supports and says so, which is more than most of this category manages. The limits come attached as well: a site visit study, not a purchase study, and not randomized. Discount it for both and it is still the strongest evidence anyone has published that getting mentioned matters. Whether a vendor's score counts your mentions correctly is a separate question with much thinner evidence behind it.

What does an AI visibility score actually measure?

Strip the packaging off and a visibility score is one thing: the share of AI answers that mention your brand, inside a sample the vendor picked. Peec AI, a GEO tracking platform, will at least tell you the formula, mentions divided by responses: "The Visibility Score measures the percentage of AI responses that mention your brand" (docs.peec.ai, accessed August 2026).

Share of voice is your slice of some total, and the total is where the number quietly stops being comparable. Conductor says it out loud: "your share of voice will tend to be much lower than your percentage of prompts with mentions," and both vendors can still be measuring honestly. Read that again before you put two numbers side by side.

Ahrefs weights Brand Radar impressions by Google search volume, then tells you not to lean on that ingredient: "This is a modeling choice, not a measured relationship: we don't claim a validated link between Google search volume and how often a query is asked inside an AI tool." That caution is about one input to one vendor's own metric, not about vendor to vendor divergence.

How many times should a tool ask before it reports a number?

The published recommendation is at least seven runs per prompt per day, and at least eight when you care which sources got cited. That comes from Julius Schulte, Malte Bleeker and Philipp Kaufmann (arXiv:2604.07585, April 2026), who ran 32 prompts across four engines over 45 days plus a simultaneous re-run set of 1,280 responses. One run carries a standard error of 0.370, so a true per brand detection rate of 50 percent "could appear anywhere from -22% to +122% in a nominal 95% interval." Two disclosures. Julius Schulte is footnoted as affiliated with Aurora Intelligence, a commercial AI visibility vendor, which supplied the study's temporal data, for a paper arguing the industry needs seven times more sampling. The work is Swiss German only and the authors decline to generalize past it.

Now the arithmetic you can check. Peec AI works its own example in public on its pricing FAQ: "If you are running 25 prompts across 3 models for 30 days, we would analyze 25x3x30 = 2250 AI answers." That is one run, per prompt, per model, per day, a seventh of what Schulte and colleagues recommend, and the lowest sampling rate any of the eleven discloses. Peec is ahead of the category on disclosure rather than sampling, since you can only check that arithmetic because they published the inputs. Scrunch writes that "you'd want to run each multiple times to get a reliable signal," then never says how many times it runs them.

We ran the same prompts for three weeks and roughly half the sources changed

We put 20 seed prompts to the engines on 11 capture dates between 2026-07-29 and 2026-08-19, logging every cited domain, and roughly half the sources cited for a given prompt turned over between one scan and the next.

Our own sample, stated the way we are asking vendors to state theirs: the sheet holds 727 logged rows, of which 634 are usable domain observations across 259 unique domains. Forty are capture failures: all 20 ChatGPT rows and all 20 Gemini rows were logged as failures rather than guessed at, so we have zero data on either engine and everything below covers Google AI Overviews and Perplexity only. One further row was excluded because it records the loss of a citation rather than the presence of one. Scans sat 3 to 14 days apart and one date, 2026-08-03, carries 382 of the rows, so coverage is uneven. A single operator captured everything, no second coder verified the domain extraction, and we never ran a prompt twice in a sitting, which makes this day to day drift, not within session variance.

Across 26 repeat scan pairs, each required to have at least 4 domains captured in both scans so a thin capture could not masquerade as churn, mean source overlap was 0.456 and the median 0.500, with a range of 0.05 to 1.00. Google AI Overviews averaged 0.439 across 12 pairs and Perplexity 0.470 across 14. Of the 259 domains cited, 144, or 56 percent, appeared in one scan and never again.

Our own result was a loss and I would rather print it. On prompt one, asked of Google AI Overviews, semoptimize.com was absent on 07-29, present on 08-03, absent again on 08-19. Three domains held all three scans, and we were not one of them.

Why do two AI engines answer the same question differently?

Each engine retrieves its own sources, and the agreement between two of them is far worse than either one's agreement with itself. We captured Google AI Overviews and Perplexity for the same prompt on the same day 37 times, under the same 4 domain floor. Mean source overlap was 0.234, the median 0.231, and the two never agreed on more than 56 percent of sources in any pair.

The one publicly documented mechanism that would produce an effect like this belongs to an engine we have no usable data on. OpenAI's ChatGPT Search help article explains that ChatGPT may work out your area from your IP address and rewrite your prompt into a search query naming that city, and that Memory feeds the rewrite as well. The rewritten query decides which pages get retrieved, so personalization sits upstream of citation. All 20 of our ChatGPT rows were capture failures, so read that as the clearest published case of a single prompt not being a single question, and not as an explanation of our 0.234.

That is the whole of our first hand evidence, and the published work gives it a place to sit. Schulte and colleagues report same day source overlap of 0.32 to 0.43 in Swiss German. We found 0.456 day to day in US English with different prompts and a different operator, and since we never repeated a prompt inside a day ours is a conservative floor.

Does source churn make a visibility score useless?

No, and the loudest critic in the category reached that conclusion before we did. Rand Fishkin and Patrick O'Donnell of SparkToro ran 2,961 runs with 600 volunteers across ChatGPT, Claude and Google's AI surfaces, published January 2026, and found "there's a <1 in 100 chance that ChatGPT or Google's AI, if asked 100X, will give you the same list of brands in any two responses." Fishkin also wrote that "any tool that gives a 'ranking position in AI' is full of baloney," then reversed half of his own starting position in the same post: "I've been swayed from my initial position and now believe visibility % across dozens to hundreds of prompts run multiple times is a reasonable metric."

That reversal is the strongest objection to our own data. What we measured is source set churn, and what a vendor sells is a brand presence rate. A presence rate averaged over hundreds of prompts and many runs can be far more stable than the source list underneath it, because averaging is exactly what fixes sampling noise. Schulte's conclusion is not to stop measuring, it is to measure seven to eight times per prompt per day and aggregate over two to four weeks. Our position is that the sampling rate is the problem, not the metric. A score blended across two engines whose source sets agree about a quarter of the time is doing a great deal of averaging on your behalf, and no dashboard we read tells you how many runs went into it.

Six checks to run on any GEO tool before you sign

Six questions separate a vendor who has thought about sampling from one who has not, and each is answerable on a sales call. They are not the same size. The first one carries the statistical argument and is worth a minute of the call, and the last two are single questions you either get a straight answer to or you do not.

1. How many times do you run each prompt, per engine, per day?

Spend real time on this one, because every other number the vendor shows you inherits its reliability from the answer. At one run per prompt per day, the Schulte arithmetic puts a brand whose true detection rate is 50 percent anywhere from -22 percent to +122 percent in a nominal 95 percent interval, which is a long way of saying that a one run number tells you nothing about that brand on its own. Seven runs per prompt per day is where the published recommendation starts, and eight if you also care which sources got cited. Ask for runs per prompt rather than answers analyzed, because a headline like 2,250 answers can mean 25 prompts asked once a day for a month across three models. If they will not answer, you have your answer.

2. Show me the error bar.

Ask for the range instead of the point estimate. Nobody publishes one by default, so the test is whether they can produce one on request.

3. Logged in, logged out, or API?

Petra Labs, a vendor in this space, ran 900 trials across those surfaces on one day in February 2026, all three on GPT-5.2, and found the same brand swinging by up to 48.3 percentage points. One model, one day, and vendor research, so ask which surface their number describes.

4. What is the denominator?

Share of voice against all brand mentions is a different quantity from the share of prompts where you appear, and Conductor documents that gap in its own product.

5. Which model version, and what happens when it changes?

None of the eleven vendors we read publishes which model version its number describes, or what it does when that version changes.

6. How is a mention counted?

Scrunch's glossary says mentions appearing only inside cited URLs or page titles do not count toward Presence unless the AI names your brand.

A vendor who answers all six cleanly is worth buying from.

Search Console now gives you a free number to check the paid one against

Google Search Console has published generative AI performance reports since 2026-06-03, with dedicated views of impressions inside AI Overviews and AI Mode, broken out by page, country and date. Data begins 2026-05-18 with no historical backfill. Check your own property before planning around any of it, because the initial rollout went to a subset of UK sites under the UK Competition and Markets Authority requirement, and we could not confirm how widely the reports are available now.

Where they do appear, the limits are real: impressions only, no clicks and no query data, and Google surfaces alone. The number is still first party, free, and owned by the engine doing the answering.

Will this instability go away?

Possibly, which is a genuine limit on everything above. Thinking Machines Lab showed in September 2025 that nondeterminism, meaning the same input producing different outputs on different runs, traces to batch size varying with server load rather than to the order in which concurrent GPU threads finish. Their batch invariant kernels then made all 1,000 of their test completions identical. If the major providers adopt fixes of that kind, the tools get more stable with no change to vendor method.

None of this is an argument for canceling anything. AI visibility is worth measuring, and the scores on sale are not yet measurements, because a measurement carries a sample size and a range with it. Ask for the sampling method and the error bar, and buy on those answers.

All insights and articles

Reading about it is the easy part

Stop guessing. Start knowing your growth potential.

Get a free, no obligation audit of your paid search account, real math from your own query data, not a sales pitch.

The Optimal Path to Conversion

With precision bidding strategy, delivering the right Keyword/Match Type, to the right Ad Message, to the right Landing Page.

Wasteful Themes Pruned - Negatives
See how queryDNA works →