Keep the pair, the exclusion, and the claim separate
This tool fits an already-collected panel where the endpoint is binary—such as confirmed brand mention, confirmed citation, or confirmed recommendation—and the same planned units exist before and after. It does not fit unrelated prompt samples, blended visibility scores, rankings, counts, sentiment values, or two independent groups.
Validate the panel before testing the change
Record the method once, then add one row per exact prompt-and-repetition unit. The entered data never leaves this browser.
Formulas and eligibility rules
- Eligible matched pair = both periods have reviewed binary evidence—confirmed presence or confirmed absence—for the same declared analysis unit.
- Stable absence = absent before and absent after. Gain = absent before and present after.
- Loss = present before and absent after. Stable presence = present before and present after.
- Before rate = (losses + stable presences) ÷ eligible pairs.
- After rate = (gains + stable presences) ÷ eligible pairs.
- Observed paired change = (gains − losses) ÷ eligible pairs, reported in percentage points.
- Exact two-sided p-value = min(1, 2 × probability that a Binomial(discordant pairs, 0.5) count is at most min(gains, losses)). Only the gain and loss cells drive this conditional test.
What the result can and cannot support
A threshold-met result says the eligible discordant pairs were sufficiently asymmetric under the exact 50/50 null and the declared alpha. It does not say the intervention caused the movement, the effect will persist, the prompt panel represents a market, or the change is commercially important.
A threshold-not-met result does not prove equality or no effect. Small discordant counts, dependence, missing pairs, a weak instrument, and insufficient sample planning can all leave a real difference unresolved. Zero discordant pairs means this panel supplied no directional transitions; it is not an equivalence test.
When the endpoint was selected after looking at results or belongs to a family of tests, this page labels the exact p-value unadjusted and withholds a raw-alpha decision. Choose and justify any multiplicity method outside this calculator; the tool does not silently impose one.
Method sources and AI-visibility context
- NIST: McNemar Test defines the paired binary 2 × 2 problem, its gain/loss comparison, and the mutually independent-pairs assumption.
- statsmodels: McNemar test of homogeneity documents the exact binomial option and identifies the exact statistic as the smaller discordant count.
- Quantifying Uncertainty in AI Visibility demonstrates why stochastic answer-engine measurement needs explicit uncertainty, sample design, and narrower claims.
- MaxAEO: controlled AI-visibility experiments recommends fixed prompt cohorts, repeated sampling, a logged intervention, control comparison, and uncertainty rather than a one-off screenshot.
These sources support the statistical and measurement boundaries. They do not validate a specific panel entered here, certify independence, select an alpha, or replace statistical review.
Frequently asked questions
How do you test whether AI visibility changed before and after an intervention?
Match the same binary analysis units across periods, keep missing evidence out of the absence category, count absence-to-presence gains and presence-to-absence losses, and use a paired test such as the exact two-sided McNemar binomial test when its assumptions are supported.
Why not use a standard A/B significance calculator?
Most A/B calculators assume two independent groups. Before-and-after observations of the same prompt, engine, market, access condition, and planned repetition are paired, so the transition within each matched unit must be preserved.
Should a failed or missing AI answer count as absence?
No. A failed, inaccessible, missing, or unreviewed answer is missing measurement, not confirmed brand absence. Exclude the pair under a prespecified rule and disclose the resulting matched coverage.
What if there are no gains or losses?
The exact p-value is 1 because there are no discordant pairs, but the useful conclusion is narrower: this eligible panel contains no directional transitions. That is not proof that the periods or underlying probabilities are equal.
Does statistical significance prove the intervention worked?
No. A paired p-value addresses gain/loss asymmetry under stated assumptions. It does not rule out platform drift, model changes, history, concurrent work, dependence, selection bias, or multiple testing, and it does not prove durability or business importance.
Does this calculator upload observation data?
No. Calculation, memo generation, copying, and Markdown download happen in your browser. The page has no upload or storage endpoint.