The short answer For matched binary observations, the change is gains minus losses divided by eligible pairs. The exact McNemar test asks whether the two discordant directions—absence to presence and presence to absence—are unusually asymmetric under a 50/50 null. Stable rows affect the rates, but not the exact test statistic.
A p-value is not a causal verdict Even a small exact p-value does not prove that your intervention caused the movement. Platform drift, model changes, concurrent work, dependent rows, selection, missingness, and multiple testing still require design evidence and human review.

Keep the pair, the exclusion, and the claim separate

Matched analysis unitThe exact prompt, engine and surface, market, access condition, and planned repetition identity must refer to the same unit in both periods.
Missing is not absenceA failed, inaccessible, unreviewed, or otherwise ineligible observation is excluded under a declared rule and remains visible in coverage.
Inference has conditionsThe exact paired result is withheld until binary classification, matching, stable conditions, pair independence, and the exclusion rule are affirmed.

This tool fits an already-collected panel where the endpoint is binary—such as confirmed brand mention, confirmed citation, or confirmed recommendation—and the same planned units exist before and after. It does not fit unrelated prompt samples, blended visibility scores, rankings, counts, sentiment values, or two independent groups.

Browser-local matched-pair review

Validate the panel before testing the change

Record the method once, then add one row per exact prompt-and-repetition unit. The entered data never leaves this browser.

01 · Analysis identity

Name the endpoint, periods, and fixed conditions

Enter the study or decision name.
Enter the organization.
Name the reviewer.
Define the binary endpoint.
Enter the before label.
Enter the before start.
Enter the before end.
Enter the after label.
Enter the after start.
Enter the after end.
Enter the engine and surface.
Enter the market and language.
Enter the access and session condition.
Describe the comparison boundary.
State the eligibility and exclusion rule.
Enter a whole number from 1 to 999.
02 · Reviewer-supported assumptions

Affirm only what the evidence supports

Unchecked statements do not stop descriptive counts. They do withhold the exact paired p-value.

03 · Matched observation rows

Preserve every planned unit and both evidence references

Use present or absent only after reviewing a complete eligible response. Use missing for failed, unavailable, or unreviewed evidence and ineligible for an observation that violates the declared rule.

RowExact prompt or questionRepetition IDBeforeAfterBefore evidenceAfter evidenceNoteAction
Blank rows are ignored. Exact prompt + repetition ID is the case-sensitive row key.
04 · Calculate and export

Run the eligibility and paired-inference checks

Examples demonstrate arithmetic and refusal branches. They are not evidence about any platform or intervention.

Complete the analysis identity and at least one matched observation row.

Formulas and eligibility rules

  • Eligible matched pair = both periods have reviewed binary evidence—confirmed presence or confirmed absence—for the same declared analysis unit.
  • Stable absence = absent before and absent after. Gain = absent before and present after.
  • Loss = present before and absent after. Stable presence = present before and present after.
  • Before rate = (losses + stable presences) ÷ eligible pairs.
  • After rate = (gains + stable presences) ÷ eligible pairs.
  • Observed paired change = (gains − losses) ÷ eligible pairs, reported in percentage points.
  • Exact two-sided p-value = min(1, 2 × probability that a Binomial(discordant pairs, 0.5) count is at most min(gains, losses)). Only the gain and loss cells drive this conditional test.
Why an ordinary A/B calculator is the wrong default A two-proportion A/B test treats the before and after groups as independent. These rows deliberately preserve within-unit transitions. Discarding that pairing changes the variance model and loses the evidence needed to distinguish gains from losses.

What the result can and cannot support

A threshold-met result says the eligible discordant pairs were sufficiently asymmetric under the exact 50/50 null and the declared alpha. It does not say the intervention caused the movement, the effect will persist, the prompt panel represents a market, or the change is commercially important.

A threshold-not-met result does not prove equality or no effect. Small discordant counts, dependence, missing pairs, a weak instrument, and insufficient sample planning can all leave a real difference unresolved. Zero discordant pairs means this panel supplied no directional transitions; it is not an equivalence test.

When the endpoint was selected after looking at results or belongs to a family of tests, this page labels the exact p-value unadjusted and withholds a raw-alpha decision. Choose and justify any multiplicity method outside this calculator; the tool does not silently impose one.

Method sources and AI-visibility context

These sources support the statistical and measurement boundaries. They do not validate a specific panel entered here, certify independence, select an alpha, or replace statistical review.

Frequently asked questions

How do you test whether AI visibility changed before and after an intervention?

Match the same binary analysis units across periods, keep missing evidence out of the absence category, count absence-to-presence gains and presence-to-absence losses, and use a paired test such as the exact two-sided McNemar binomial test when its assumptions are supported.

Why not use a standard A/B significance calculator?

Most A/B calculators assume two independent groups. Before-and-after observations of the same prompt, engine, market, access condition, and planned repetition are paired, so the transition within each matched unit must be preserved.

Should a failed or missing AI answer count as absence?

No. A failed, inaccessible, missing, or unreviewed answer is missing measurement, not confirmed brand absence. Exclude the pair under a prespecified rule and disclose the resulting matched coverage.

What if there are no gains or losses?

The exact p-value is 1 because there are no discordant pairs, but the useful conclusion is narrower: this eligible panel contains no directional transitions. That is not proof that the periods or underlying probabilities are equal.

Does statistical significance prove the intervention worked?

No. A paired p-value addresses gain/loss asymmetry under stated assumptions. It does not rule out platform drift, model changes, history, concurrent work, dependence, selection bias, or multiple testing, and it does not prove durability or business importance.

Does this calculator upload observation data?

No. Calculation, memo generation, copying, and Markdown download happen in your browser. The page has no upload or storage endpoint.