The short answerAn aggregate can rise because visibility improved within prompt strata, because more observations came from high-rate strata, or both. This calculator preserves all three views: crude movement, the exact within-rate/composition bridge, and two rates standardized to the same supplied mix.
Adjustment is not repairA common mix makes one declared categorical distribution comparable. It cannot rescue changed prompt identities inside a stratum, missing-not-at-random answers, different model or access conditions, dependence, confounding, or an invalid time series.

Preserve the aggregate, declare the standard, narrow the diagnosis

Expose both compositionsKeep planned and usable shares separate. Differential failures can change the effective mix even when the planned panel stays fixed.
Use one common standardApply the same reviewer-declared weights to both periods. Record the source, owner, version, timing, and rationale.
Reconcile exactlyWithin-rate movement plus composition movement must equal the crude percentage-point difference to numerical tolerance.

Choose one categorical stratification dimension whose levels are mutually exclusive and exhaustive—for example prompt intent, engine, market, buyer stage, or branded versus unbranded. Do not combine several changing dimensions into ambiguous row labels and call the remainder controlled.

Browser-local composition record

Diagnose whether the rate moved or the mix moved

Name the aggregate and standard, affirm only supported comparison gates, then add one row per stratum with planned, usable, present, and evidence fields for both periods.

01 · Analysis identity

Name one aggregate, endpoint, and stratification

Enter the study or decision name.
Enter the organization.
Name the reviewer.
Define one binary endpoint.
Name one stratification dimension.
Name the aggregate.
Enter the before label.
Enter the before start.
Enter the before end.
Enter the after label.
Enter the after start.
Enter the after end.
Enter the engine and surface scope.
Enter the market and language.
State the decision or claim.
02 · Common standard provenance

Document whose mix both periods will use

The page supplies no default weights and does not normalize a total that misses 100%.

Name the standard source.
Enter the source version and date.
Name the standard owner.
Document the standard rationale.
Document the missing-evidence rule.
03 · Reviewer-supported gates

Affirm comparability without concealing arithmetic

Unchecked gates leave the descriptive bridge visible but label the interpretation diagnostic only.

04 · Stratum evidence rows

Keep planned, usable, present, and standard counts separate

Counts are whole observations and must satisfy 0 ≤ present ≤ usable ≤ expected, with expected and usable above zero in both periods. Duplicate keys exclude every duplicate instance.

Supplied before-and-after stratum evidence and standard weights
#Stratum keyStratum labelBefore evidenceBefore expectedBefore usableBefore presentAfter evidenceAfter expectedAfter usableAfter presentStandard weight %Action
Maximum 100 supplied strata. Weight zero is explicit; missing weight is incomplete.
05 · Calculate or challenge the diagnosis

Run the supplied record or load a worked branch

Complete the identity, standard record, and at least one valid stratum row.

Formula and reconciliation

For stratum i, the observed binary visibility rates are p₀ᵢ = present₀ᵢ ÷ usable₀ᵢ and p₁ᵢ = present₁ᵢ ÷ usable₁ᵢ. The crude usable-observation weights are w₀ᵢ = usable₀ᵢ ÷ Σusable₀ and w₁ᵢ = usable₁ᵢ ÷ Σusable₁.

Symmetric two-component identityCrude change = Σ[(p₁ᵢ − p₀ᵢ)(w₀ᵢ + w₁ᵢ)/2] + Σ[(w₁ᵢ − w₀ᵢ)(p₀ᵢ + p₁ᵢ)/2]. The first sum is within-rate movement; the second is composition movement.

For supplied standard weights qᵢ that total exactly 100%, the directly standardized rates are Σ(qᵢp₀ᵢ) and Σ(qᵢp₁ᵢ). The page never substitutes either period's observed mix, picks a favorable standard, or silently rescales incomplete weights.

What the result can and cannot support

A crude/standardized direction conflict means the historical aggregate and the same-mix comparison point in opposite directions. A Simpson-style warning means the crude direction is opposite every nonzero stratum-specific rate direction. Both are prompts to inspect the evidence and reporting design—not automatic proof that one result is true and another false.

Standardization balances only the declared categorical mix. It does not repair changed prompt wording or identities inside strata, missing-not-at-random observations, differential classification error, platform or model drift, dependence, confounding, selection, sampling error, or an undisclosed series break. The decomposition is an arithmetic identity, not a causal attribution model.

The tool provides no significance test, interval, sample-size rule, universal standard, preferred weights, causal effect, acceptable-change threshold, certification, or instruction to replace the preserved crude series. Consequential interpretation needs the underlying evidence and qualified review.

Method sources

Those sources establish general arithmetic. They do not endorse this independent page, select an AI-visibility stratification or standard, validate supplied observations, or transfer public-health sampling assumptions to AI answer measurements.

Frequently asked questions

How can prompt mix change an AI visibility rate?

An aggregate is a weighted average. If high-visibility strata become a larger share of usable observations, the aggregate can rise without any within-stratum gain; the reverse can produce an apparent decline.

How do you adjust AI visibility for a changed prompt mix?

Apply one justified common set of stratum weights to both periods' stratum-specific rates, then compare the two standardized results alongside the crude rates, coverage, and evidence.

What is a Simpson-style reversal?

On this page it means the crude aggregate moves in one direction while every nonzero stratum-specific rate change moves in the other. It flags a composition-sensitive aggregate; it does not by itself establish a cause.

Should failed or unavailable answers count as absence?

No. Expected, usable, and present counts remain separate. Failed, unavailable, refused, unreviewed, and ineligible outcomes do not become confirmed absence.

Does this calculator upload observation data?

No. Validation, arithmetic, memo generation, copying, and Markdown download happen locally in the browser.