Compare the method before comparing the number
Five fields gate direct comparison: the measured construct and unit, formula and denominator, prompt universe, platform or surface coverage, and reporting period or window. The remaining fields explain how much confidence a narrower directional comparison deserves.
Build the two-method record
Use vendor documentation, raw exports, report footnotes, and written answers. “Looks similar” is not evidence of alignment. If a material field is missing, label it not disclosed.
Not generated
Values are preserved, not combined. The worksheet performs no averaging, normalization, conversion, gap percentage, ranking, or winner selection.
Critical comparability findings
Evidence requests
Field-level reconciliation record
| Field | Method A | Method B | Relationship |
|---|
Portable Markdown memo
Memo ready.
Worked example: 30% and 57/100 are not a 27-point gap
Method A reports mention coverage: the brand appeared in 18 of 60 eligible answers, so its unit is a percentage of answers. Method B reports a 0–100 composite that mixes presence, prominence, sentiment, and source authority with undisclosed weights. Even if both reports show “July” and include ChatGPT, they measure different constructs and Method B cannot be reconstructed.
The loads this case into the worksheet so you can inspect the generated blockers and evidence requests.
Questions to ask each AI visibility provider
- What exact quantity does the headline number measure, and what is its unit?
- What are the numerator, denominator, weights, exclusions, and failed-run rules?
- Which prompts, versions, intents, markets, languages, platforms, products, surfaces, and access modes entered this report?
- How many observations were planned and completed per cell, and what variability or uncertainty accompanies the estimate?
- Can we export the full prompts, answers, citations, UTC timestamps, conditions, classifications, nulls, and failures?
- Which methodology version produced the report, and what changed since the prior baseline?
Use the separate AI visibility methodology disclosure generator when either provider first needs to document one method in full. Use this reconciliation worksheet only after there are two records to compare.
Frequently asked questions
Why do two AI visibility tools show different scores?
They may measure different events, prompt populations, AI surfaces, markets, sessions, repetition schemes, classifiers, denominators, formulas, weights, or time windows. Even aligned protocols can differ because generated answers and retrieval results vary across repeated runs.
Can I average two AI visibility scores?
Not unless the measured construct, unit, population, collection conditions, calculation, and material comparison fields align. Averaging a mention percentage with a composite or readiness score manufactures a third number with no documented interpretation.
Does “comparable record” mean the providers should agree exactly?
No. It means the methods are documented and materially aligned enough to interpret the difference. Sample size, uncertainty, model updates, retrieval, and ordinary nondeterminism still matter.
Which provider is right?
This worksheet does not choose. A transparent method can still be unsuitable for a particular decision, and two valid methods can answer different questions. Compare shared raw components under matched conditions if the decision requires a common basis.
Does the page upload or retain report details?
No. Classification, memo generation, copying, and download happen locally in the browser. Closing or refreshing the tab clears unsaved entries.