The short answer A 42% mention rate, a 57/100 composite, a technical-readiness score, and a referral-traffic count are four different quantities. The fact that each dashboard calls its number “AI visibility” does not put them on one scale.
What this worksheet does—and does not do It records the supplied methods, classifies their documented relationship, exposes missing disclosures, and preserves both reported values exactly. It does not inspect private code, verify provider claims, determine which vendor is correct, normalize a proprietary score, certify decision fitness, or predict future mentions and citations.

Compare the method before comparing the number

Five fields gate direct comparison: the measured construct and unit, formula and denominator, prompt universe, platform or surface coverage, and reporting period or window. The remaining fields explain how much confidence a narrower directional comparison deserves.

Non-comparableA critical gate is different, has no overlap, is not disclosed, or cannot be reconstructed. Preserve both values separately.
Directional onlyCritical fields align at least partially, but material conditions or rules still differ. Discuss components and limitations, not a precise gap.
Comparable recordCritical and material fields are documented and aligned. Differences can be interpreted, but probabilistic outputs may still vary.
Browser-local reconciliation worksheet

Build the two-method record

Use vendor documentation, raw exports, report footnotes, and written answers. “Looks similar” is not evidence of alignment. If a material field is missing, label it not disclosed.

01 · Report identity

Preserve what each report actually says

02 · Field-level method comparison

Classify the documented relationship

Choose Aligned only when the supplied details support it. “Partial overlap” is not a softer version of aligned: it limits the output to directional interpretation.

Comparison fieldMethod A evidenceMethod B evidenceDocumented relationship
Measured construct and unitMention rate, citation rate, recommendation rate, share of voice, composite, readiness, rank, traffic, or another quantity.Critical gate
Formula, numerator, and denominatorInclude weights, normalization, bands, exclusions, and what happens when an engine or observation fails.Critical gate
Prompt universe, selection, and weightsPanel size, source, taxonomy, intent, branded mix, ownership, versions, and weighting.Critical gate
Platforms, surfaces, models, and access laneSeparate API, consumer app, search-grounded mode, AI Overview, AI Mode, and other surfaces.Critical gate
Reporting period, collection window, and aggregationA “July” label can mean one scan, a rolling week, a calendar month, or a stored answer index.Critical gate
Entity matching and competitive universeBrand aliases, product/company boundaries, false positives, competitors, and category membership.
Locale, language, account, session, and personalizationMarket and session conditions can change retrieval and generated answers.
Runs, repeats, cadence, and samplingOne answer, repeated draws, paraphrases, daily panels, and stored indexes have different variance.
Valid, null, failure, refusal, and retry rulesA timeout is not automatically a brand absence; hidden exclusions can move a rate.
Mention, citation, recommendation, prominence, and sentiment rulesA linked source is not the same event as an unlinked mention or recommendation.
Source and URL normalizationCanonical URLs, redirect targets, domains, source classes, and source-weight judgments.
Variability, uncertainty, and change thresholdA point estimate without its sample and variability can overstate a small difference.
Raw evidence, traceability, and exportCan a reviewer trace the headline value back to observations and retain the record?
Method version, change log, and rebaseline policyA vendor-side method change can move a score even when the measured brand does not change.
03 · Reviewer conclusion

Separate the observation from the decision

Start with the provider documentation. Every unassessed field stays visible as not disclosed.

All entered details stay in this browser tab. This page has no upload endpoint and does not persist the record.

Worked example: 30% and 57/100 are not a 27-point gap

Method A reports mention coverage: the brand appeared in 18 of 60 eligible answers, so its unit is a percentage of answers. Method B reports a 0–100 composite that mixes presence, prominence, sentiment, and source authority with undisclosed weights. Even if both reports show “July” and include ChatGPT, they measure different constructs and Method B cannot be reconstructed.

Correct reconciliation Preserve “30% mention coverage” and “57/100 composite” as separate observations. Label the headline values non-comparable, request Method B’s full formula and denominator, compare any shared raw mention-rate component separately if one exists, and avoid describing the result as a 27-point disagreement.

The loads this case into the worksheet so you can inspect the generated blockers and evidence requests.

Questions to ask each AI visibility provider

  1. What exact quantity does the headline number measure, and what is its unit?
  2. What are the numerator, denominator, weights, exclusions, and failed-run rules?
  3. Which prompts, versions, intents, markets, languages, platforms, products, surfaces, and access modes entered this report?
  4. How many observations were planned and completed per cell, and what variability or uncertainty accompanies the estimate?
  5. Can we export the full prompts, answers, citations, UTC timestamps, conditions, classifications, nulls, and failures?
  6. Which methodology version produced the report, and what changed since the prior baseline?

Use the separate AI visibility methodology disclosure generator when either provider first needs to document one method in full. Use this reconciliation worksheet only after there are two records to compare.

Frequently asked questions

Why do two AI visibility tools show different scores?

They may measure different events, prompt populations, AI surfaces, markets, sessions, repetition schemes, classifiers, denominators, formulas, weights, or time windows. Even aligned protocols can differ because generated answers and retrieval results vary across repeated runs.

Can I average two AI visibility scores?

Not unless the measured construct, unit, population, collection conditions, calculation, and material comparison fields align. Averaging a mention percentage with a composite or readiness score manufactures a third number with no documented interpretation.

Does “comparable record” mean the providers should agree exactly?

No. It means the methods are documented and materially aligned enough to interpret the difference. Sample size, uncertainty, model updates, retrieval, and ordinary nondeterminism still matter.

Which provider is right?

This worksheet does not choose. A transparent method can still be unsuitable for a particular decision, and two valid methods can answer different questions. Compare shared raw components under matched conditions if the decision requires a common basis.

Does the page upload or retain report details?

No. Classification, memo generation, copying, and download happen locally in the browser. Closing or refreshing the tab clears unsaved entries.