What this calculator answersGiven two reviewers’ original labels for the same preserved responses, how consistently did they apply one frozen codebook to brand mention, brand/domain citation, and active recommendation? The answer is three endpoint-specific records—not one composite visibility score.
Agreement is not correctnessHigh agreement does not prove that either label is true, that the codebook is valid, that the reviewers were independent, that the prompt sample represents a market, or that the underlying AI answer is accurate. Adjudication can create a final working label, but it cannot be recycled as independent agreement.
Evidence-preserving QA record

Measure consistency without hiding the rows

Enter the original independent labels. The page calculates each endpoint separately, excludes incomplete or non-binary pairs visibly, and keeps any adjudicated disposition outside the agreement arithmetic.

IncludedSame stable response ID, same preserved evidence, complete row identity, and yes/no labels from both reviewers for that endpoint.
Excluded—not convertedUnknown, ineligible, missing, incomplete, mismatched-evidence, or duplicate-ID rows stay visible and never become “no.”
Never inferredTruth, validity, reviewer independence, audit quality, a universal acceptance threshold, or correctness of the answer itself.
01 · Review identity and codebook

Name the people, scope, rules, and intended decision

Enter the audit or study name.
Enter the organization.
Enter the review owner.
Name reviewer A.
Name reviewer B.
Name the codebook.
Enter a codebook version.
Enter the codebook date.
Enter the window start.
Enter the window end.
Declare the engine and surface scope.
Declare the market and language.
State the purpose and boundary.
02 · Endpoint definitions and supplied policy

Freeze three distinct yes/no questions

Define brand mention.
Define brand/domain citation.
Define active recommendation.
A reviewer policy input, not a universal cutoff.
Evaluated only when kappa is defined.
Applied only to this endpoint.
03 · Reviewer-supported process controls

Affirm only what the review record supports

Unchecked controls do not hide the descriptive arithmetic. They withhold endpoint policy conclusions and remain visible in the memo.

04 · Same-response evidence rows

Enter both reviewers’ original endpoint labels

Evidence references must match exactly. Duplicate response IDs exclude every duplicate instance until the identity conflict is resolved.

RowResponse IDExact prompt or questionEngine / surfaceRun time or repetitionReviewer A evidence refReviewer B evidence refA mentionB mentionA citationB citationA recommendationB recommendationOptional adjudication dispositionRow noteAction
All form processing stays in this browser. This page has no account, upload, analytics payload for field values, or storage endpoint.
05 · Worked branches

Challenge the arithmetic and the interpretation

Examples are illustrative. Their supplied thresholds are fictional internal policies, not recommended cutoffs.

06 · Endpoint diagnostics and review queue

Reviewer agreement result

Complete the review identity and at least one evidence row, then calculate.

How to read the diagnostics

For each endpoint, the calculator builds a 2 × 2 table: both yes (a), reviewer A yes and B no (b), A no and B yes (c), and both no (d). Observed agreement is (a + d) / n. Expected agreement uses the two reviewers’ observed marginals. Cohen’s kappa is (observed − expected) / (1 − expected) when that denominator is defined.

Kappa can look poor even when raw agreement is high if nearly every eligible response falls into one category. That is why this page also reports each reviewer’s positive prevalence, positive agreement, negative agreement, prevalence index, bias index, eligible count, exclusions, and coverage. Kappa is never the only result.

Display warnings are not acceptance standardsThe interface calls attention to fewer than 20 eligible pairs, coverage below 80%, reviewer positive prevalence below 10% or above 90%, and a large bias index. These are transparent investigation prompts chosen for this interface—not universal statistical cutoffs and not substitutes for a sampling plan.

A supplied minimum kappa is evaluated only for its named endpoint, only when kappa is defined, and only when the independent-label, codebook, same-evidence, pre-adjudication, and exclusion controls are affirmed with no row-integrity conflict. The result means “meets this supplied policy” or “does not meet this supplied policy,” never accredited, valid, correct, or decision-grade.

Method sources and AI-visibility context

These references support the calculation and claim boundaries. They do not validate a codebook entered here, endorse this implementation, set a universal threshold, or certify an AI-visibility audit.

Frequently asked questions

How do you measure agreement between two AI visibility reviewers?

Have both reviewers independently classify the same preserved responses under the same frozen codebook. For each binary endpoint, report the full 2 × 2 table, observed and expected agreement, Cohen’s kappa when defined, reviewer prevalences, positive and negative agreement, exclusions, and the actual disagreement rows.

Why calculate mention, citation, and recommendation separately?

A brand can be named without its domain being cited, and it can be cited without being actively recommended. Combining those non-exclusive judgments into one score hides which rule the reviewers applied differently.

Should unknown or ineligible labels count as “no”?

No. Unknown means the reviewer cannot support a binary judgment from the preserved evidence; ineligible means the row falls outside the declared endpoint rule. Neither establishes absence. This page excludes the affected endpoint pair and reports the coverage loss.

What if both reviewers label every eligible row “no”?

Observed and negative agreement are 100%, but expected agreement is also 100%, so Cohen’s kappa is undefined. Report the raw table, single-category limitation, eligible count, coverage, and negative agreement instead of coercing kappa to zero or one.

Can I adjudicate disagreements?

Yes, but preserve the two original independent labels. Record any discussion, third-reviewer decision, or final working label in the disposition field. Do not overwrite the source columns or recalculate the adjudicated result as independent agreement.

Does high agreement prove the labels are correct?

No. Agreement measures repeatability between reviewers. It does not establish truth, accuracy, validity, absence of shared bias, reviewer independence, audit quality, or correctness of the underlying AI answer.

Does the calculator upload review evidence?

No. Calculation, memo generation, copying, and Markdown download happen locally in the browser. The page has no upload or storage endpoint.