Measure consistency without hiding the rows
Enter the original independent labels. The page calculates each endpoint separately, excludes incomplete or non-binary pairs visibly, and keeps any adjudicated disposition outside the agreement arithmetic.
How to read the diagnostics
For each endpoint, the calculator builds a 2 × 2 table: both yes (a), reviewer A yes and B no (b), A no and B yes (c), and both no (d). Observed agreement is (a + d) / n. Expected agreement uses the two reviewers’ observed marginals. Cohen’s kappa is (observed − expected) / (1 − expected) when that denominator is defined.
Kappa can look poor even when raw agreement is high if nearly every eligible response falls into one category. That is why this page also reports each reviewer’s positive prevalence, positive agreement, negative agreement, prevalence index, bias index, eligible count, exclusions, and coverage. Kappa is never the only result.
A supplied minimum kappa is evaluated only for its named endpoint, only when kappa is defined, and only when the independent-label, codebook, same-evidence, pre-adjudication, and exclusion controls are affirmed with no row-integrity conflict. The result means “meets this supplied policy” or “does not meet this supplied policy,” never accredited, valid, correct, or decision-grade.
Method sources and AI-visibility context
- SAS: More than Just the Kappa Coefficient provides the binary agreement, positive/negative agreement, prevalence-index, and bias-index formulas used here.
- PubMed: Bias, prevalence and kappa explains why kappa is affected by marginal prevalence and why prevalence and bias diagnostics should accompany it.
- CDC/NIOSH: Accuracy and Precision distinguishes consistency from accuracy and defines independent classification as rating without another reader or knowledge of that reader’s classification.
- FDA: Statistical Guidance on Reporting Agreement warns that agreement is not correctness: two methods can agree and both be wrong.
- IAB: Measuring Visibility in the AI Era separates mention, citation, recommendation, and accuracy concepts and emphasizes disclosed classification, validation, stability, and reproducibility.
These references support the calculation and claim boundaries. They do not validate a codebook entered here, endorse this implementation, set a universal threshold, or certify an AI-visibility audit.
Frequently asked questions
How do you measure agreement between two AI visibility reviewers?
Have both reviewers independently classify the same preserved responses under the same frozen codebook. For each binary endpoint, report the full 2 × 2 table, observed and expected agreement, Cohen’s kappa when defined, reviewer prevalences, positive and negative agreement, exclusions, and the actual disagreement rows.
Why calculate mention, citation, and recommendation separately?
A brand can be named without its domain being cited, and it can be cited without being actively recommended. Combining those non-exclusive judgments into one score hides which rule the reviewers applied differently.
Should unknown or ineligible labels count as “no”?
No. Unknown means the reviewer cannot support a binary judgment from the preserved evidence; ineligible means the row falls outside the declared endpoint rule. Neither establishes absence. This page excludes the affected endpoint pair and reports the coverage loss.
What if both reviewers label every eligible row “no”?
Observed and negative agreement are 100%, but expected agreement is also 100%, so Cohen’s kappa is undefined. Report the raw table, single-category limitation, eligible count, coverage, and negative agreement instead of coercing kappa to zero or one.
Can I adjudicate disagreements?
Yes, but preserve the two original independent labels. Record any discussion, third-reviewer decision, or final working label in the disposition field. Do not overwrite the source columns or recalculate the adjudicated result as independent agreement.
Does high agreement prove the labels are correct?
No. Agreement measures repeatability between reviewers. It does not establish truth, accuracy, validity, absence of shared bias, reviewer independence, audit quality, or correctness of the underlying AI answer.
Does the calculator upload review evidence?
No. Calculation, memo generation, copying, and Markdown download happen locally in the browser. The page has no upload or storage endpoint.