What this calculator answersFor the supplied, scoped reference set, how did one reviewer’s labels perform on brand mention, brand/domain citation, and active recommendation? The answer is three endpoint-specific confusion records—not one composite visibility or reviewer score.
A reference standard is not absolute truthEven a locked, adjudicated reference can contain uncertainty, incorporate shared bias, or fail to represent a different prompt population, engine, market, language, or evidence window. This page reports performance against the supplied reference set and preserves corrections separately; it does not certify correctness or reviewer competence.
Reference-set QA record

Measure annotation performance without hiding the errors

Enter the locked reference labels and the reviewer’s original labels. The page calculates each endpoint separately, excludes incomplete or non-binary pairs visibly, and keeps corrections outside the original confusion arithmetic.

IncludedSame stable response ID, same preserved evidence, complete row identity, and yes/no labels from the reference set and reviewer for that endpoint.
Excluded—not convertedUnknown, ineligible, missing, incomplete, mismatched-evidence, or duplicate-ID rows stay visible and never become “no.”
Never inferredAbsolute truth, reference correctness, sample representativeness, reviewer competence, audit quality, or a universal acceptance threshold.
01 · Validation identity and reference provenance

Name the reviewer, reference set, scope, and intended decision

Enter the audit or study name.
Enter the organization.
Enter the review owner.
Name the reviewer or classifier.
Name the codebook.
Enter a codebook version.
Enter the codebook date.
Name the reference set.
Enter the reference-set version.
Enter the freeze date.
Enter the window start.
Enter the window end.
Declare the engine and surface scope.
Declare the market and language.
Document reference-set construction and adjudication.
Document the population, sampling, and limits.
State the purpose and boundary.
02 · Endpoint definitions and optional supplied policy

Freeze three distinct yes/no questions

Define brand mention.
Define brand/domain citation.
Define active recommendation.
Pair with a named metric; not a universal cutoff.
Evaluated only when the named metric is defined.
Applied only to this endpoint.
03 · Reference-process controls

Affirm only what the validation record supports

Unchecked controls do not hide the descriptive arithmetic. They withhold optional policy conclusions and remain visible in the memo.

04 · Same-response evidence rows

Enter reference and reviewer endpoint labels

Evidence references must match exactly. Duplicate response IDs exclude every duplicate instance until the identity conflict is resolved.

RowResponse IDExact prompt or questionEngine / surfaceRun time or repetitionReference evidence refReviewer evidence refReference mentionReviewer mentionReference citationReviewer citationReference recommendationReviewer recommendationOptional correction or dispositionRow noteAction
All form processing stays in this browser. This page has no account, upload, analytics payload for field values, or storage endpoint.
05 · Worked branches

Challenge the arithmetic and the interpretation

Examples are illustrative. Their supplied thresholds are fictional internal policies, not recommended cutoffs.

06 · Endpoint diagnostics and error review

Reference-set accuracy result

Complete the validation identity and at least one evidence row, then calculate.

How to read the diagnostics

For each endpoint, the calculator builds a 2 × 2 confusion table with the reviewer as the evaluated classifier and the supplied reference set as the comparison standard. A true positive is reviewer yes/reference yes; a false positive is reviewer yes/reference no; a false negative is reviewer no/reference yes; and a true negative is reviewer no/reference no.

Accuracy is (TP + TN) / n. Sensitivity is TP / (TP + FN); specificity is TN / (TN + FP); precision is TP / (TP + FP); NPV is TN / (TN + FN); balanced accuracy averages defined sensitivity and specificity; F1 is 2TP / (2TP + FP + FN); and MCC is the signed four-cell correlation when its product denominator is nonzero. Every unavailable denominator stays visibly undefined.

Accuracy can look excellent when the reference set is imbalanced even though a reviewer misses every positive. Keep reference prevalence, reviewer-positive rate, sensitivity, specificity, predictive values, balanced accuracy, F1, MCC, the complete confusion table, eligible count, exclusions, and coverage together.

Display warnings are not acceptance standardsThe interface calls attention to fewer than 20 eligible pairs, coverage below 80%, reference prevalence below 10% or above 90%, and a sensitivity–specificity gap of at least 30 percentage points. These are transparent investigation prompts chosen for this interface—not universal cutoffs and not substitutes for a sampling plan.

An optional supplied minimum is evaluated only for its named endpoint and named metric, only when that metric is defined, and only when reference locking, reference-label blinding, same-evidence comparison, original-label preservation, and exclusion controls are affirmed with no row-integrity conflict. The result means “meets this supplied policy” or “does not meet this supplied policy,” never correct, competent, representative, certified, accredited, or decision-grade.

Method sources and AI-visibility context

These references support the calculation and claim boundaries. They do not validate the supplied reference set or codebook, endorse this implementation, set a universal threshold, or certify an AI-visibility audit.

Frequently asked questions

How do you validate AI visibility annotations against a reference set?

Compare the reviewer and reference labels for the same preserved response. For each binary endpoint, report the full confusion table, coverage, reference prevalence, reviewer-positive rate, accuracy, sensitivity, specificity, precision, NPV, balanced accuracy, F1, MCC when defined, exclusions, and the actual false-positive and false-negative rows.

Why calculate mention, citation, and recommendation separately?

A brand can be named without its domain being cited, and it can be cited without being actively recommended. Combining those non-exclusive judgments into one score hides which rule the reviewers applied differently.

Should unknown or ineligible labels count as “no”?

No. Unknown means the reviewer cannot support a binary judgment from the preserved evidence; ineligible means the row falls outside the declared endpoint rule. Neither establishes absence. This page excludes the affected endpoint pair and reports the coverage loss.

What if the reviewer labels every eligible row “no”?

Accuracy can still look high when positive reference rows are rare. Sensitivity will reveal missed positives, while precision and MCC may be undefined depending on the four cells. Keep the complete table and defined metrics; never coerce an undefined denominator to zero or one.

Can I correct an annotation after review?

Yes, but preserve both original columns. Record any correction, adjudication, retraining note, or final working label in the disposition field. Do not overwrite the source labels or quietly recompute the original validation result.

Does high performance prove the labels are correct?

No. The result describes performance against this supplied reference set. It does not establish absolute truth, validity, absence of shared bias, reference correctness, reviewer competence, population representativeness, audit quality, or generalization beyond the declared scope.

Does the calculator upload review evidence?

No. Calculation, memo generation, copying, and Markdown download happen locally in the browser. The page has no upload or storage endpoint.