The repeat-count answer Three observations per question and surface—two immediate identical runs and one delayed identical run—are enough to reveal some short-window instability. They are not enough to estimate a universal visibility probability. Five fixed runs can be a practical directional planning default, but “3 of 5 observed” is the claim; “60% true visibility” is not. Formal inference needs a separately designed sample and uncertainty analysis.
Free local tool · no signup

Plan repeat runs and compare two ordered result lists

First calculate the exact workload before collecting anything. Then paste two completed native-search URL lists to test whether their order is identical and where membership or positions changed. All calculations stay in this browser; nothing is uploaded or saved.

01 · Plan the observations

How many ChatGPT searches will this audit require?

Keep prompt breadth, surfaces, immediate repetition, and delayed checks visible. Combining them into one unexplained “sample size” hides what you actually observed.

Distinct frozen questions
Report each separately
Identical back-to-back runs
Same wording, later
Inputs are capped only to prevent accidental browser-sized numbers.
30Total completed observations
3Per question and surface
20Immediate observations
10Delayed observations
Minimum instability screen

Run two identical checks back-to-back for each question and surface, then repeat once later. This plan can expose instability; it does not estimate a universal probability.

02 · Compare the evidence

Did two ChatGPT web-search result lists actually match?

Paste one complete URL per line in displayed order. Blank lines are ignored. Exact-order comparison retains duplicates and query strings; Jaccard overlap uses unique exact URL strings and ignores order.

Enter both completed URL lists to calculate the comparison.

Interpret the numbers narrowly Exact full order asks whether every displayed URL occupied the same position. Jaccard asks only how much unique membership overlapped: shared URLs divided by the union. Neither metric proves that an answer, citation, recommendation, or brand position was stable unless that is what the pasted lists represent.

Collected three or more comparable runs? Use the free AI citation stability calculator to calculate every pairwise overlap and each URL's persistence across the complete panel.

Measure four different outcomes

Teams often compress every appearance into an “AI visibility score.” That hides the useful distinction. Record these outcomes separately for every run:

StateWhat happenedWhat it proves
AbsentYour company and domain do not appear.No visibility in that completed response.
MentionedThe answer names your company without presenting it as a suitable choice.Entity recognition, not recommendation.
RecommendedThe answer offers your SaaS as a relevant option for the buyer's task.Recommendation in that response, not a durable rank.
LinkedA result block or citation links to a canonical page on your domain.That exact page entered the returned source set.

A company can be mentioned but not recommended, recommended but not linked, or linked as background without being the recommended product. Preserve the evidence instead of upgrading one state into another.

A repeatable six-step ChatGPT visibility audit

1. Declare the buyer-question set before searching

Write 10–20 questions a real buyer could reasonably ask. Use language from sales calls, support tickets, and product comparisons. Include a mix of category, problem, audience, constraint, and comparison intent. Keep most questions non-branded; branded questions diagnose entity clarity but do not measure discovery.

2. Freeze the conditions you can control

Record the date and time, product or interface, model when visible, whether web search is enabled, account/session context, locale, and the exact question. You cannot eliminate every source of variation. You can make later runs comparable and disclose what remains unknown.

3. Repeat identical requests

Use two identical back-to-back runs plus one identical delayed run as a minimum instability screen, then declare a larger fixed budget when you need directional monitoring. Do not rewrite a query after an unfavorable answer. When measuring native web-search placement, use the identical search request and keep the returned result blocks in displayed order. The repeat planner makes the resulting workload explicit.

4. Save the complete response

Log the full answer or ordered search-result set, not only whether your domain appeared. Record your canonical URL and position when linked, all named competitors, cited sources, and the four-state classification above. Negative runs belong in the dataset.

5. Inspect why the current sources answer the question

For each useful winner, inspect its title and H1, opening answer, audience, evidence, dates, internal context, publisher clarity, and page type. This turns “we are absent” into a page or entity hypothesis you can actually test.

6. Repeat on a fixed cadence

Re-run the same set weekly for directional monitoring and after a meaningful site change. Compare rates and source membership, not just one rank. Keep the questions frozen for the measurement window; add new questions as a separately labeled cohort.

Copyable audit worksheet

A spreadsheet is enough for a first audit. Use one row per completed response, not one row per question.

FieldExample
Question IDQ07
Exact questionWhat project-management tool works for a five-person remote agency?
Declared intentCommercial recommendation · small team
UTC timestamp2026-08-23T21:08:03Z
EnvironmentChatGPT · web search on · model shown in UI
Run number2 of 3
StateAbsent / mentioned / recommended / linked
Canonical URL and positionBlank unless actually linked
Named competitorsPreserve displayed order
Linked sourcesPreserve every URL in displayed order
NotesAnswer emphasized integrations over team size
Prefer a working tracker? Use the free local-first LLM citation tracker to add completed responses, calculate separate mention, recommendation and linked-citation rates, and export the full log as CSV. No signup or data upload.

Starter questions for an early-stage B2B SaaS audit

Replace the bracketed category and constraints with facts about your market. These are patterns, not magic wording:

  1. What is the best [product category] for an early-stage B2B SaaS?
  2. Which [product category] works for a five-person team with no dedicated administrator?
  3. What [product category] should a startup choose before product-market fit?
  4. Which tools solve [specific buyer problem] without [real constraint]?
  5. What are the best alternatives to [named competitor] for [audience]?
  6. How should a [role] solve [problem] on a limited budget?
  7. Which [category] products support [required integration or compliance need]?
  8. What should I compare when choosing a [category] for [use case]?
  9. Which company offers a managed [service] for [audience]?
  10. How can I verify whether an AI assistant recommends my company?

Do not add your brand to a discovery question. Do not invent a unique phrase only your site uses. A qualifying query should remain useful if your company does not appear.

What 110 native-search calls taught us about measurement

On August 23, 2026, Quoted First ran an empirical study of native web-search result blocks: 110 calls covering 85 distinct query strings, 15 immediate identical repeats, 10 delayed repeats, 40 unrelated-category searches, and 20 substantive page inspections.

10 / 15

Immediate repeat pairs returned the exact same full URL order.

Observed, single-session study
5 / 6

Full-sequence delayed comparisons were exact after roughly 16 minutes.

One retained 9 of 10 leaders
0 / 25

Relevant branded and category queries returned Quoted First before implementation.

Null result retained

Three behaviors mattered most. First, adding a real audience, budget, stage, or problem often changed several leading sources. Second, focused pages with a direct answer could outrank larger publishers. Third, exact wording was powerful but noisy: literal phrase matches sometimes returned the wrong entity or product type.

This supports a practical audit rule: freeze natural query families, repeat them, and inspect the pages that satisfy the intent. It does not reveal a proprietary provider, index, crawler, or ranking architecture.

Limitations and honest interpretation

DIY audit or managed audit?

Run the worksheet yourself if you have a bounded question set and can preserve complete responses consistently. A managed audit is useful when you need competitor classification, page-level diagnosis, a prioritized implementation plan, and repeated measurement owned by one team.

Quoted First's audit is a fixed-scope baseline and implementation roadmap. It separates evidence from inference and includes null results. See the complete service scope, deliverables and price, or email hello@quotedfirst.com.

Frequently asked questions

How do I measure whether ChatGPT recommends my SaaS?

Declare a fixed set of natural, non-branded buyer questions before testing. Run each question repeatedly under the same conditions, save the complete answers and linked sources with timestamps, classify whether your SaaS was absent, mentioned, recommended, or linked, and repeat the same set on a schedule.

How many times should I repeat a ChatGPT search for an AI visibility audit?

There is no universal run count. Two identical back-to-back runs plus one identical delayed run form a minimum instability screen, not a probability estimate. For directional monitoring, choose and disclose a larger fixed run budget, preserve every null result, and report observed counts with their denominator.

What is the difference between a mention and a citation?

A mention names the company in answer text. A recommendation presents it as a suitable option. A citation or linked result points to a specific page. Record these states separately.

Method status Version 1.1, updated August 24, 2026. This revision separates a minimum instability screen from directional run planning and adds the local repeat planner and ordered-list comparator. Publication dates are not refreshed cosmetically.