AI search visibility is a set of descriptive rates measured across a sampled set of AI answers. It counts what those answers show, which sources they cite, and which visits you can observe. It is not demand, reach, or market share. Keep visibility separate from qualified inquiries and business outcomes, and put a defined denominator beside every rate.
This is a repeatable measurement protocol, not a benchmark. It contains no original dataset, no client results, and no universal visibility score. Running the same protocol can return different answers, so the method preserves evidence rather than promising reproducible numbers.
Jump to the measurement worksheet, see the research overview, or start with the definitions below.
What AI search visibility measures, and what it does not
AI search visibility measures how often a sampled set of AI answers names you, cites you, or recommends you. It does not measure demand.
A visibility rate describes the answers you sampled and nothing else. It is not market share. It is not audience reach. It is not query volume. It is not the probability that an arbitrary buyer will see you.
Three confusions distort most AI visibility reporting. A mention is not a recommendation. A citation is not proof the cited page supports the claim. Crawler access is not evidence that any answer cited you. Each of these is a separate measurement, and collapsing them creates numbers that look precise and mean little.
What this does not cover. Visibility says nothing about whether an inquiry qualified or a matter closed. Those are downstream events with their own definitions and owners, covered later in this guide.
The universal-score myth. Many tools report a single AI visibility score or share-of-voice figure. The number persists because one figure is easy to sell and easy to chart. In reality these are descriptive rates over a convenience sample, and averaging different platforms into one score hides more than it shows. Report each platform and each rate separately, with its denominator, and do not present a blended score as performance.
Measure six signals separately
Six signals carry different meaning: mention, citation, recommendation, referral, qualified inquiry, and business outcome. Define and report each on its own.
Mention. An answer names the target person, product, or organization. Record whether the description is accurate. A name match alone is not a positive assessment.
Citation. An answer attaches a source link. Record the destination and the claim it appears to support. A citation to a third-party page differs from a citation to your own site, and a link is not proof of support.
Recommendation. The answer presents the target as a candidate for the user’s stated need. Inclusion in a list, a disclaimer, or a negative example is not a recommendation. Keep the passage and flag ambiguous cases for review.
Referral. An observable website visit tied to an identifiable source under your analytics rules. This is measured from your website records, not inferred from a citation in a sampled answer.
Qualified inquiry. A unique inquiry that meets your documented qualification criteria. Intake or sales decides qualification. A form fill, call, signup, or demo request does not establish it.
Business outcome. A separately recorded downstream event such as a retained matter, a sales opportunity, or a closed sale. State which outcome you mean and who validates it. Do not treat revenue and pipeline as interchangeable.
Put a denominator beside every rate
Every rate needs a numerator, a denominator, a time period, and a statement of what the sample excludes. Show all four or do not report the rate.
Assessable-answer yield. Assessable answers divided by all planned attempts. Separately report no-answer responses, refusals, technical failures, incomplete captures, and attempts not run, so missing work cannot disappear.
Sampled mention rate. Answers with a confirmed mention divided by answers assessable for a mention in the same platform, mode, wave, and prompt segment. Count each answer once.
Sampled own-site citation rate. Answers citing at least one URL you own divided by answers assessable for citation presence in that segment. Count third-party citations about you as a separate rate. Exclude unrelated sources from the numerator.
Sampled recommendation rate. Answers with a confirmed recommendation divided by answers assessable for recommendation status in the relevant-intent segment. Define which intents are eligible before you collect.
Answer-feature appearance. Captured searches showing the feature divided by all successfully captured searches in that segment. For AI Overviews, report this rate and any visibility rate that assumes an Overview appeared. They answer different questions.
Show the numerator and denominator next to every percentage. A zero denominator produces “not measurable,” not zero percent. State missing and uncertain counts. Never rerun only the unfavorable answers.
The method in one line
Sample AI answers against a fixed prompt panel, code six signals with defined denominators, and keep visibility, referrals, qualified inquiries, and outcomes as separate measures. The rest of this guide covers how to sample, how to capture, what platform documentation actually supports, and how to connect referrals to qualified demand without inventing causation.
Choose the decision and freeze the sample before testing
Start with the decision the measurement should inform, then choose the platforms and prompts that fit it. Prompts you write are test inputs, not evidence of real query volume.
- Build a prompt register from documented buyer questions, intake and sales themes, and search research. Record each prompt’s source.
- Give each prompt an ID, audience, topic, intent stage, language, and branded or unbranded status. Keep discovery, evaluation, and decision questions distinguishable.
- Freeze the prompt list, target aliases, dates, platforms, modes, and repeats per prompt before a wave. There is no universal minimum that makes a convenience sample representative, so set repeats by capacity and report the number.
- Keep a fixed panel for comparisons. Label exploratory prompts separately. Record every addition and removal.
- Use one session protocol: fresh conversation where available, recorded account state, locale, device, and model label when visible. Record unknown settings as unknown.
Capture one record per attempt
An attempt is one prompt run in one platform, mode, and session at a recorded time. Give every planned attempt a unique ID and preserve the evidence.
Store the exact prompt, the answer, the visible sources, the timestamp with time zone, and a screenshot or permitted export. Keep follow-up turns separate from first answers.
Record a response status for each attempt: assessable answer, no answer feature, refusal, technical failure, incomplete capture, or not run. A fully captured answer without citations is an observed zero for citation presence. A truncated or inaccessible answer is missing evidence, not a zero. For Google Search, record whether an AI Overview appeared, and do not discard successful searches that returned none.
For each citation, keep a source row linked to the attempt ID: displayed link, resolved destination, source owner and type, related passage, and a support assessment of supports, partly supports, does not support, or not checked. Do not describe unchecked citations as verified. Have a second reviewer check ambiguous mentions and recommendations where possible, and keep unresolved cases out of the rate while disclosing how many you excluded.
What platform documentation says, and what it does not
Documentation describes a platform feature at a review date. Observation shows what happened in a captured attempt. Recommendation is an analyst’s proposed action. Label these separately in every report.
The summaries below were checked on September 9, 2026. They are documentation, not test findings about any business, and platform behavior changes.
Google. Google’s generative AI performance reports in Search Console show impressions for your pages inside AI Overviews, AI Mode, and generative features in Discover, broken out by pages, countries, dates, and device for Search. As of August 31, 2026 the reports reached all sites worldwide, and they exclude click data, per Google’s announcement of the reports. Use the dedicated view, record its actual fields and filters, and do not read overall Web totals as isolated AI performance. This establishes impression reporting only. It does not justify a separate click, query, recommendation, or lead measure, and the AI data also folds into overall Performance, so do not add the two together. Check the Search Console data anomalies log before comparing periods.
ChatGPT. OpenAI documents OAI-SearchBot for surfacing sites in ChatGPT search and GPTBot for training, controlled independently in robots.txt. Crawler access is not evidence that any answer cited or recommended you. Record the observed ChatGPT experience and whether search is evidenced in the response, and do not infer it from the product name. See OpenAI’s crawler documentation.
Claude. Anthropic’s help documentation describes web search as an enabled capability and states that every web-search response includes source citations. Record the mode and the visible search activity in each capture. A response with no evidenced web search belongs in a separately labeled condition. See Anthropic’s web search documentation.
Perplexity. Perplexity documents PerplexityBot for surfacing and linking sites and a separate user-triggered fetcher for on-demand visits. These access descriptions do not establish a ranking formula or the source of any individual recommendation. Capture the answer and its sources directly. See Perplexity’s crawler documentation.
Connect referrals to qualified demand carefully
Keep the answer-capture log separate from analytics and intake or CRM records. Each system describes a different population, and merging them without a valid relationship produces false rates.
Record the analytics property, report, source and medium rule, reporting window, time zone, and session or user definition. Exclude identifiable staff testing and document the exclusion. Missing referrers and consent gaps stay unknown.
Never divide website referrals by your sampled citations to produce a click-through rate. Those records describe different populations. A platform impression report is also not a denominator for a cross-platform prompt sample.
Where permitted records can be linked, use a pseudonymous inquiry ID and an explicit linkage rule. Define qualification rate as unique qualified inquiries divided by unique linked inquiries whose status has been assessed within the stated window. Report the full linked cohort, the pending or unknown count, unlinked records, and qualification dates separately. Do not count pending inquiries as unqualified, and do not divide inquiries by sessions and call it a qualification rate.
For a law firm, intake decides whether an inquiry qualifies and whether a matter is retained. For SaaS, sales defines a qualified demo or opportunity. Downstream rates need a defined cohort with enough follow-up time. Keep client, matter, and sensitive lead details out of prompt logs.
Repeat the protocol and explain what changed
Run the next wave against the fixed prompt panel with the same modes, repeats, coding rules, and window where practical. Record changed conditions such as model labels, account state, features, source content, and missing attempts. Compare matched segments and keep new prompts separate.
Report percentage-point changes between comparable rates with their counts. Two small samples that differ are not a trend. Repeated answers to related prompts are not independent observations, so a conventional confidence interval does not remove selection bias.
Keep an intervention log of page edits, technical changes, campaigns, and releases with dates. A before-and-after comparison can show an association. It cannot isolate a content change from platform updates, demand shifts, competitors, or other work.
The measurement worksheet
Copy the three worksheets into your own document. Complete the protocol first, then duplicate an attempt record for each planned run and a source record for each citation. A blank field means not collected. Enter zero only for an observed absence.
Worksheet A, protocol.
- Project and decision
- Protocol version, owner, reviewer
- Target entity, accepted aliases, owned domains
- Audience, topic, intent segments
- Prompt register version, prompt IDs, exact text, source
- Platforms, modes, visible model labels
- Language, locale, device, account and personalization state
- Wave dates, time zone, repeats, session reset rule
- Planned attempts, eligibility, missing-data rules
- Coding definitions, uncertainty resolution, weighting rule
Worksheet B, attempt and source records.
- Attempt ID, prompt ID, repeat number, wave
- Exact prompt, timestamp, platform, mode, session conditions
- Status, feature present, capture reference
- Target mention: yes, no, uncertain, or not assessable
- Own-site citation: yes, no, uncertain, or not assessable
- Third-party citation about the target: yes, no, uncertain, or not assessable
- Recommendation: yes, no, uncertain, or not assessable, with supporting passage
- Source row: attempt ID, source ID, displayed URL, resolved URL
- Source owner and type, cited claim, support assessment
- Coder, second review, adjudication, exclusions
Worksheet C, report and business linkage.
- Segment, wave, protocol version
- Planned attempts, assessable counts by metric, missing, uncertain
- Mentions: numerator, denominator, rate
- Own-site citations: numerator, denominator, rate
- Third-party citations about the target: numerator, denominator, rate
- Recommendations: numerator, denominator, rate
- Feature appearances: numerator, denominator, rate
- Platform report, actual fields, filters, export date, anomaly notes
- Referral report, period, source rule, session definition, exclusions
- Linked inquiry cohort, qualification criteria, assessed count, qualified count, pending or unknown, missing links
- Outcome definition, cohort, follow-up period, observed count
- Changed conditions, confounders, interpretation, proposed next action
Methodology limitations
- A selected prompt panel cannot represent every buyer, language, location, or conversation. Publish the selection logic and the coverage gaps.
- The same protocol can return different answers. Preserve evidence instead of promising reproducible numbers.
- Mentions can be inaccurate, citations can fail to support claims, and recommendations can be ambiguous. Keep those quality checks separate from presence counts.
- Exports, screenshots, and dashboards use different units and collection rules. Do not merge them without a valid relationship.
- Referrals omit unobserved visits and later discovery. Linked leads exclude unlinked records. Disclose both.
- Observed change is not causation. No visibility rate guarantees qualified demand or a business outcome.
Sources and review status
Version 2.0 was prepared September 9, 2026 using the primary platform documentation linked beside each claim, verified live on that date. No original answer sample or client dataset was collected for this guide. The worksheet and counting protocol are recommendations for review and application, not validated performance findings. Record a new review date when the method or platform guidance is checked and revised, and do not refresh the date for cosmetic edits.
Written by Chuck Price, Founder of Measurable SEO. Read about the author and published work, or return to the research overview.
Discuss a measurement question
For help defining the question, the evidence, and the measurement boundaries for your firm, describe your measurement problem. The engagement context sits on the AI search consulting page.