How to Measure AI Search Visibility Without Confusing It With Demand

Methodology guide · Version 1.1
Prepared by OpenAI Codex for Chuck Price. Prepared and documentation reviewed: September 7, 2026.

Measure AI search visibility by recording what answers show, what sources they cite, and what visits you can observe—then assess qualified inquiries and business outcomes separately. Every reported rate needs a defined unit, denominator, time period, and explanation of what the sample leaves out.

This guide proposes a repeatable measurement process. It contains no original dataset, client results, platform benchmark, or universal AI visibility score. Reproducing the process does not guarantee identical answers.

Apply the measurement worksheet or begin with the definitions below.

Define six different signals

Mention. A captured answer names the target person, organization, or product. Record whether the description is accurate; a name match alone does not establish a positive assessment.

Citation. An answer attaches a source link to support its response. Record the actual destination and the claim it appears to support. A citation to a third-party page about the target differs from a citation to the target’s own website. A link is not proof that the source supports the claim.

Recommendation. The answer presents the target as a candidate for the user’s stated need. Mere inclusion in a list, a disclaimer, or a negative example is not automatically a recommendation. Retain the relevant passage and flag ambiguous cases for review.

Referral. An observable website visit associated with an identifiable referring source under your analytics rules. This is measured from website records, not inferred from a citation in a sampled answer.

Qualified inquiry. A unique inquiry that meets the organization’s documented qualification criteria. Intake or sales determines qualification. A form event, call, signup, or demo request alone does not establish it.

Business outcome. A separately recorded downstream event, such as a retained matter, sales opportunity, or completed sale. State which outcome you mean and who validates it. Do not combine stages or treat revenue and pipeline as interchangeable.

Choose the decisions and sample before testing

Start with the decision the measurement should inform: whether important product information is misrepresented, whether buyer questions find useful sources, or whether observed referrals warrant further investigation. Choose platforms and search experiences relevant to that decision.

  • Build a prompt register from documented buyer questions, sales or intake themes, and relevant search research. Record each prompt’s source. Prompts you design are test inputs, not evidence of actual query volume.
  • Assign each prompt an ID, audience, topic, intent stage, language, and branded/unbranded status. Keep discovery, evaluation, and decision questions distinguishable.
  • Freeze the prompt list, target entity aliases, dates, platforms, modes, and repeats per prompt before a measurement wave. Set the number of repeats according to capacity and report it; there is no universal minimum that makes a convenience sample representative.
  • Keep a fixed panel for comparisons. Label exploratory prompts separately. Record additions and removals rather than quietly changing the test.
  • Use a consistent session protocol: fresh conversation where available, recorded personalization/account state, locale, device, and model label when visible. Record unknown settings as unknown.

Capture one record per planned attempt

An attempt is one prompt run in one specified platform/mode/session at a recorded time. Give every planned attempt a unique ID. Preserve the exact prompt, answer, visible sources, timestamp with time zone, and a screenshot or permitted export reference. Keep follow-up turns separate from first answers.

Record the response status: assessable answer, no answer feature, refusal, technical failure, incomplete capture, or not run. A fully captured answer without citations is an observed zero for citation presence. An inaccessible or truncated answer is missing evidence, not a zero. For Google Search, record whether an AI Overview appeared; do not silently discard successful searches without one.

For each citation, keep a separate source row linked to the attempt ID. Store the displayed link and resolved destination when observable, source owner/type, related passage, and support assessment: supports, partly supports, does not support, or not checked. Do not describe unchecked citations as verified evidence.

Use target-specific yes/no/uncertain coding for mentions and recommendations. Separate own-site citations from third-party references. Have a second reviewer check ambiguous cases when possible; retain the initial code, adjudication, reviewer, and reason. Keep unresolved cases out of the applicable rate and disclose how many were excluded.

Put a denominator beside every rate

Assessable-answer yield: assessable answers ÷ all planned attempts. This measures usable-answer coverage, not technical completion. Report no-feature responses, refusals, technical failures, incomplete captures, and attempts not run separately so missing work cannot disappear.

Sampled mention rate: answers with a confirmed target mention ÷ answers assessable for target mention in the same platform, mode, wave, and prompt segment. Count each answer once, regardless of repeated mentions.

Sampled own-site citation rate: answers citing at least one target-owned URL ÷ answers assessable for citation presence in that same segment. Multiple citations in one answer still count as one positive answer. For third-party citations about the target, count answers with at least one such citation and divide by answers assessable for that specific citation category. Exclude unrelated third-party sources from this numerator.

Sampled recommendation rate: answers with a confirmed recommendation ÷ answers assessable for recommendation status in the predefined relevant-intent segment. Define which intents are eligible before collection.

Answer-feature appearance: successful searches showing the specified feature ÷ all successfully captured searches in that test segment. For AI Overview analysis, report both this rate and any visibility rate conditional on an Overview appearing. They answer different questions.

Show numerator and denominator alongside the percentage. A zero denominator produces “not measurable,” not 0%. State missing and uncertain counts. Equal planned repeats prevent one prompt from dominating through extra reruns; if completed repeats differ, report per-prompt coverage and use a disclosed weighting rule. Never rerun only unfavorable answers.

These are descriptive rates for the sampled answers. They are not market share, audience reach, query-volume estimates, or the probability that an arbitrary buyer will see the target. Do not average platforms into an unexplained universal score.

Separate documentation, observations, and recommendations

Documented behavior means a linked primary source describes a platform feature at the review date. Observation means a dated capture shows what happened in a particular attempt. Recommendation means an analyst proposes an action based on the evidence and its limits. Label these separately in every report.

The platform descriptions below were checked on September 7, 2026. They are documentation summaries, not test findings about any business. Product behavior and reporting can change.

Google: Google’s June 2026 announcement, updated August 31, describes dedicated generative-AI performance views for Search and Discover, including impressions, pages, countries, dates, and Search device information. The data also contributes to overall performance reporting. Use the dedicated view and record its actual fields and filters; do not treat overall Web totals as isolated AI performance. Google’s reporting announcement.

This announcement establishes impression reporting; it does not justify assuming a separate click, query, recommendation, or lead measure. Do not add the dedicated report to overall totals, which would double count overlapping data. Check Google’s data-anomaly log before comparing periods, and annotate affected dates.

ChatGPT: OpenAI currently distinguishes OAI-SearchBot’s search purpose from GPTBot’s training control. Crawler access is not evidence that a particular answer cited or recommended a site. Record the observed ChatGPT experience and whether search is evidenced in the response; do not infer that from the product name alone. OpenAI crawler documentation.

Claude: Claude’s help documentation describes web search as an enabled capability and describes source citations in its web-search responses. Record the mode and visible search activity used in each capture. A response without established web-search use belongs in a separately labeled condition. Claude web-search documentation.

Perplexity: Perplexity documents PerplexityBot for surfacing and linking sites in search results and a separate user-requested fetcher. These access descriptions do not establish a ranking formula or the source of any individual recommendation. Capture the answer and sources directly. Perplexity crawler documentation.

Repeat the protocol and explain what changed

Run the next wave against the fixed prompt panel with the same planned modes, repeats, coding rules, and sampling window where practical. Retain changed conditions: model labels, account state, features, source content, and missing attempts. Compare matched segments and keep new prompts separate.

Report percentage-point changes between comparable rates, plus their counts. Do not claim a meaningful trend merely because two small samples differ. Repeated answers to related prompts are not independent random observations; a conventional confidence interval does not remove selection bias.

Keep an intervention log with page edits, technical changes, campaigns, releases, and their dates. A before/after comparison can reveal an association; it cannot isolate the effect of a content change from platform updates, demand shifts, competitors, or other work.

Connect referrals with qualification carefully

Keep the answer-capture log separate from analytics and intake/CRM records. Record the analytics property, report, source/medium rule, reporting window, time zone, and session or user definition. Exclude staff testing where identifiable and document the exclusion. Missing referrers and consent-related measurement gaps remain unknown.

Never divide website referrals by your manually sampled citations to calculate click-through rate: those records describe different populations. Likewise, a platform impression report is not a denominator for a cross-platform prompt sample.

If permitted records can be connected, use a pseudonymous inquiry ID and an explicit linkage rule. Define qualification rate as unique qualified inquiries divided by unique linked inquiries whose qualification status has been assessed within the stated observation window. Report the full linked cohort, pending or unknown qualification counts, unlinked records, and qualification dates separately; do not silently count pending inquiries as unqualified. Do not divide inquiries by sessions and call the result a qualification rate.

For a law firm, intake decides whether an inquiry qualifies and whether a matter is retained. For SaaS, sales defines a qualified demo or opportunity. Downstream rates need a defined cohort with enough follow-up time. Self-reported discovery sources can add context but should remain labeled self-report. Avoid transferring client, matter, or sensitive lead details into prompt logs.

Apply the measurement worksheet

Copy the following blank worksheet into your own document. Complete the protocol first, then duplicate an attempt record for every planned run and a source record for each citation. Blank fields mean not collected; enter zero only for an observed absence. No registration or submission is required.

Worksheet A — protocol

  • Project / decision: ______
  • Protocol version / owner / reviewer: ______
  • Target entity / accepted aliases / owned domains: ______
  • Audience / topic / intent segments: ______
  • Prompt register version / prompt IDs / exact text / source: ______
  • Platforms / modes / visible model labels: ______
  • Language / locale / device / account and personalization state: ______
  • Wave dates / time zone / repeats / session reset rule: ______
  • Planned attempts / eligibility / missing-data rules: ______
  • Coding definitions / uncertainty resolution / weighting rule: ______

Worksheet B — attempt and source records

  • Attempt ID / prompt ID / repeat number / wave: ______
  • Exact prompt / timestamp / platform / mode / session conditions: ______
  • Status / feature present / capture reference: ______
  • Target mention: yes / no / uncertain / not assessable: ______
  • Own-site citation: yes / no / uncertain / not assessable: ______
  • Third-party citation about the target: yes / no / uncertain / not assessable: ______
  • Recommendation: yes / no / uncertain / not assessable; supporting passage: ______
  • Source row: attempt ID / source ID / displayed URL / resolved URL: ______
  • Source owner/type / cited claim / support assessment: ______
  • Coder / second review / adjudication / exclusions: ______

Worksheet C — report and business linkage

  • Segment / wave / protocol version: ______
  • Planned attempts / assessable counts by metric / missing / uncertain: ______
  • Mentions: numerator ______ / denominator ______ / rate ______
  • Own-site citations: numerator ______ / denominator ______ / rate ______
  • Third-party citations about the target: numerator ______ / denominator ______ / rate ______
  • Recommendations: numerator ______ / denominator ______ / rate ______
  • Feature appearances: numerator ______ / denominator ______ / rate ______
  • Platform report / actual fields / filters / export date / anomaly notes: ______
  • Referral report / period / source rule / session definition / exclusions: ______
  • Linked inquiry cohort / qualification criteria / assessed count / qualified count / pending or unknown / missing links: ______
  • Outcome definition / cohort / follow-up period / observed count: ______
  • Changed conditions / confounders / interpretation / proposed next action: ______

Methodology limitations

  • A selected prompt panel cannot represent every buyer, language, location, or conversation. Publish selection logic and coverage gaps.
  • The same protocol can return different answers. Preserve evidence rather than promise exact answer reproduction.
  • Mentions may be inaccurate, citations may fail to support claims, and recommendations may be ambiguous. Keep those quality checks separate from presence counts.
  • Tool exports, screenshots, and platform dashboards have different units and collection rules. Do not merge them without a valid relationship.
  • Referrals omit unobserved visits and later discovery. Linked leads exclude unlinked records; disclose that selection limitation.
  • Observational changes do not prove causation. No visibility rate guarantees qualified demand or a business outcome.

Sources, version, and review status

Version 1.1 was prepared September 7, 2026 using the primary documentation linked beside the platform claims. No original answer sample or client dataset was collected for this guide. The worksheet and counting protocol are recommendations for review and application, not validated performance findings.

Draft preparation and documentation review: OpenAI Codex. Prepared for Chuck Price, Founder of Measurable SEO. Chuck’s editorial review and final author attribution are pending. Record a new substantive review date when the method or platform guidance is checked and revised; do not refresh dates for cosmetic edits.

About Chuck Price · Research

Discuss a measurement question

For help defining the question, evidence, and measurement boundaries, discuss your measurement approach. The AI search consulting page explains the available engagement context.