Module 08 · AI discovery and answer visibility

Measure AI visibility with a repeatable prompt-and-citation study

Build a small, auditable research study that records what answer products visibly say, link, cite, and omit over time.

Lesson 28AI discovery and answer visibility · Practical courseLast updated
78% through the course

What you will learn

AI visibility measurement is research, not a dashboard magic trick. One prompt can produce a different answer because of wording, mode, location, model version, time, personalization, retrieval availability, or simple variation. A useful study controls what it can, records what it cannot, and reports observations without claiming the system's hidden reasoning.

By the end of this lesson: Design a repeatable prompt-and-citation study with a fixed universe, clear coding rules, evidence snapshots, and honest reporting limits.

Why this matters

01

Without a fixed prompt set and a clear grading method, teams collect memorable screenshots: a competitor mention feels alarming, a brand mention feels like proof of success, and neither can be compared next month. A study turns anecdotes into a pattern that can guide content, product, reputation, or research work.

02

Citation and mention data can also reveal a useful gap. If answer products repeatedly surface inaccurate third-party facts, unclear owned documentation, or unsupported comparisons, the response may be to improve the source ecosystem—not to demand a model recommend the brand.

Keep the boundary clear: Do not claim that a prompt study measures all AI users, total market share, model training data, or causal ranking factors. Respect product terms, do not automate access where prohibited, and never present screenshots as proof of a guaranteed outcome.

The prompt-study evidence loop

Field note 28The prompt-study evidence loopCollect comparable evidence

A fixed universe makes comparison possible. Capture the exact product context and answer. Code visible mentions, citations, claims, sentiment, and gaps using rules written before scoring. Compare the same study over time, then choose a modest underlying improvement instead of reacting to one output.

Core concepts

01

Prompt universe

A prompt universe is a deliberately chosen list of customer questions, stages, markets, languages, and intents. It should include branded and non-branded questions, comparison questions, problem questions, and questions where the honest answer may be “not applicable.”

Use it when: Can another researcher explain why each prompt belongs in the study?

02

Capture context

Record product, model or mode if visible, logged-in state, geography, language, date and time, exact prompt, response text, links or citations, and any browsing or search mode. Save allowed evidence so the observation can be audited later.

Use it when: Could the team reproduce the conditions closely enough to compare a future run?

03

Coding rules

Define what counts as a mention, recommendation, citation, linked source, accurate claim, competitor reference, negative statement, omission, and unavailable result. Code uncertain cases as uncertain rather than forcing a yes or no.

Use it when: Would two reviewers reach a similar result from the same answer?

04

Directional interpretation

A change in observed mentions may reflect model behavior, prompt sensitivity, source availability, market context, or randomness. Treat it as a signal to investigate with source, traffic, and customer evidence—not as a direct performance metric.

Use it when: Does the conclusion describe the observation and its limits separately?

The practical method

  1. 01

    Write the research question

    Ask something practical, such as: “When people compare secure file-sharing tools for regulated teams, which decision criteria and sources appear?” Avoid the impossible question “How do we rank first in AI?”

  2. 02

    Build the prompt sample

    Use customer research, search queries, sales calls, support tickets, and category language. Balance lifecycle stages and include prompts where your product should not be the answer.

  3. 03

    Freeze the protocol

    Set the product, mode, locale, language, login state, run cadence, capture method, and rules for retries or unavailable responses. Version the protocol when it changes.

  4. 04

    Capture evidence consistently

    Run prompts manually or through approved methods. Save response text, cited links, screenshots or exports where permitted, timestamps, and a unique observation ID. Do not edit the answer to make coding easier.

  5. 05

    Code with a rubric

    Have one or two trained reviewers code outcome fields. Spot-check disagreement, keep a notes field for ambiguity, and preserve the original evidence beside the coded result.

  6. 06

    Report patterns and next questions

    Show prompt coverage, observed mention and citation patterns, source domains, accuracy issues, omissions, variance, and limitations. Pair every recommendation with the evidence that supports it and a follow-up test.

Guided workshop

Measure AI visibility as a bounded research program

This section turns the lesson into a bounded working session. It is designed to leave you with an AI visibility observation board that combines first-party signals where available, fixed customer scenarios, source analysis, accuracy review, and decisions.

Practice scenario

Practice scenario: A B2B cybersecurity company wants a monthly AI visibility report. Leaders ask for a single score that compares assistants, search experiences, and cited sources. The analyst knows the systems change, outputs vary by question and context, and not every response shows why it selected a source.

The team defines a modest research program. It uses a small, stable prompt set based on real customer stages, records the environment and output, checks cited URLs and source types, and adds first-party reporting only as one bounded signal where it exists. Subject experts review accuracy before the report is shared.

The report does not claim a universal position. It tells the team where customers may lack a clear source, which answers are incomplete or misleading, and what source, evidence, or customer explanation deserves attention next.

Build it step by step

01

Define the decision the study supports

Choose a practical question such as whether to improve a comparison page, clarify a product claim, build a source asset, or investigate a recurring customer objection. Keep the study tied to an action.

Make it tangible: Save a study decision statement. It helps the team decide what the report should help the team choose. Check starting with a demand for a single AI rank before moving forward.

02

Build a stable scenario set

Use real customer questions across discovery, evaluation, objections, local needs, and post-purchase support where relevant. Preserve wording, context, market, and follow-up order so observations remain comparable.

Make it tangible: Save a versioned scenario library. It helps the team decide which questions will be observed each cycle. Check changing the prompt set whenever the result is inconvenient before moving forward.

03

Capture evidence and source context

Save date, product, language, market, account or device context when material, response, cited URLs, source type, direct mentions, and missing information. Label direct observation separately from interpretation.

Make it tangible: Save an observation board. It helps the team decide what changed and what merely appears different. Check reporting a screenshot without the conditions that produced it before moving forward.

04

Add first-party signals carefully

Where a provider offers reporting, record what it measures, its limits, coverage, and date range. Use it to complement the study, not to generalise beyond the reporting product.

Make it tangible: Save a first-party signal note. It helps the team decide how a provider metric fits the evidence set. Check treating a citation trend as a universal answer ranking before moving forward.

05

Review accuracy and customer safety

Ask experts to check material claims, omissions, sources, and advice. Flag harmful or confusing responses and decide whether a clearer owned source, correction request, or customer communication is needed.

Make it tangible: Save an accuracy review record. It helps the team decide which finding deserves urgent source work. Check celebrating a mention before checking whether it helps the customer before moving forward.

06

Make one bounded improvement

Choose a source-quality, access, explanation, or evidence change. Repeat the stable study after an appropriate interval and report what was observed without claiming direct causation.

Make it tangible: Save a change-and-retest card. It helps the team decide what the program learned from the next cycle. Check changing several variables and calling the result a proof before moving forward.

Working template

Use these fields in a document, task, or spreadsheet. Keep the evidence close to the decision.

  1. Study decision: State the product, content, or evidence decision the research should inform. A leader should know what action the report could trigger.
  2. Customer scenario: Record prompt, buyer context, market, follow-up sequence, and expected customer task. A researcher should be able to repeat the observation.
  3. Observation: Save date, conditions, answer, cited URLs, source type, and notable missing information. A reviewer should distinguish the output from its interpretation.
  4. First-party signal: Describe any official reporting, coverage, limits, and period used. An analyst should avoid extending the metric beyond its stated scope.
  5. Accuracy review: Mark claims, citations, omissions, and customer-risk notes with expert input. A subject owner should approve the assessment for high-stakes topics.
  6. Next action: Choose one source, content, access, or evidence improvement and a retest window. The team should be able to learn from a controlled change.

Quality review before you ship

Use these checks while the evidence, owners, and customer context are still easy to correct.

  1. Frame each research check as a real buyer question with a context and decision stage. A broad prompt may be interesting, but it cannot tell a team which page, fact, or relationship deserves work.
  2. Save the response, source set, date, locale, account state, and query wording. Without those conditions, another reviewer cannot tell whether a difference reflects the site, the system, or the test setup.
  3. Only act on a recurring, explainable gap that your team can improve responsibly. A research output should lead to stronger sources and clearer facts, not to attempts to manipulate wording or citations.

Decision rules for the real world

The study sees no brand mention

Do: Inspect the customer task and source gap before deciding whether a mention is even the useful outcome.

Avoid: Do not force the brand into prompts or claim the system is biased.

A cited source is inaccurate

Do: Correct your own source where relevant, document the risk, and use appropriate reporting channels if available.

Avoid: Do not treat a citation as an endorsement of the answer.

Official reporting and prompts differ

Do: Report each method's scope and investigate whether the scenarios or time windows differ.

Avoid: Do not average unrelated signals into one score.

Leaders want a single number

Do: Offer a short evidence summary with trends, limits, customer scenarios, and next decision.

Avoid: Do not invent precision to satisfy a reporting format.

Coach notes

  • The aim is not to win a fictional leaderboard. It is to improve the sources customers and systems can use.
  • A small stable study is easier to review honestly than a large, changing set of prompts.
  • Accuracy review belongs in the workflow because a visible but wrong answer can harm a customer.

A payroll platform studies regulated-business questions

A payroll platform wants to understand answer visibility for nonprofit organizations. Instead of checking one prompt, it creates 30 questions across setup, compliance, pricing, migration, and comparison stages. It includes questions where specialist nonprofit tools are likely to be a better fit. Each run records product mode, US locale, exact prompt, answer, citations, brand mentions, stated criteria, and factual errors.

Two reviewers find that the brand is rarely named on broad comparison questions, but its public documentation is cited in migration answers. They also find a recurring inaccurate claim about nonprofit tax filing made on several third-party pages. The team creates a clearer migration guide, corrects its own public terminology, contacts the sources it can responsibly correct, and re-runs the same study on schedule without treating a later mention as proof of causality.

What changed: The study produces a credible source-and-content roadmap rather than a misleading single “AI rank.”

Make it stronger

Use inter-rater reliability

For a consequential study, independently code a sample of answers and compare agreement. Refine unclear definitions before publishing a trend. If reviewers cannot agree, the metric is not mature enough for strong conclusions.

Preserve a changelog

If you add prompts, change locale, switch account state, or alter coding rules, report the break in comparability. Keep the old series visible or restart the trend honestly.

Segment cited sources

Classify citations as owned, publisher, community, regulator, academic, marketplace, or unknown. This helps distinguish an owned-content gap from a reputation or data-ecosystem gap.

Guard against confirmation bias

Include prompts that do not name your brand, competitors your team would rather ignore, and questions where the correct result is a caution or no recommendation. A study that only seeks wins teaches little.

Current guidance

Use a visibility evidence ladder

No tool gives a universal AI rank. Start with first-party product data when it is available, then use a repeatable prompt study to understand what people can actually see.

  • Use Search Console's Generative AI performance report when it is available for your site. Google is rolling it out gradually; it can show generative-feature impressions and the pages, countries, devices, and dates involved.
  • Use Bing Webmaster Tools' AI Performance report for aggregate citations, cited pages, and grouped grounding queries. It is a public preview and is not a ranking, authority, click, or causal-impact report.
  • Use a manual prompt study to inspect answer quality, cited sources, omissions, and context. Freeze the product, mode, language, market, date, and coding rules before comparing runs.
  • Use analytics and the CRM to learn whether discovery became useful customer behavior. A citation or mention alone does not prove traffic or revenue.
  • Change one controllable thing at a time, then report the observation and its limits separately from any recommendation.

Use this before you publish

  • The report labels each signal as first-party platform data, a manual observation, or a business outcome.
  • Prompt studies preserve the exact prompt, context, response, cited links, and date.
  • No chart is labelled as an ‘AI rank’ or presented as proof of causation.

Current field note

Measure AI visibility without inventing an AI rank

Answer systems do not provide one stable, universal ranking. Where first-party reporting is available, use it as one bounded signal, then pair it with a repeatable prompt study to compare customer scenarios, source types, accuracy, and useful next steps.

  • Use Bing Webmaster Tools' AI Performance reporting for cited-URL and citation-trend signals where it is available; do not treat it as a universal AI ranking report.
  • Build a small prompt set from real discovery, comparison, and objection questions.
  • Separate direct observation from interpretation in the report.
  • Prioritize source quality and customer usefulness over attempts to force a mention.

Official reference: Bing Webmaster: AI Performance ↗

Lesson artifact

AI visibility observation board

Design a ten-prompt pilot study that another person could run next month.

Research question: State the customer decision and what you want to learn—not the answer you hope to get.

Prompt universe: List ten prompts across need, comparison, evaluation, and support stages.

Protocol: Record product, mode, locale, language, account state, cadence, and retry rules.

Capture fields: Define the exact evidence to save for each observation.

Rubric: Write coding rules for mention, citation, accuracy, source type, and ambiguity.

Quality control: Choose a double-coded sample and disagreement-resolution process.

Decision use: Name the kinds of source, content, product, or reputation improvements the study could justify.

Done looks like this: The pilot is reproducible, auditable, and narrow enough that its results will not be exaggerated.

Before you move on

  • The prompt set represents customer decisions, not only branded queries or desired wins.
  • Every observation records product context, exact prompt, date, response, and visible sources.
  • Coding definitions distinguish mention, recommendation, citation, accuracy, and uncertainty.
  • Protocol changes are logged so trend comparisons remain honest.
  • Recommendations improve controllable information and evidence rather than claiming control over model outputs.

Put the lesson into practice.

Create a free Spacebrain account and use the SEO suite with your own data providers.

Start for free →