A fair AI visibility test uses a defined set of buyer questions, records the conditions of each answer and repeats the same checks over time. It separates being mentioned, being recommended and being cited. One favourable answer is evidence of that answer, not a permanent rank or proof that an edit caused a change.
This is an editorial testing protocol. Apply it to the evidence your collector actually provides; do not infer undisclosed settings or searches. It helps you decide what to investigate and improve without turning a noisy result into an unsupported success claim.
Define the decision before choosing prompts
Start with a real buying task: choosing a local lead-monitoring tool, comparing costs or reporting AI visibility to a client. Write the question without inserting your brand unless you are deliberately measuring a branded task. Keep branded and unbranded questions in separate groups.
Set the intended market and language. A result collected for the United States does not establish Australian coverage. A country setting also does not prove that every answer reflects a particular suburb. Record the actual setting, not the audience you hoped to reach.
Use a small fixed baseline and a separate exploratory group. The baseline shows whether results change for the same questions. The exploratory group helps find new customer tasks. Adding easier questions to the baseline halfway through a test can make the headline percentage rise without any improvement on the original questions.
Record the conditions of every run
| Field | What to retain privately |
|---|---|
| Question | Exact wording and its parent topic; deduplicate identical variants. |
| Surface and method | The named engine or product surface and how the collector obtained the answer. |
| Context | Session history, memory or personalisation settings when observable; otherwise mark unknown. |
| Market and time | Configured geography, language, timestamp and timezone. |
| Evidence | Original answer, cited URLs and any explicitly disclosed search queries. |
| Run status | Successful, failed, missing or completed without the requested AI feature. |
Use fresh, controlled context for a baseline and record any limitations. An answer generated after a personal conversation about your company belongs in a separate investigation. Exclude contaminated results using a stated rule and retain the exclusion reason; do not quietly remove only unfavourable observations.
Keep product surfaces distinct. Google documents Search grounding for the Gemini API, including access to web information and citations. That does not make an API response interchangeable with an answer in a consumer interface. Record which method you actually tested.
Count outcomes at the right level
- Brand mention: the answer explicitly refers to the intended business, after resolving ambiguous names.
- Recommendation: the answer presents the business as a relevant option, rather than only discussing it.
- Citing answer: the answer includes at least one citation to your site.
- Cited page: a distinct URL used within an answer; count repeat passage links to the same page once for page coverage.
- Disclosed search: a query the provider actually exposes. No visible query is not proof that no retrieval happened.
Report each engine separately, with the number of eligible completed answers. For example, two citing answers out of 20 is a 10% observed citation rate for that panel. Six links spread across those two answers are still two citing answers. Track failed runs alongside the rate so reduced coverage cannot disappear into a flattering percentage.
Read the cited page and its role in the answer
Open the exact source URL. Identify the passage it supports: a definition, a comparison, a price, an example or a claim about a competitor. A link to your domain may support another company’s description. That can still be useful source visibility, but it is different from a recommendation of your product.
Then compare that reader task with your own closest page. Is a decision table missing? Is the price basis unclear? Does the article omit an important limitation? Use the fan-out content workflow to turn the gap into an editorial change. Treat the reason a source was selected as a hypothesis unless the evidence explicitly establishes it.
Repeat the panel after a documented change
Record the URL, the substantive edit and the actual deployment date. Check that the page is live before starting the after-period. Compare repeated batches at similar intervals and retain an unchanged topic or page as a reference where practical. Engine changes, indexing delays and competitor updates can still affect both groups.
Our suggested checkpoints are seven, fourteen and twenty-eight days after deployment. These are review dates, not guaranteed indexing or ranking deadlines. At each checkpoint, show the counts, engine mix, prompt coverage and any collection changes before interpreting movement.
Google’s guidance for AI features says the usual search fundamentals apply and that eligibility does not guarantee crawling, indexing or serving. Focus on useful, accessible content and internal links. A special file or extra posting frequency is not a guarantee of inclusion.
Separate visibility from business results
Track cited answers, referred visits, qualified enquiries and trial registrations as distinct outcomes. An engine can cite a page without sending a visitor, and a visitor can arrive through another route. For Google AI Overviews and AI Mode, Search Console reports their traffic within its overall Web search reporting, so a change in that total alone cannot identify the exact AI surface.
A useful conclusion is specific: the same panel produced more citing answers after a documented page change, with these limitations. Continue checking before describing a durable trend. Use the AI visibility buyer’s guide to evaluate whether a tool provides enough evidence to support that conclusion.
Sources
- AI features and your websiteGoogle Search Central · accessed 2026-09-18 · primary source
- Grounding with Google SearchGoogle AI for Developers · accessed 2026-09-18 · primary source