How to Prepare an AI Visibility Report: Responses per Engine, Margin of Error and Raw Data

By Ömer Cenk Tokgöz · Published: · Updated:

How to Prepare an AI Visibility Report: Responses per Engine, Margin of Error and Raw Data

A single visibility score hides more than it shows. Using our own two scans, we show how to report engines, response counts, margins of error and raw data — and why an unmeasured engine is not 0%.

Correction (6 October 2026): Our measurement provider withdrew two data sets used in this post. The 5 October 2026 scan was only partially measured: the provider ran out of credit during the scan; ChatGPT answered all 300 questions, while Gemini, Perplexity, Claude and Grok answered only 42–44. The Gemini and Google AI Overview zeros in the 15 September 2026 scan were also a measurement error. Those numbers in this post are invalid. The 15 September ChatGPT measurement (249 answers, 11.2%, ±3.9 points) stands. We will update this post when a complete new scan is available.

Short answer: A reliable AI visibility report does not stop at a single score. It shows which engines were actually measured, how many responses were collected per engine, the margin of error for each rate, and the raw data. An engine that was not measured is reported as "not measured", not "0%". A difference smaller than the margin of error is not counted as a change.

In this post we walk through the report we prepared for our own site. Every figure comes from two scans of botfusions.com (September 15, 2026 and October 5, 2026) and from our second measurement line on October 5, 2026.

Why a single score is not enough

In the October 5, 2026 scan our overall score was 38/100 and our mention rate was 14.8%. Those two numbers tell you something, but not what is working and what is not. Read engine by engine, the same scan looks different:

Engine Responses Mention rate Margin of error
ChatGPT 300 16.0% ±4.1 pp
Gemini 42 11.9% ±9.9 pp
Perplexity 44 13.6% ±10.2 pp
Grok 44 15.9% ±10.7 pp
Claude 44 9.1% ±8.8 pp
Google AI Overview — not measured —

Source: measurement platform scan, October 5, 2026, 474 responses.

ChatGPT's 16.0% and Claude's 9.1% are not the same kind of number. The first rests on 300 responses, the second on 44. That is why a report has to show the response count and the margin on every row.

Step 1: State the engine set

The first line of the report answers "which engines did we measure?" A measurement platform can change the engine set from scan to scan. Ours did:

  • September 15, 2026 scan: ChatGPT, Gemini, Google AI Overview.
  • October 5, 2026 scan: ChatGPT, Gemini, Perplexity, Claude, Grok. Google AI Overview did not run.

Writing "0%" in the Google AI Overview row of the October 5 report would mislead the reader. That engine was not queried at all that day. The correct label is "not measured".

When the engine set changes, the overall scores of the two scans cannot be compared either. Putting the September 15 score next to the October 5 score and saying "it went up" is like comparing grades from different exams. We compare only engines that ran in both scans, one engine at a time.

Step 2: State the number of responses per engine

A mention rate is a sample proportion: the share of answers from that engine that include the brand. When the denominator is small, the rate swings. Gemini's 11.9% in the October 5, 2026 scan means 5 out of 42 responses. One more mention would have made it 14.3%.

Put the response count next to every rate. A reader who sees "11.9%" should know it is 5/42.

Step 3: State the margin of error and do not count small differences as change

The platform reports a margin of error in percentage points for each engine in the October 5, 2026 scan. We recomputed these margins ourselves. They match the standard formula for a sample proportion at a 95% confidence level.

When we compare two scans we use the combined margin (the square root of the sum of the squared margins):

  • ChatGPT: 11.2% on September 15, 2026 (249 responses, ±3.9), 16.0% on October 5, 2026 (300 responses, ±4.1). The difference is 4.8 points. The combined margin is ±5.7 points. The difference is inside the margin, so we do not count it as an improvement.
  • Gemini: 0% on September 15, 2026 (379 responses, ±0.5), 11.9% on October 5, 2026 (42 responses, ±9.9). The difference exceeds the combined margin (±9.9). But 42 responses is a small sample. We report it as a signal to confirm in the next scan, not as a result.

The rule is simple: if the difference is smaller than the margin, the report says "no change". Good news and bad news get weighed on the same scale.

Step 4: Use a fixed, unbranded query set

A question that contains the brand name ("What is Botfusions?") brings up the brand by itself. The questions that actually measure visibility are unbranded, such as "Which GEO agency in Istanbul is best?"

Write the query set at the top of the report and repeat the same set under the same conditions in later measurements. If the query set changes, the comparison breaks just as it does when the engine set changes.

Step 5: Add measurement conditions and raw data

For someone else to verify a rate, they need: the model version, whether web search was on, a timestamp for each response, and the full text of the response.

On October 5, 2026 we used a second measurement line to document these conditions: 8 queries, 4 models, 32 responses in total. Each row has the model name, the web search flag, a timestamp and the full response text.

Model Web search Mentioned in 7 unbranded queries "Do you know Botfusions?"
openai/gpt-4.1-mini No 0/7 Does not know us
anthropic/claude-haiku-4.5 No 0/7 Does not know us
google/gemini-2.5-flash No 0/7 Knows our old positioning
perplexity/sonar Yes 3/7 Knows us correctly

Source: second measurement line, October 5, 2026, 32 responses.

This is a small sample. It is not a statistical result; it shows a direction: the model that searches the web finds us, while models that answer from training data alone do not yet. The report states this limit plainly.

Step 6: Write the interpretation and next steps separately

The measurement section says what you saw. The interpretation section says what you will do. Keep them apart. The interpretation section of our October 5, 2026 report was:

  1. No meaningful change on ChatGPT. Continue the current content work.
  2. Confirm the Gemini increase in the next scan. Do not present it as a win before that.
  3. No new information on Google AI Overview. Ask for it to be measured in the next scan.
  4. We do not appear in models that do not search the web. Content will not change that quickly. Set expectations accordingly.

Report skeleton

You can build your own report in this order:

  1. Summary: three sentences. What was measured, what was found, what comes next.
  2. Scope: scan date, engine set, engines not measured, query set.
  3. Engine table: engine, responses, mention rate, margin of error.
  4. Comparison with the previous scan: only engines that ran in both, with the combined margin.
  5. Measurement conditions: model version, web search, timestamps.
  6. Interpretation and next steps: numbered list.
  7. Appendix: raw data file (query, engine, response text, mentioned yes/no).

Common mistakes

  • Writing 0% for an engine that was not measured. If the engine is missing from the scan, mark the row "not measured".
  • Comparing overall scores of two scans with different engine sets. Compare engine by engine.
  • Leaving out the margin of error. A rate without a margin leads readers to treat small swings as change.
  • Reporting only the good news. Showing the Gemini increase while hiding that the ChatGPT difference is not meaningful destroys trust in the report.
  • Withholding the raw data. Without raw data a rate cannot be verified.

We covered the tracking setup in our LLM citation tracking guide, and how we read AI traffic itself in our zero-click search post.

What we claim and what we do not

What we claim: every figure in this post comes from our own measurements and is given with its date, response count and margin of error.

What we do not claim: that these figures generalize. The Gemini, Perplexity, Grok and Claude rows rest on 42-44 responses. They are a starting point, not a trend.

How to Prepare an AI Visibility Report in 6 Steps

A reporting workflow that shows engine set, responses per engine, margin of error and raw data, based on our own scans of September 15 and October 5, 2026.

  1. Step 1: State the engine set

    List the engines that actually ran in the scan. Mark any engine that did not run as "not measured", never as 0%.

  2. Step 2: State responses per engine

    Put the number of responses next to every mention rate so readers can see how large the sample is.

  3. Step 3: State the margin of error

    Give each rate a ± margin in percentage points. When comparing scans, use the combined margin and do not count smaller differences as change.

  4. Step 4: Use a fixed unbranded query set

    Measure with questions that do not contain the brand name and repeat the same set under the same conditions.

  5. Step 5: Add conditions and raw data

    Record model version, web search, timestamp and full response text for every answer, and attach the raw file.

  6. Step 6: Separate interpretation from measurement

    Write what you observed and what you will do next in separate sections, with next steps as a numbered list.

Ömer Cenk Tokgöz

Ömer Cenk Tokgöz — Founder, Botfusions

Founder of Botfusions; focuses on Generative Engine Optimization (GEO) and AI / answer-engine visibility. Author of the AI Visibility methodology.

Author profile & areas of expertise → · Editorial standards →

Related Posts

All Posts