An AI search visibility report should tell a client what was observed, how many observations support it, what changed, and what decision follows. It should not compress sampled answers, crawler requests, and human visits into one impressive but unauditable score.

The test is simple: can a skeptical reviewer trace every headline back to the prompts, answers, URLs, analytics rows, or supplied logs that support it?

This template works with any collection tool. Use the agency AEO tools guide for procurement and the sample AI visibility report to inspect the finished format.

The report outline to copy

Use these named sections in this order. The sequence moves from instrument to observation to decision, so a conclusion never appears before its evidence.

1. Scope and measurement window

State:

  • client, brand, domain, and market in scope;
  • first and last date included;
  • report generation date and time zone;
  • sampled answer products and disclosed models or modes;
  • analytics property and session scope, when connected;
  • edge or server log source, when supplied; and
  • the comparison window, if this report makes a change claim.

Do not label an integration “connected” when the report did not receive usable data from it during the window. Say not connected, connected with no usable rows, or connected with observations. Those states mean different things.

2. Prompt panel and selection method

List the stable buyer questions or provide an appendix. Give every prompt an ID, version, intent cluster, priority state, and reason for inclusion.

Explain whether the panel came from customer interviews, sales objections, search research, landing-page behavior, category comparisons, or another source. A panel is an instrument designed by the reporter, not a census of every private question buyers ask.

Keep branded and neutral prompts distinct. A question that supplies the client’s name tests aided treatment. A neutral category question tests unaided inclusion. Combining them without a label can flatter the result.

3. Coverage

Report planned samples and completed usable answers as a fraction:

Completed N of M planned answer samples during the window.

Then list failures, refusals, timeouts, and blocked runs by engine and reason. Do not silently remove them. Appearance incidence may use completed answers as its denominator, but the report must also disclose completed coverage out of planned coverage.

If one engine failed disproportionately, say so before presenting a blended number. The observed mix may no longer match the planned mix.

4. Appearance incidence

For each engine and question cluster, show:

Brand appeared in N of M completed answers.

If you add a percentage, keep the fraction beside it. The denominator is not a footnote; it is part of the metric.

The free AI visibility check demonstrates this reporting format with three OpenAI questions that include your domain. Its result is a prompted-domain sample, not a measure of unaided brand discovery or a substitute for the client panel.

Do not present a position inside a generated list as a durable rank. The answer may not contain an ordered list at all, and another run can compose a different response. Appearance incidence answers a narrower question: how often did the brand appear in this disclosed sample?

5. Citation incidence and cited URLs

Separate brand appearance from own-domain citation. Report:

  • answers citing the client’s domain out of completed answers;
  • exact own-domain URLs cited;
  • recurring third-party domains and pages;
  • the claim or answer section each citation supported; and
  • brand-absent answers where a competitor or source recurred.

Count answers containing a source when reporting incidence. One response containing several links to the same domain is still one answer citing that domain. The detailed workflow for preserving raw and normalized URLs lives in the guide to tracking AI search engine citations.

6. Competitor presence

Report named competitor appearances out of the same completed-answer denominator, segmented by prompt cluster and engine. Attach each claim to the answer rows behind it.

Competitor presence is not market share. Use the selected panel to identify repeated decision contexts and sources, not to declare a universal category leader.

7. AI-referred human visits

Where analytics is connected, show recognizable AI-referred sessions out of all recorded sessions in the same date window. Retain the source/medium rows and relevant landing pages.

This is a human-visit measure. It is not evidence that the brand appeared in a sampled answer, that a particular citation caused the click, or that an AI crawler fetched the page. If attribution is missing, report what the analytics system recorded rather than estimating the visits it may have missed.

8. Verified crawler access

Include this section only when suitable server or edge evidence exists. State the log source, retention window, bot-identity method, verification-date source, named agent, requested URL, and observed response outcome.

A user-agent string alone is a candidate request, not verified identity. A successful fetch is infrastructure evidence, not proof of indexing, citation, recommendation, or a human visit.

9. Change since the previous comparable window

Show the previous and current fractions with the corresponding dates and coverage. Keep engine and prompt-cluster detail available beneath any rollup.

Use changed, increased in this sample, or decreased in this sample. Do not say an intervention worked merely because a later sample moved. Retrieval, models, other web sources, and ordinary run-to-run variation can all change between windows.

10. Method changes

Give method changes their own section. Record:

  • added, removed, or edited prompts;
  • engine, model, product mode, location, or account changes;
  • sampling-frequency or repetition changes;
  • classifier changes for brands, citations, or referrals;
  • analytics-property or consent changes;
  • log-source, WAF, or identity-rule changes; and
  • missing data or outages.

A chart annotation saying “panel v2 begins” is more honest than drawing an uninterrupted line through a changed instrument.

11. Limitations

State what the report cannot know: every private prompt, every answer, future output, the exact cause of a change, the downstream use of a crawler fetch, unattributed influence, and causation from an ordinary before-and-after observation.

Limitations tell the client where a decision can be strong and where it remains provisional.

12. Next tests

End with a short queue. Each item should contain:

  1. the observed gap;
  2. the proposed single change;
  3. the pages and prompt cluster affected;
  4. the next comparable observation window; and
  5. the result that would cause the team to retain, revise, or stop the change.

Avoid an activity list detached from evidence. “Publish more content” is not a test. “Add a dated specifications section to the page absent from recurring comparison citations, then resample the unchanged comparison cluster” is.

Definitions block: every metric gets a denominator

Place this block near the front of the report so the client does not have to infer what a chart means.

MetricDefinition
Collection coverageCompleted usable answers ÷ planned answer samples
Appearance incidenceCompleted answers naming the tracked brand ÷ completed answers in the stated segment
Own-domain citation incidenceCompleted answers citing at least one tracked-domain URL ÷ completed answers in the stated segment
Competitor presence incidenceCompleted answers naming the competitor ÷ completed answers in the stated segment
Source incidenceCompleted answers citing the source domain ÷ completed answers in the stated segment
AI-referred session shareRecognizable AI-referred sessions ÷ all recorded sessions in the same analytics window
Verified crawler accessVerified requests from the named agent ÷ the relevant retained request set, with source and window stated

A metric without a denominator does not go in the report. A raw count can still be useful, but its eligible population and time window must remain visible.

Keep the three evidence lanes separate

Keep three independent lanes:

  1. Sampled answers: what selected answer systems returned for a disclosed prompt panel.
  2. Verified crawler access: what identified infrastructure requested in supplied server or edge evidence.
  3. AI-referred human visits: which sessions analytics attributed to recognizable assistant sources.

One can occur without the others. A citation can receive no click; a fetch may never be observed in an answer; a referral may come from outside the panel.

How Brandvane measures AI visibility defines the claim boundary for each lane. Use that boundary rather than creating a composite index whose weights conceal the events underneath it.

For the evidence framework, see how Brandvane measures AI visibility. Automated weekly AI reports for ChatGPT and Google AI Overviews are coming soon; the planned reports will keep the two surfaces separate and show completed-check counts.

Comparability rules for a defensible series

Editing a prompt breaks the series

Punctuation correction is one thing; changing the subject, constraint, brand cue, or buyer stage is another. If meaning changes, assign a new prompt version and begin a new segment. Preserve the old observations.

Do not rewrite history by applying the new prompt label to old answers. The report should let a reviewer see exactly what was asked at the time.

Changing engine coverage changes the denominator

If a window sampled OpenAI, Anthropic, and Perplexity but the next sampled only two, the blended figures are not directly equivalent. Report each engine separately and annotate the coverage change. Never leave the client to assume the engine set remained stable.

Brandvane’s current subscription fresh checks sample OpenAI on demand. Automated weekly AI reports for ChatGPT and Google AI Overviews are coming soon. API answers and observations of consumer products must be labeled separately; neither provides access to every private answer.

Product and model changes are confounders

A model snapshot, retrieval behavior, or consumer-product update can move the series even when the client changes nothing. Record the disclosed mode when available and annotate known changes; do not pretend an outside observer can reconstruct their full effect.

Classifier changes need versioning too

If the team expands a brand alias list, changes URL normalization, or adds a referral-domain marker, later totals may capture events earlier logic missed. Version the classifier and, if raw evidence permits, recalculate both windows with the same rule. If it does not, state the break.

Failed runs are part of coverage

Never backfill a failed run silently or drop it from planned coverage. If a retry occurs, preserve the attempt relationship. The report can use successfully completed answers for content incidence while still showing whether the intended instrument operated as planned.

Answer the questions clients actually ask

“Are we winning?”

The honest answer is not a universal yes or no.

Say where the brand appeared within the disclosed panel, where competitors recurred, which pages were cited, and whether recorded referrals or verified access changed in their own lanes. Then name the business decision: protect a recurring source, investigate an absent cluster, repair access, or improve the panel.

If the executive wants a category comparison, make clear that competitor incidence describes the same selected panel—not total assistant demand or market share.

“What’s our number?”

Give the fraction that answers the question:

The brand appeared in N of M completed answers during this window; collection completed M of P planned samples.

Then provide the engine and cluster breakdown. Do not replace this with an undisclosed 0–100 score. A single index forces hidden choices about prompt weights, engine weights, appearance, citation, sentiment, and time.

“Why did it drop?”

Often the defensible answer is: we observed a drop, but the evidence does not isolate its cause.

Check for prompt, engine, model, retrieval, coverage, classifier, and page changes. Compare the movement with the variation observed across unchanged runs. If this window moved less than the run-to-run variation the panel normally shows, say so. Do not invent a cause to make the commentary sound decisive.

“What did you change?”

Name the exact intervention, page, publication date, and hypothesis. Separate actions completed from outcomes observed. A technical repair can be complete when the page becomes fetchable even if later citation incidence does not move. A new comparison section can be published even if the sampling window is not yet complete.

The client should be able to distinguish work delivered, evidence collected, and effect not yet established.

What changes for leadership and boards

Compression changes. Evidence does not.

An executive or board version can be one slide containing:

  • collection coverage: completed out of planned samples;
  • appearance and citation incidence with denominators;
  • the largest recurring source or content gap;
  • recorded human referrals when available, with their traffic denominator;
  • material method or coverage changes; and
  • one decision requested from leadership.

The request might be approval for subject-matter review, log access, analytics cleanup, or a focused content test. It should not be “approve a higher visibility score.”

Do not turn the evidence into a single index merely because the audience is senior.

Anti-patterns to remove before sending

An undisclosed 0–100 visibility score

If inputs, weights, coverage, and method changes are not visible, the score conceals more than it communicates. Keep the underlying events.

A “rank” or “position” for an assistant answer

Generated answers do not provide a stable results list for the brand to occupy. Preserve the answer as an observation; do not promote its wording or list order into a durable property.

A percentage without its denominator

“Citation visibility rose” is not a result. State answers citing the domain out of completed answers and completed samples out of planned samples.

One sample per week charted as a trend

One changing answer can draw a dramatic line. Without replication, the chart cannot show whether the point is an ordinary alternate response. Call it a sequence of observations, not a stable trend.

Blended engines without disclosed coverage

An aggregate can rise because the sampled engine mix changed. Show which engines ran and retain engine-level results.

Referral traffic presented as citation proof

Analytics can show an attributed arrival. It usually cannot show the exact answer or citation behind it. Keep the referral and citation records separate unless a controlled journey supplies that link.

Recommendations with no observable stop rule

Every action sounds strategic when no later result could disconfirm it. Define the next sample and what would cause the team to keep, revise, or stop the work.

The report’s job

An AI search visibility analysis report is not a certificate of winning. It is a compact, auditable account of a selected panel and window: what completed, what appeared, what was cited, which competitors recurred, which people arrived, what verified infrastructure accessed, and what remains unknown.

When a client challenges the number, the report should get stronger—not collapse. That happens when the denominator, raw evidence, method version, and claim boundary were part of the deliverable from the beginning.