AI crawler analytics starts with a narrow question: did identifiable AI infrastructure request a page, and what did the site return?
Server and edge logs can answer that question. They can show a time, source address, user-agent string, URL, status code, response size, and sometimes a CDN’s verified-bot classification. With the right fields, they can distinguish a successful request from a block, redirect loop, rate limit, or origin error.
They cannot show that the page appeared in an AI answer. They cannot show that a person saw the answer. And a 200 response does not, by itself, show that a crawler executed the page’s JavaScript or extracted its important content.
Those boundaries are why Brandvane treats verified crawler access as the Do lane in its three-signal AI visibility method: evidence about what infrastructure did, separate from what sampled answers said and which AI-referred people arrived.
If the job is to change access rather than measure it, the companion guide to blocking AI crawlers maps search, training, and user-directed agents to the controls and trade-offs each provider documents.
An llms.txt file is a different mechanism again: the guide to whether you should create an llms.txt file explains what that proposal specifies without treating publication as evidence of crawler use.
What verified crawler access means
A verified crawler-access record joins three pieces of evidence:
- Identity: the request matches more than a self-declared user-agent string.
- Request: a server, CDN, or edge service recorded the URL and time.
- Outcome: the infrastructure recorded the response status and, ideally, bytes served, cache action, and security action.
“Verified” needs to describe the identity check actually performed. Depending on the provider and your infrastructure, that may be a current provider-published IP range, a forward-confirmed reverse-DNS procedure, a cryptographic bot signature, or a CDN’s retained verified-bot field.
A log row that only says GPTBot or ClaudeBot is a candidate bot request, not verified identity. Anyone operating an HTTP client can put a familiar name in the User-Agent header.
Cloudflare’s current verified-bot documentation describes published IP lists, reverse DNS, and Web Bot Auth as identity methods. Its IP-validation reference also couples the network check with a user-agent pattern. That is a stronger standard than trusting the header alone.
Keep crawler purposes separate
One provider may operate different bots for different jobs. Pooling them into “AI traffic” erases the reason for the request.
As documented on August 28, 2026:
- OpenAI uses
OAI-SearchBotfor search discovery,GPTBotfor potential model-training collection, and user-initiated agents for requests made on a person’s behalf. OpenAI publishes current network ranges for some crawlers and tells publishers to use those ranges or a provider-level verified-bot system. Its crawler documentation also treats search inclusion and training controls as separate choices. - Anthropic documents
Claude-SearchBot,ClaudeBot, andClaude-Useras separate search, model-development, and user-directed agents. Anthropic’s crawler guidance says it does not currently publish IP ranges because the bots use service-provider public IPs. A raw origin log therefore may not contain enough evidence to independently verify an Anthropic identity; a retained CDN verification signal can be stronger. - Perplexity distinguishes
PerplexityBot, used to surface and link sites in search, fromPerplexity-User, which supports a user’s request. Its crawler documentation publishes separate IP-range files and recommends combining those ranges with the matching user agent.
These names, ranges, and policies can change. Store the provider documentation URL and the date on which the verification rule was applied. Do not silently apply today’s IP list to an old request and imply that the old range was valid at the time.
How to check AI crawler access in server or edge logs
The exact query depends on the host, but the evidence workflow is portable.
1. Preserve enough fields
Export or retain, where available:
- timestamp and time zone;
- client IP or the trusted edge’s original-client field;
- full user agent;
- HTTP method, host, path, and query;
- response status and bytes served;
- cache status and origin status;
- WAF, bot-management, rate-limit, or challenge action;
- verified-bot name, category, or detection identifier; and
- request ID for tracing the same event across edge and origin.
Privacy and retention rules still apply. A crawler investigation rarely needs cookies, form bodies, or other user data. Collect the smallest log surface that can support the access claim.
2. Find candidates without declaring them verified
Filter the full user-agent field for documented bot tokens such as OAI-SearchBot, GPTBot, Claude-SearchBot, ClaudeBot, PerplexityBot, and their user-directed counterparts.
Call the result candidate requests until the identity step is complete. Keep unmatched and failed-verification rows available for security review, but do not add them to verified crawler totals.
3. Verify identity using the strongest available signal
For a provider with a published IP file, compare the source address with the current documented CIDR ranges and require the corresponding user-agent token. Automate range refreshes if this becomes an ongoing control; copying a list once creates a verification rule that quietly decays.
If a provider documents reverse-DNS verification, perform both the reverse lookup and the forward lookup back to the original address. A PTR hostname alone is not sufficient.
If your CDN supplies a verified-bot or signed-agent field, retain that field in the export and document what the CDN’s classification means. Do not reconstruct a “verified” label later from the user agent if the original security metadata was discarded.
When no provider-supported network or platform signal is available, the honest result is identity not independently verified. That is missing evidence, not proof that the request was fake.
4. Classify the access outcome
Group verified requests by outcome before counting URLs:
| Observed result | Defensible reading | Investigation |
|---|---|---|
200 with a plausible response size | The server or edge returned a successful response | Inspect the actual response body and JavaScript dependency |
301 or 302 | The requested URL redirected | Follow the chain and verify the final target |
401 or 403 | Authentication, authorization, WAF, or bot policy denied access | Review the matching security event and robots policy |
404 or 410 | The requested resource was unavailable at that URL | Check stale links, canonicals, and intended retirement |
429 | A rate limit applied | Inspect crawl rate and rule scope before widening access |
5xx | The edge or origin failed during the request | Trace the request ID and server health |
A response served from cache is still an observed response. Note it, because the origin may have no matching row. Conversely, an origin log may contain only requests the edge forwarded, omitting crawler requests the CDN blocked or answered itself. Edge and origin logs describe different checkpoints.
5. Report a dated window, not a timeless status
A useful crawler-access statement looks like this:
During [start date]–[end date], the retained edge logs contained [N] requests verified as [provider and bot] across [U] unique URLs. [S] returned
2xx; [R] redirected; [B] were blocked or rate-limited. Identity was checked using [method] against [source and verification date].
The placeholders matter. A count without a time window or verification method is hard to audit. “AI crawlers can access the site” overstates what one request to one page can establish.
Can AI crawlers execute JavaScript?
There is no safe universal “yes.” Crawler, search fetcher, user-directed agent, and browser agent are different instruments, and providers can change their infrastructure.
The OpenAI, Anthropic, and Perplexity crawler pages reviewed on August 28, 2026 describe access controls, bot purposes, and in some cases IP verification. We did not find a public promise on those pages that every named crawler executes client-side JavaScript like a full browser, waits for asynchronous requests, accepts consent state, or exposes the resulting DOM to the same downstream system.
Google is a useful contrast, not evidence about another provider. Google explicitly documents a separate rendering stage that runs JavaScript, while also warning that not all bots can run JavaScript. OpenAI’s current advertiser guidance separately warns that JavaScript challenges and human-verification logic can block automated crawlers. Neither fact establishes how an AI crawler processed your page.
Treat JavaScript support as unverified for the crawler and path you care about unless the provider documents it or you can observe stronger evidence.
A successful fetch does not prove the crawler saw the page
Consider a client-rendered product page whose initial response contains only:
<div id="app"></div>
<script src="/assets/app.js"></script>
The server can return 200 and record thousands of bytes. The product name, price, comparison copy, links, and structured data may still depend on a later script request and API call.
From the first log row alone, you do not know whether the requester:
- downloaded the script;
- executed it successfully;
- waited for network calls;
- received the API response;
- accepted or bypassed a consent gate;
- rendered content hidden behind interaction; or
- used any retrieved text in an index or answer.
Separate requests for JavaScript or API resources can strengthen an access diagnosis, but they still do not reveal the final DOM or prove downstream use. A headless-browser test that you run yourself shows what your test browser rendered, not what a provider’s crawler rendered.
Test the content path, not just the status code
For each commercially important URL:
- Fetch the page without executing JavaScript and save the response HTML.
- Check whether the title, canonical, main heading, core explanatory copy, primary links, and structured data are present in that response.
- List the resources required to produce any missing content.
- Check logs for verified access to those resources, while preserving the same uncertainty about execution.
- Test blocked cookies, empty local storage, no authentication, and a fresh session.
- Move essential public content into static HTML, server rendering, or prerendered output when feasible.
This is not a recommendation to build a different page for bots. The goal is equivalent, accessible content in the initial response, not cloaking. Progressive enhancement is the safer default: JavaScript can improve the experience without being the only place the evidence a buyer or crawler needs exists.
What crawler analytics can and cannot support
Verified logs can support statements such as:
- a named, verified bot requested a named URL during a stated window;
- the edge allowed, redirected, challenged, rate-limited, or blocked that request;
- the response status and size recorded at that checkpoint; and
- repeated failures cluster around a path, security rule, or response type.
They do not support statements such as:
- the page was indexed;
- the content was understood;
- JavaScript executed;
- the page was cited or recommended;
- a sampled answer came from that fetch; or
- a human visited because of it.
For the evidence framework, see how Brandvane measures AI visibility. Automated weekly AI reports for ChatGPT and Google AI Overviews are coming soon; the planned reports will keep the two surfaces separate and show completed-check counts.
Brandvane’s current SEO and Pro+ plans do not include automated crawler-log ingestion. Today, customers can supply suitable logs for a hand-built audit. If no suitable log evidence is supplied, the crawler lane is reported as not connected, not estimated from citations or GA4.
The bottom line
Good AI crawler analytics does not turn a bot-shaped user agent into a visibility score. It verifies identity where the available evidence permits, keeps search, training, and user-directed agents separate, reports the actual response outcome, and tests whether essential content exists before JavaScript runs.
A verified fetch is useful infrastructure evidence. Keeping it that narrow is what makes it useful.