You should create an llms.txt file when your site has documentation, reference material, or another agent-facing surface that benefits from a short, curated index. It is inexpensive to publish and can make that material easier for a compatible agent to navigate.
You should not create one expecting it to get the site indexed by an assistant, improve search visibility, or increase citations. Those outcomes are not part of the file’s published contract, and the major crawler documentation reviewed for this article does not document using a site’s llms.txt when answering a user’s question.
The useful decision begins by separating four questions that are often collapsed into one:
- What does the llms.txt proposal actually specify?
- Who demonstrably publishes or consumes it today?
- What do major AI crawler documents say about it?
- Does it help SEO or appearance in generated answers?
Those questions have different evidence. A clear answer to one should not be stretched into an answer to the others.
See where your brand stands How we measure
Read the transcript
Should you make an llms.txt file? Only if you have documentation worth curating — and someone to keep it accurate.
It’s a short, hand-picked index of your best pages. Root of the site, or under a path if it’s scoped.
What it isn’t is a way to get indexed by an LLM. Google says you don’t need AI text files for its AI answers.
And don’t let a tool write it unchecked. A generator will hand you a tidy file full of links that don’t exist and summaries that oversell the page.
A wrong index is worse than no index — you’ve handed an agent a map to the wrong places. Fix the pages first. Keep it small. Give it an owner.
See where your brand actually stands. Free check at brandvane.ai. Takes a minute, no signup.
What the llms.txt proposal actually specifies
The llms.txt proposal describes a Markdown file that gives an agent a curated overview of a site or a section of a site. Version 2 was originally published September 3, 2024 and was last modified August 10, 2026.
The proposal calls itself a proposal. It is not a ratified web standard, an access-control mechanism, a sitemap replacement, or a documented inclusion requirement for the answer engines Brandvane samples.
Its format is intentionally small:
- an H1 containing the project or site name—the only required section;
- an optional blockquote summary;
- optional non-heading Markdown context;
- zero or more H2 sections containing Markdown links with optional notes; and
- an optional H2 named
Optionalfor material an agent can skip when it needs a shorter context.
A minimal file for a software documentation site could look like this:
# Northstar API
> Reference material for integrating the Northstar scheduling API.
Use the current v3 reference for production integrations. The migration guide
explains differences from v2.
## Documentation
- [API overview](https://example.com/docs/): Authentication, requests, and errors
- [v3 reference](https://example.com/docs/reference/): Current endpoint reference
- [Migration guide](https://example.com/docs/migrate/): Moving from v2 to v3
## Optional
- [Release notes](https://example.com/changelog/): Detailed product history
The example is a navigation aid, not a list of claims about what an assistant must index. Its value comes from honest curation: a short explanation of what the site contains, which pages are authoritative, and which material is optional.
Does llms.txt go at the site root or somewhere else?
It can go at the root or at any path within the site. The proposal says a file covers the pages beneath its path, and the most specific applicable file wins.
That means:
/llms.txtcan describe the whole site;/docs/llms.txtcan describe everything beneath/docs/; and/docs/sdk/llms.txtcan provide more specific guidance for/docs/sdk/.
If a site has one compact body of public reference material, a root file is the simplest choice. If different product areas have distinct owners, terminology, or documentation sets, scoped files can be clearer. Do not create a maze of nearly identical files merely because nested placement is allowed. Each file should reduce ambiguity for the material beneath it.
The proposal also describes clean Markdown companions for HTML pages, using forms such as page.md or page.html.md. An HTML page can advertise that representation with rel="alternate" and type="text/markdown"; rel="describedby" can describe the relationship as well. Those relations can appear in HTML <link> elements or an HTTP Link response header.
That companion-page idea is separate from the index file. You can publish a useful llms.txt without maintaining a Markdown twin for every page. If you do publish alternate representations, keep their material facts consistent with the visible page. An easier-to-fetch contradiction is still a contradiction.
Who demonstrably uses llms.txt today?
The proposal’s own version 2 page says that thousands of sites publish the file, that documentation platforms generate it automatically, and that Chrome’s Lighthouse audits for one as part of its agentic-browsing checks. It also points to AI labs publishing llms.txt files for their own developer documentation.
Those are meaningful observations about adoption around documentation and agent tooling. They are not proof of a general visibility effect.
In particular, an AI lab publishing an llms.txt for its own developer-docs site demonstrates that the lab sees value in presenting those docs in the proposed format. It does not demonstrate that the lab’s search crawler reads your file, that its consumer assistant consults the file for every answer, or that publishing the file changes the likelihood of a citation.
The same boundary applies to automatic generation. A documentation platform can produce a technically correct file because it is cheap and useful to compatible tools. The existence of the generator does not establish which external systems used the result.
When evaluating an adoption claim, ask what was actually observed:
| Observation | What it supports | What it does not support |
|---|---|---|
A site publishes llms.txt | The publisher made the file available | Any agent fetched or used it |
| A documentation platform generates it | The platform supports the proposal | Major answer engines rely on it |
| A tool audits for the file | The tool treats presence as a check | The file changes answer citations |
| A verified request reaches the file | Identified infrastructure fetched that URL in that window | The file influenced indexing or an answer |
The last row is where AI crawler analytics becomes useful. Logs can show a request to /llms.txt, subject to identity and retention limits. Even then, a fetch remains infrastructure evidence—not proof of downstream use.
What major crawler documentation says
The provider documents reviewed on August 28, 2026 describe crawler purposes, user-agent tokens, access controls, and—in some cases—published IP information. They do not document consuming a site’s llms.txt file when answering user queries.
- OpenAI’s crawler documentation distinguishes
OAI-SearchBot,GPTBot,ChatGPT-User, andOAI-AdsBotand treats their controls independently. - Anthropic’s crawler article distinguishes
ClaudeBot,Claude-User, andClaude-SearchBotand explains the visibility trade-offs of blocking user-directed and search access. - Perplexity’s crawler documentation distinguishes
PerplexityBotfromPerplexity-User, documents independent settings, and publishes network information for WAF configuration.
This is scope-limited review language. Documentation can change, and the absence of a statement on these pages is not proof that no internal or future system will ever consult the file. It does mean a publisher should not turn the current crawler documents into a promise they do not make.
It also means llms.txt is not the place to allow or block these crawlers. Access preferences belong in robots.txt and, when enforcement is required, in edge or server controls. The practical guide to blocking AI crawlers keeps those mechanisms and their costs separate.
Will an llms.txt file help SEO?
There is no defensible basis here for promising an SEO improvement.
Google’s AI-features guidance says there are “no additional requirements” or special optimizations needed to appear in AI Overviews or AI Mode. It also says publishers do not need to create new machine-readable files, AI text files, or special markup for those features.
That is a Google Search statement about Google Search features. It is not a statement about OpenAI, Anthropic, or Perplexity, and Google AI Overviews and AI Mode are not Brandvane sampled engines. It does answer the Google-specific SEO question directly: Google does not present an AI text file as an eligibility requirement for those features.
For other answer systems, the reviewed provider crawler pages establish useful access choices but no llms.txt outcome. The honest conclusion is therefore narrower than the usual tactic pitch:
- publishing the file makes a proposed, structured resource available;
- a compatible agent may find that resource useful;
- current reviewed crawler documentation does not establish an answer-inclusion benefit; and
- no observed benefit should be claimed until it is measured in a disclosed design.
Do not describe the file as a way to “get indexed by an LLM.” Indexing, retrieval, answer generation, citation, and human referral are different events. A file’s availability alone does not reveal which of those events occurred.
What to do before you create the file
First make sure the underlying pages deserve to be curated. An llms.txt that points to thin, contradictory, gated, or JavaScript-dependent material does not repair the material.
Review whether:
- important pages return useful HTML at stable canonical URLs;
- the opening copy answers the named question directly;
- specifications, dates, prices, and definitions agree across pages;
- primary evidence appears on or is linked from the relevant page;
- internal links connect overview, detail, and decision pages; and
- crawler access choices match the publisher’s actual search and training policy.
Those checks have observable outputs: a successful response, visible HTML, reconciled facts, a traceable source, or a documented access rule. They do not promise a citation. The testable AI-citation checklist shows how to turn a content or access gap into a before-and-after test without guaranteeing the result.
How to publish one without creating maintenance debt
Keep the first version small. Choose the pages a knowledgeable maintainer would hand to an agent trying to understand the site, not every URL in the sitemap.
Use a simple release checklist:
- put the file at the narrowest path that matches its scope;
- include the required H1 and a concise, factual summary;
- group links by task or subject rather than by marketing priority;
- use canonical, successful URLs;
- omit private, obsolete, duplicate, or access-controlled material;
- label the optional section deliberately;
- assign an owner and review the file when referenced pages move; and
- confirm that its contents do not conflict with
robots.txtor site policy.
A stale curated index can be worse than no index because it sends a compatible agent toward retired or contradictory material. If no one owns the update path, defer publication until the file can be maintained with the documentation it describes.
How to find out whether it did anything
Treat publication as a test with an uncertain outcome, not as a launch announcement about visibility.
Before publishing, freeze a buyer-question panel and record:
- exact prompt wording and version;
- sampled engine and disclosed mode;
- planned and completed answers;
- brand appearances out of completed answers;
- own-domain citations out of completed answers; and
- exact cited URLs.
The free AI visibility check will not serve as that baseline—three prompts on one engine is too thin to detect a change—but it does show the appearance-and-citation format you will be recording, before you build the fuller panel.
After publishing, wait for the review window you defined, repeat the same independently designed panel under comparable conditions, and report the later incidence with its denominator. Keep failed or refused runs in the coverage record.
Brandvane’s current subscription fresh checks return one on-demand OpenAI API answer per approved question; they do not schedule this experiment or include a multi-engine panel. Repetitions require separate approvals and fresh-check units. Provider-recorded mentions remain a distinct evidence source. See current AI features.
If the site appeared in N of M completed answers before and X of Y afterward, that is the observation. It is not automatically evidence that llms.txt caused the change. Model, retrieval, web, and product changes can move the sample. The Brandvane measurement method explains why replication reduces the influence of one unusual answer without making the estimate universal.
Logs can provide a separate observation: whether verified infrastructure requested the file during the window. Do not combine that fetch with answer incidence into a causal story. The file could be fetched without influencing an answer, and an answer could cite the site through another retrieval path.
For the evidence framework, see how Brandvane measures AI visibility. Automated weekly AI reports for ChatGPT and Google AI Overviews are coming soon; the planned reports will keep the two surfaces separate and show completed-check counts.
The decision
Publish an llms.txt when you have a real agent-facing information architecture to curate and can maintain it cheaply. Put it at the root for site-wide guidance or beneath a path for scoped guidance; when several files apply, the most specific one governs the proposal’s intended scope.
Do not publish one expecting a measurable change in citations, visibility, rankings, or traffic. Do not let it displace fetchable HTML, deliberate crawler permissions, primary evidence, consistent facts, or a sampling design that can show what actually happened.
The file is a modest interface. Treating it modestly is not pessimism—it is the difference between implementing a useful proposal and promising an outcome the available evidence cannot support.