What is llms txt? An honest guide to the llms.txt proposal
If you searched for llms txt, here is the short answer: llms.txt is a proposed Markdown file, placed at the root of a website, that gives AI tools a short, curated map of the site's most useful pages. It was suggested by Jeremy Howard in September 2024 and is documented at llmstxt.org. It is a community proposal, not an official web standard, and no major search engine or AI company has announced that it uses the file to rank or cite websites.
Why does it get so much attention? You spend weeks crafting a piece of content. It ranks reasonably well on traditional search engines, but when you ask ChatGPT or Claude a question about your niche, your brand is nowhere to be found, or the answer repeats older, less accurate information. You are no longer optimizing only for human readers and Google's indexer; you also need to think about how to optimize for AI search, a field often called generative engine optimization.
For years, we relied on standard technical SEO practices to ensure visibility. We built XML sitemaps and meticulously managed our crawl budgets. But Large Language Models (LLMs) consume data differently. They do not want to parse complex HTML, nested JavaScript menus, or visually heavy layouts. They want raw, semantic text and clear instructions on what matters most. This friction between how we build websites for humans and how AI models read them is what the llms.txt proposal tries to address. This guide breaks down what the file is, what it can and cannot do, how to write it in the correct format, and how to check in your own server logs whether anything actually reads it.
What is llms.txt?
llms.txt is a proposed Markdown file placed at the root of a website (for example https://example.com/llms.txt) that summarizes what the site is about and links to its most important pages, so that language models and AI tools can find clean, relevant context quickly. It is a suggestion published at llmstxt.org by Jeremy Howard of Answer.AI in 2024, not a standard approved by any standards body.

The idea came from developers who were frustrated that AI coding assistants struggled to read complex project documentation. According to llmstxt.org, thousands of sites publish the file, including the developer documentation of OpenAI, Anthropic and Google's Gemini API. Publishing a file for your own docs is not the same as using other sites' files to rank them, though, and publishing it is voluntary, and reading it is voluntary for every AI provider too. While we have long used specific files to control bot behavior, this format serves a different purpose.
To understand its role, you must look at how it differs from the established standards. The most common confusion lies in the robots.txt vs llms.txt debate.

| Feature | robots.txt | llms.txt |
|---|---|---|
| Primary Function | Access control (Allow/Disallow) | Content curation and context mapping |
| Status | Long-established convention, published as RFC 9309 in 2022 | Community proposal from 2024, not an official standard |
| Format | Specific directive syntax (User-agent, Disallow) | Markdown: H1 title, blockquote summary, H2 sections of links |
| Target Audience | All crawlers, including AI crawlers such as GPTBot | AI tools and language models that choose to read it |
| Outcome | Asks crawlers to stay out of specific paths | Offers a curated reading list; no guarantee it is used |
Think of your website as an expansive public library. The robots.txt file is the security guard at the front door holding a clipboard, telling visitors which rooms are strictly off-limits and which are open to the public. The llms.txt file, on the other hand, is the head librarian standing at the reference desk, handing out a carefully curated reading list of the most important, accurate, and easily readable books in the building.
The Meaning and Purpose in the AI Era
The proposal targets a real bottleneck. When an AI tool fetches a standard web page, it has to strip away headers, footers, sidebars, inline styles, and navigation menus to find the core text, and a language model's context window is too small to hold a whole website. A short, Markdown-formatted directory at the root offers a shortcut to the pages that matter. Whether a given AI system takes that shortcut is up to its provider.

In the broader picture of technical SEO, this file sits downstream from your foundational architecture. You still need proper robots.txt configuration rules to manage crawl budgets and secure administrative areas. You still need XML sitemaps for discovery. The new file is an optional, additive layer. If you are deciding where it fits next to classic SEO, the overview of how GEO differs from SEO is a useful companion.
The Difference Between Search AI and Standalone LLMs
A major gap in understanding this concept is confusing traditional search engines that use AI with standalone language models. They operate on entirely different paradigms.

Search AI, like Google's AI Overviews, is built directly on top of Google's massive, existing search index. When a user queries Google, the AI Overview is generated by synthesizing information from pages that Google has already crawled, indexed, and ranked using traditional signals. Therefore, optimizing for AI Overviews means leaning heavily into established best practices: securing backlinks, establishing topical authority, and implementing structured JSON-LD schema.
Standalone assistants, like ChatGPT or Claude, work differently. When you use advanced ChatGPT prompting workflows to ask it to research a company, it may fetch live pages on your behalf. OpenAI documents separate user agents for this: GPTBot for training data, OAI-SearchBot for search, and ChatGPT-User for fetches triggered by a user. If such a fetch hits a complex, JavaScript-heavy page, it might fail to extract the right context. None of these providers has publicly stated that it reads llms.txt automatically. Where the file clearly helps today is when a developer or a tool deliberately points an AI at it, for example by pasting the llms.txt URL into a coding assistant or an IDE plugin.
So the realistic framing is this: publishing the file is a low-cost way to hand AI tools a clean summary of your site if they look for one. It is not a switch that makes AI models describe your brand correctly.
When You Do Not Need It Yet
Despite the hype, you do not need an llms.txt file if your website is a simple, five-page local business portfolio (like a local plumber or bakery). If your site lacks deep, informational content, technical documentation, or extensive blog archives, creating this file is a waste of time. AI bots can easily parse a basic about us page without a markdown map. In these scenarios, focus your resources on local SEO and basic on-page optimization instead of worrying about advanced LLM ingestion protocols.
Evaluating the Value and Tangible Benefits
Publishing the file can add some value, mostly around narrative control and internal efficiency, as long as you keep expectations realistic.
Business Value: Brand Safety and Narrative Control
For a business, the primary value is narrative control. When language models synthesize answers about your brand, products, or industry, they rely on the clearest data available. If your competitor publishes clear, well-structured explanations of their product and your site only offers convoluted marketing copy trapped in complex HTML, an AI tool that reads both will find it easier to quote the competitor. The fix starts with clear pages; llms.txt only points to them.

Used that way, the file supports brand safety. It gives any tool that looks for it a canonical list of your current, accurate pages, instead of leaving it to stumble on third-party reviews or outdated press releases.
Illustrative example: An enterprise SaaS company selling cybersecurity software.
- Context: They noticed ChatGPT was consistently describing their flagship product using outdated feature lists from three years ago, leading to confused enterprise leads.
- Actions taken: They created an llms.txt file pointing directly to a clean markdown version of their current product spec sheet and their core whitepapers. They strictly defined their current capabilities.
- Hurdles: Initially, they included links to complex, gated PDF reports, which the bot could not read, resulting in failed ingestion errors in their logs. They fixed this by replacing the PDF links with public HTML summary pages.
- Results: The file alone did not change what ChatGPT said. What helped was the work it forced: rewriting the public spec pages so they were current and easy to parse. Over the following months, AI answers that cited those pages became more accurate, while answers based on older third-party articles still needed separate outreach.
Operational Benefits: Internal Efficiency
For the individuals executing technical marketing, the most concrete benefit is internal. When you build Retrieval-Augmented Generation (RAG) systems or custom AI agents for your own company, a single, curated entry point reduces development work. You do not need complex scraping scripts; you point your internal tools at your own root file and the Markdown pages it lists.
| Possible Benefit | What to Measure | Honest Caveat |
|---|---|---|
| Easier reading by AI tools | Requests to /llms.txt and the linked .md pages in server logs | A request shows a fetch, not that the content was used in an answer |
| More accurate brand answers | Spot checks of AI answers about your brand, repeated monthly | Many other sources influence the answer |
| Faster internal AI projects | Time to connect a RAG tool or assistant to your docs | Only applies if you build such tools |
Managing Scraping Risks and Content Protection
While visibility is valuable, there is a massive hidden risk: unauthorized content scraping. The very nature of this file—gathering your most valuable, high-signal content into one easily digestible map—makes it incredibly easy for competitors to steal your intellectual property. If you put links to your proprietary research, paywalled articles, or unique methodologies into this file, you are essentially handing over your business moat to anyone with a basic python script.

The strategy here must be defensive. You must deliberately choose what to expose and what to hide. Use the file to highlight top-of-funnel marketing content, feature summaries, and public API documentation. Deliberately exclude links to premium content, deep proprietary research, or anything that gives you a unique competitive advantage. You want the AI to know what you do, not exactly how you do it.
Discover Orova.vn – a Biz AI Agent platform with OROVA SEO, a complete solution for every website. The system supports search engine optimization from A to Z with features including keyword research, writing new SEO-ready articles, optimizing existing content, rank tracking, plus competitor analysis and in-depth technical analysis. Sign up today to experience OROVA SEO completely free (offer valid through July 7, 2027).
How It Works: A Technical Breakdown
Understanding how to implement this requires looking under the hood at the file's anatomy, how to verify its usage, and how to structure it for different business models.
Anatomy of the File
The proposal asks for a plain Markdown file named llms.txt, usually served at the root path (for example https://yourdomain.com/llms.txt). The current version of the proposal also allows a file at a subpath, such as /docs/llms.txt, which then covers only the pages under that path. Unlike many config files, it has no YAML front matter. The sections must appear in this order:

- An H1 heading with the name of the site or project. This is the only required section.
- A blockquote (a line starting with >) holding a short summary with the key information needed to understand the rest of the file.
- Optional free text: zero or more paragraphs or lists with extra details, but no headings.
- Zero or more H2 sections, each a list of links. Every list item is a Markdown link name, optionally followed by a colon and a short note about the page.
- By convention, a final section titled ## Optional for secondary links that a tool can skip when it needs a shorter context.
A minimal, valid file looks like this:
# Example Coffee Roasters > Example Coffee Roasters sells single-origin coffee beans online and publishes brewing guides for home baristas. All prices and stock levels change often; check the shop pages for current details. ## Guides - [Pour-over brewing guide](https://example.com/guides/pour-over.md): Step-by-step method, grind size and water ratio - [Espresso basics](https://example.com/guides/espresso.md): Equipment, dosing and common mistakes ## Shop - [All coffee beans](https://example.com/shop/beans): Current single-origin range by roast level ## Optional - [Company history](https://example.com/about): Background on the roastery
The proposal also suggests two companions. Pages can offer a clean Markdown copy at the same URL, either with .md appended (guide.html.md) or with the extension replaced (guide.md); URLs without a file name use index.html.md or index.md. Some sites additionally publish llms-full.txt, a single file containing the full text of the key pages; that variant is a common convention among documentation tools rather than part of the core format.
Server Log Analysis: Proving Bots Read Your File
One of the biggest gaps in current SEO advice is the lack of verification. Anyone can upload a text file, but how do you know if it is actually providing value? You cannot rely on Google Analytics or Google Search Console, because AI crawlers often do not execute the JavaScript required to trigger traditional analytics pixels.

You must look at your raw server access logs. OpenAI, Anthropic and Perplexity each publish the user-agent names of their crawlers, which is what you search for in the logs. Keep in mind that a user-agent string can be faked, so the providers that publish IP ranges also let you verify the source address.
If you are using a Linux server (like Ubuntu) running Apache or Nginx, you can use command-line tools to query your logs. The grep command allows you to search through massive text files for specific patterns.
To find all requests to your new file, you would SSH into your server and run a command targeting your access log. For Nginx, the command looks like this: grep "GET /llms.txt" /var/log/nginx/access.log
This command filters the log to only show lines where a bot requested that specific URL. However, you want to know which bots are reading it. You need to look for specific AI user agents, such as GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, or PerplexityBot. To show only one bot:

grep "GET /llms.txt" /var/log/nginx/access.log | grep -i "GPTBot"
To count hits per AI crawler, extract the known bot names with grep -o, then pipe the result to sort and uniq:
grep "GET /llms.txt" /var/log/nginx/access.log | grep -o -i -E "GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|PerplexityBot" | sort | uniq -c
Illustrative example: A technical documentation startup.
- Context: The technical SEO lead implemented the file but faced pushback from the engineering team who claimed it was a waste of time and wouldn't be read.
- Actions taken: The SEO lead accessed the server via SSH. They ran a complex grep command piped to awk to extract just the User-Agent strings from hits on the /llms.txt path over a 7-day period. They compiled the output into a simple CSV.
- Hurdles: The main access.log file was over 50GB, and running a raw grep command crashed their terminal session. They fixed this by running the command against the rotated, compressed logs (zgrep) from the previous day instead of the live log.
- Results: The report showed that several AI crawlers had fetched the file during the week, while some popular assistants never requested it. That settled the argument in a useful way: the file cost almost nothing to maintain, but it could not be credited for any change in AI visibility on its own.
How to Structure the File for Different Business Models
There is no one-size-fits-all approach. The structure of your markdown links must reflect the nature of your content. Here are templates for different website types; each keeps the required order (H1, blockquote, then H2 link lists) and uses the optional : note after a link where it adds context.

The E-commerce Template
For e-commerce, AI bots do not need to read 10,000 individual product pages. They need to understand your categories, your brand positioning, and your distinct value proposition.
# Brand Name E-commerce Map > This file provides a structured overview of our product categories and brand guidelines. ## About Us - [Brand History and Mission](https://domain.com/about) - [Sustainability Practices](https://domain.com/sustainability) ## Core Product Categories - [Running shoes](https://domain.com/category/running): Road and trail running shoes, sizing advice - [Hiking gear](https://domain.com/category/hiking): Boots, packs and outdoor accessories ## Customer Policies - [Shipping and Returns Policy](https://domain.com/policies)
The goal here is high-level semantic understanding, utilizing semantic search optimization techniques to ensure the AI grasps the broad entities your store represents.
The SaaS and Software Template
SaaS companies have the clearest use for the format, because developers often paste documentation into AI coding assistants to write code or integrate APIs. A curated list of the right reference pages saves them from pasting the wrong ones.
# SaaS Platform Documentation > Technical documentation and API references for our platform. ## Getting Started - [Platform Overview](https://domain.com/docs/overview) - [Quickstart Guide](https://domain.com/docs/quickstart) ## API Reference - [Authentication Endpoints](https://domain.com/docs/api/auth) - [Data Ingestion Webhooks](https://domain.com/docs/api/webhooks) ## Architecture - [Security and Compliance (SOC2)](https://domain.com/security) ## Optional - [Changelog](https://domain.com/changelog): Release history
Notice how direct these links are. They point to the most dense, factual pages on the site, allowing the bot to ingest the technical specifications without getting lost in marketing fluff.
The Blog and Publisher Template
For content-heavy sites, listing every article is impossible and counterproductive. You must curate your cornerstone content.
# Publisher Name Cornerstone Content > A curated list of our most comprehensive research and editorial standards. ## Editorial Guidelines - [Fact-Checking Policy](https://domain.com/editorial-standards) ## Cornerstone Research - [The 2026 State of Digital Marketing Report](https://domain.com/reports/digital-2026) - [Comprehensive Guide to AI SEO](https://domain.com/guides/ai-seo) ## Main Categories - [Technical SEO Archives](https://domain.com/category/technical) - [Content Strategy Archives](https://domain.com/category/strategy)
The Agency Template
Service businesses need to highlight their methodologies and case studies to build trust with AI synthesis engines.
# Agency Name Services Map > Overview of our consulting methodologies and service offerings. ## Our Approach - [The Proprietary Growth Framework](https://domain.com/methodology) ## Core Services - [Enterprise SEO Consulting](https://domain.com/services/enterprise-seo) - [Conversion Rate Optimization](https://domain.com/services/cro) ## Verified Case Studies - [Client A case study](https://domain.com/case-studies/client-a): Technical SEO project for a retail brand
With OROVA.VN and the OROVA SEO module, you put an end to the exhausting days of manual work for good. Instead of struggling for hours to write articles and compile reports, the entire process is now optimized and completed in just 5 minutes.
What to Do Now: Practical Steps by Role
Adapting requires different strategies depending on your role and resources. You must prioritize actions that yield the highest clarity for AI crawlers without exposing your proprietary data.
Small Business Owners
If you manage your own small business website, your approach should be minimal but precise.

- Identify your core truth: Write down the 3-5 pages on your site that perfectly explain what you do, who you serve, and why you are different.
- Create the basic file: Open a plain text editor, write a brief markdown heading, and paste the links to those 3-5 pages.
- Upload to root: Place the file in the public HTML folder of your server (where your robots.txt lives).
- Do not overcomplicate: Avoid trying to list every blog post. Keep it strictly to foundational business information.
Content Managers and Editors
For those managing larger content operations, the task is about integration and process updates.
- Audit cornerstone content: Identify the most critical, high-traffic informational pages.
- Strip the formatting: If possible, create alternate, markdown-only versions of these core pages (e.g., domain.com/guide.md) to link to within your root file.
- Establish a curation process: Make updating the file a quarterly task. When a massive new piece of research is published, add it to the directory.
- Implement risk reviews: Before adding any link to the file, have a stakeholder confirm that the content is safe for public, unlimited AI ingestion.
SEO Agencies and Consultants
Agencies must view this as a new technical deliverable and an analytical challenge.
- Add to technical audits: Include the presence and optimization of this file in every initial client audit.
- Setup log monitoring: Create scripts to regularly parse client server logs for AI user-agents hitting the file, providing concrete reporting on AI visibility.
- Map entities: Use the file to reinforce the semantic entities the client wants to own in their space.
- Educate clients on scraping: Clearly explain the risks of over-exposure and help them build a defensive content strategy.
| Common Mistake | Immediate Consequence | How to Avoid It |
|---|---|---|
| Linking to paywalled content | Competitors use bots to extract premium data | Keep premium URLs out of the directory |
| Using complex HTML in the file | Bots fail to parse the document | Stick strictly to plain text Markdown |
| Listing hundreds of low-value links | Crawlers dilute their focus | Curate only the top 5% of cornerstone content |
Illustrative example: A mid-sized digital marketing agency.
- Context: The agency wanted to position themselves as thought leaders in AI marketing, but AI models rarely cited their blog.
- Actions taken: They created an llms.txt file but made the mistake of listing all 500 of their blog posts in chronological order.
- Hurdles: Server logs showed GPTBot accessing the file, but subsequent AI queries still ignored their content because the file was too large and lacked hierarchical structure. They fixed this by deleting 90% of the links, leaving only their 10 most comprehensive, authoritative pillar pages organized by category.
- Results: The shorter file was easier for their own team to keep current, and the cleanup exposed several thin pillar pages that they rewrote. Over the next quarter, those rewritten pages started appearing more often in AI answers, which the team attributed mainly to the better pages rather than to the file. For the bigger picture of how these tactics fit together, see how SEO, AEO and GEO work as one strategy.
Author's Predictions for the Next Few Years
The intersection of technical SEO and AI ingestion is evolving rapidly. As of 2026, we are still in the early, experimental phases of how websites communicate with language models. Based on current trajectories, I believe the landscape will shift dramatically in three specific ways.
A Shared Convention, but Slowly
Right now, this file is a grassroots proposal that AI companies have not formally committed to. I think some form of machine-readable site summary for AI tools is likely to settle into a shared convention eventually, the way robots.txt did over many years, but it may not be llms.txt in its current shape, and I do not expect sites without one to be penalized. My advice is to treat it as a cheap, low-risk experiment, not a mandatory piece of your technical stack.
The Rise of Monetized Ingestion and Paywalls
I lean toward the view that files like this will become part of how publishers negotiate access with AI companies. Today we use it to give context away for free. I suspect publishers will increasingly want a machine-readable way to state licensing terms or paid access for AI crawlers, and a root-level file is a natural place for that. Whether it grows out of llms.txt or a separate convention is still open.
Native CMS Integration as a Default Feature
Today, creating this file usually means manual editing or a plugin. I expect content management systems to offer it as a built-in option sooner or later, much as they generate XML sitemaps today, so that marking a post as cornerstone would add it to the file. Until then, the main advantage of managing it by hand is the discipline of deciding which pages truly represent your business.
However, this assumes AI models will continue to rely on external crawling rather than closed, proprietary datasets, which could shift depending on copyright legislation.
Frequently Asked Questions
Is llms.txt an official web standard?
No. It is a proposal published at llmstxt.org in 2024. It has no approval from a standards body, and AI providers are free to ignore it. Treat any claim that it is required, or that it directly boosts rankings, with caution.
Is llms.txt a replacement for robots.txt?
No. They serve completely different functions. robots.txt is an access tool used to ask crawlers to stay out of specific areas of your server. llms.txt is a curation tool used to guide crawlers toward your best content. Many sites publish both, and they do not conflict.
Do I still need this file if search engines already use AI?
It is optional either way. Google's AI features rely on its existing search index, and Google has not announced any use of llms.txt in Search. Standalone assistants such as ChatGPT or Claude can fetch live pages when users ask research questions, but their providers have not confirmed reading the file automatically. Its clearest use today is giving developers and AI tools a ready-made summary when they look for one.
Will this file guarantee my brand is mentioned by ChatGPT?
Absolutely not. At best it makes your key pages easier to find for a tool that chooses to read it. Whether the AI actually mentions your brand depends on the authority, uniqueness, and relevance of the content you provide, not just the presence of the file itself.
Can this file hurt my SEO rankings on Google?
There is no known mechanism by which it would. It is a small text file that Google has not said it uses for ranking. Just make sure you do not accidentally block important pages in your robots.txt while setting it up, and avoid creating thin duplicate .md pages that compete with your HTML pages for indexing.
How often should I update the links inside the file?
Treat it like a curated reading list. You do not need to update it every time you publish a minor blog post. You should update it quarterly, or whenever you publish a massive, authoritative piece of content (like a yearly industry report or a major new product feature) that you explicitly want AI models to learn from.
Where to Start?
Jumping into AI optimization can feel paralyzing. Your next step depends entirely on the current state of your technical SEO and content architecture. Do not try to execute everything at once; follow a single, logical next step.
If you have nothing in place (no file, no strategy): Your immediate goal is simply establishing a presence. Open a basic text editor today. Write a two-sentence description of your company, paste the URLs to your homepage, your 'About' page, and your primary product/service page. Format it as an H1 title, a blockquote summary and an H2 list of links, save it as llms.txt, and upload it to your root directory. It is a short task, and it gives you a baseline you can measure.
If you have a file, but it is fragmented (too many links, messy structure): Your task is aggressive curation. In your next work session, review the file and delete at least 50% of the links. Remove anything that isn't absolute, foundational truth about your business. Ensure every remaining link points to a page that is text-heavy and logically structured. AI favors depth and clarity over sheer volume.
If you have implemented it, but are unmeasured (you don't know if it works): Your next step is technical verification. You must gain access to your server logs. Spend one afternoon learning the basic grep command. Download yesterday's access log and run a search for the file path. If you cannot prove that bots are hitting the file, you cannot justify spending further resources optimizing it. Measurement must precede expansion.
Run your business with AI Agents
Orova is the always-on Biz AI Agent — it plans, runs, and optimizes the work for you.
Save time, unlock productivity.