- Jeremy Howard proposed the llms.txt standard in September 2024 as a markdown file that tells AI crawlers which pages matter most on your site.
- According to Semrush, Search Engine Land ran their own llms.txt from March 2025 and recorded zero visits from GPTbot, PerplexityBot, ClaudeBot, or Google-Extended between mid-August and late October 2025.
- According to Ahrefs, 97% of the roughly 38,000 domains with a valid llms.txt received zero crawler requests for it in May 2026 — adoption is real, but AI readership is not yet confirmed at scale.
- The file itself takes one evening to write: 20–50 lines of markdown listing 10–30 key pages (The Prompt Bench, 2026).
- Google included llms.txt in their Agent2Agent (A2A) protocol in April 2025, which is the strongest signal yet that the standard has a future.
Introduction
The llms.txt file is a plain markdown document placed at yourdomain.com/llms.txt that lists your site's most important pages along with short descriptions — essentially a curated reading guide for AI crawlers and large language models. Jeremy Howard published the proposal on September 3, 2024, and the idea spread quickly through technical SEO circles: if robots.txt tells crawlers what to ignore, llms.txt tells AI what to prioritise.
For an e-commerce operator, the pitch is straightforward. Generative search engines — ChatGPT, Perplexity, Google's AI Overviews — increasingly answer shopping queries by synthesising content from across the web rather than returning a ranked list of links. If an AI model is deciding whether to cite your product page or a competitor's, the quality and accessibility of your structured content matters. An llms.txt file is one mechanism for making that content easier to parse.
The honest caveat: the data on whether AI systems are currently reading these files is thin. The numbers below show both the opportunity and the gap between theory and current practice.
What Exactly is the llms.txt File and Why Should E-commerce Stores Care?
The llms.txt file is a markdown-formatted text file, hosted at the root of your domain, that provides AI language models with a structured overview of your site's content hierarchy. Jeremy Howard published the proposal for llms.txt on September 3, 2024 (GreenGeeks, 2024), and the specification has been updated actively since — reaching version 1.7.0 by May 2026.
For an e-commerce store, the practical value is in curation. A typical 10,000-SKU catalogue has thousands of crawlable URLs, but an AI model building a response about "best running shoes under $100" does not need to parse all of them. Your llms.txt can point directly to your category landing pages, your size-guide content, your returns policy, and your brand story — the pages that contextualise your products for a language model building a shopping recommendation.
The analogy to robots.txt is useful but imprecise. Robots.txt is a blocking mechanism; llms.txt is an invitation and a map. Where robots.txt says "do not crawl /checkout/", llms.txt says "here are the 20 pages that best represent what we sell and why customers trust us." The two files coexist and serve different purposes.
One more structural reason e-commerce stores should care: product pages are often thin on prose. A product listing with a title, three bullet points, and a price gives a language model very little to work with when composing a recommendation. An llms.txt that links to richer editorial content — buying guides, comparison pages, use-case articles — gives AI crawlers a path to the substance behind your catalogue.
How is llms.txt Adoption Faring in the E-commerce Landscape?
Adoption is growing but remains a minority behaviour. According to SE Ranking, 10.13% of nearly 300,000 analyzed domains had an llms.txt file in 2025. A separate study cited by Ahrefs put the figure at 28% of domains studied, though that sample likely skews toward tech-forward publishers rather than the broader e-commerce population.
The traffic-tier breakdown from SE Ranking is instructive: high-traffic domains (over 100,001 visits) had an adoption rate of 8.27%, while mid-tier domains (1,001–5,000 visits) came in at 10.54%. The pattern suggests that llms.txt adoption is not yet a behaviour driven by scale or sophistication — smaller stores are adopting it at roughly the same rate as large ones, which means there is no first-mover disadvantage to waiting, but also no clear competitive signal that the biggest players have decided it is essential.
As of July 2025, only 951 domains had published an llms.txt file according to Semrush (Semrush, 2025) — a figure that predates the SE Ranking study's much larger count, reflecting how quickly the standard gained traction in the second half of 2025. Mintlify rolled out llms.txt support to thousands of sites in November 2024, which likely accounts for a large share of early technical-documentation domains in the count.
Across the stores we manage at MirandaMedia, the pattern we keep seeing is that llms.txt implementation happens in a single afternoon when someone on the team owns it — and then sits unreviewed for months, which undermines most of the potential value.
Are AI Search Engines Actually Reading llms.txt Files?
The honest answer is: not reliably, not at scale, and not confirmed by any major provider. According to Ahrefs, 97% of the roughly 38,000 domains with a valid llms.txt received zero requests for it in May 2026 (Ahrefs, 2026). That is the sharpest single number in this entire debate.
Search Engine Land implemented their own llms.txt in March 2025 and tracked bot traffic carefully. According to Semrush, the llms.txt page received zero visits from Google-Extended bot, GPTbot, PerplexityBot, or ClaudeBot between mid-August and late October 2025 (Semrush, 2025). That is a seven-week window of silence from the four bots most likely to benefit from reading the file.
The provider-side picture is equally cautious. As of mid-2026, major AI providers including Anthropic, OpenAI, Google, and Perplexity have not made public commitments to reading llms.txt. In late May 2026, Google's guide on optimising for generative AI features stated that machine-readable files like llms.txt are not needed — a direct signal from the dominant search provider that the file is not part of their current ingestion pipeline.
There is one counterpoint worth noting. There are reports that OpenAI appears to be crawling llms.txt files on live websites as often as every 15 minutes, though this is not an official confirmation from OpenAI and the crawl behaviour may reflect content discovery rather than structured llms.txt parsing. Google also included llms.txt in their Agent2Agent (A2A) protocol, launched in April 2025 (Ahrefs, 2025), which is a meaningful architectural signal even if the practical crawl data does not yet reflect it.
The practical conclusion for operators: implement the file because the cost is low and the upside is real if provider behaviour shifts — but do not treat it as a confirmed traffic lever today.
What Are the Best Practices for Implementing an Effective llms.txt File?
A well-structured llms.txt file takes one evening to build. According to The Prompt Bench, implementing a working llms.txt typically requires 20–50 lines of markdown. The file format is simple: a top-level H1 with your brand name, a short paragraph describing what your store sells, and then a series of markdown link lists grouped by content type.
A useful llms.txt file should contain between 10 and 30 key pages — not an exhaustive index of 200 entries. The goal is curation, not comprehensiveness. An AI model reading your llms.txt should come away with a clear picture of your product categories, your key editorial content, and your trust signals (returns policy, about page, reviews). A 200-link dump gives it none of that clarity.
For an e-commerce store, a practical structure looks like this:
| Section | What to include | Why it matters |
|---|---|---|
| Products | 3–5 top category pages | Orients the model to your catalogue scope |
| Editorial | Buying guides, comparison pages | Provides prose context AI can quote |
| Trust | Returns policy, about page, FAQ | Answers the questions buyers ask AI assistants |
| Technical | Sitemap URL, structured data notes | Helps crawlers find the rest of your content |
Keep descriptions under 15 words per link — enough to tell a language model what it will find, not enough to pad the file into noise. Update the file when you launch new categories or retire old ones; a stale llms.txt pointing to 404 pages actively degrades the signal you are sending.
Tools exist to accelerate the build. The Ecommerce llms.txt template from EasyLLMsTxt can be generated in about 3 seconds with their 1-Click Generator, which is a reasonable starting point before you customise for your specific catalogue and content structure.
How Does llms.txt Impact AI Visibility and Generative Search Discovery?
The llms.txt file is one input into a broader practice called Generative Engine Optimisation (GEO) — the discipline of structuring your content so that AI-powered answer engines cite and recommend your store. Research published on arXiv found that GEO techniques can boost visibility by up to 40% in generative engine responses. The llms.txt file is not the whole of GEO, but it is the most implementable single action an operator can take in under a day.
The mechanism is indirect. An llms.txt file does not directly inject your content into an AI model's training data. What it does is reduce the friction for AI crawlers to find your highest-value pages — particularly the editorial and contextual content that language models use when composing recommendations. A store with a well-maintained llms.txt pointing to a thorough buying guide is more likely to have that guide indexed and cited than a store where the same guide is buried three clicks deep in a blog archive.
Scale compounds the effect. According to TNGShopper, 100 products multiplied by 50 locations can create 5,000 unique, crawlable, AI-optimised entry points. Your llms.txt cannot list all 5,000, but it can point to the category and location index pages that allow a crawler to discover them systematically.
The relationship between llms.txt and traditional SEO is additive, not substitutive. Your robots.txt, sitemap.xml, structured data markup, and page speed all remain relevant for conventional search. The llms.txt layer sits on top of that foundation and speaks specifically to the retrieval behaviour of large language models — which increasingly operate on a "find the most useful structured page, summarise it, cite it" loop rather than a ranked-list model.
What is the Current Status and Future Outlook for the llms.txt Standard?
The llms.txt standard sits in what The Prompt Bench describes as the "useful but under-honored" stage as of mid-2026 — the infrastructure exists, a meaningful minority of sites have adopted it, but the demand side (AI providers actually reading it) has not yet caught up with the supply side (sites publishing it).
The specification itself is actively maintained. It reached version 1.7.0 on May 11, 2026, which signals that the technical community behind it is treating it as a living standard rather than an abandoned proposal. The most significant institutional signal remains Google including llms.txt in their Agent2Agent (A2A) protocol in April 2025 (Ahrefs, 2025) — even if Google's own guidance simultaneously says the file is not required for generative AI optimisation. Those two positions are not necessarily contradictory: the A2A inclusion may be forward-looking infrastructure for agentic use cases rather than current generative search.
For operators, the cost-benefit calculation is clear. The file costs one evening and 20–50 lines of markdown. The downside if AI providers never adopt it at scale is zero — you have a well-organised content inventory that is useful for your own internal audits. The upside if provider behaviour shifts — as it plausibly will as agentic AI use cases grow — is meaningful discoverability advantage over competitors who waited.
The analogy to early structured data adoption is instructive. Schema markup was optional for years before Google began surfacing rich snippets prominently. The stores that had clean structured data in place when that shift happened captured the benefit immediately. The llms.txt window may follow a similar pattern.
Editor's Take — Michal Baloun, Co-founder
The data in this article tells a story that is easy to misread in both directions. The pessimistic read: 97% of domains with a valid llms.txt received zero crawler requests in May 2026, Google says the file is not needed, and Search Engine Land's own seven-week experiment showed no bot traffic. That sounds like a dead standard.
The optimistic read: Google built it into their A2A protocol, the specification is on version 1.7.0 and still being updated, and OpenAI appears to be crawling these files at high frequency on at least some sites. That sounds like infrastructure being laid quietly before the demand becomes obvious.
My position is closer to the optimistic read, but with a specific caveat about what "implementing llms.txt" actually means in practice. Most stores I see publish a file once and forget it. They list 150 URLs with no descriptions, include pages that have since been restructured, and point to category pages that return thin content anyway. That is not an llms.txt that helps a language model — it is noise.
The stores that will benefit when AI providers do start reading these files systematically are the ones that treat the file as a living editorial document: 15–25 carefully chosen pages, updated when the catalogue changes, pointing to the content that actually answers buyer questions. That is a different task from "generate an llms.txt and tick the box."
If your store has fewer than 50 pages of real editorial content — buying guides, comparison articles, detailed product narratives — fix that first. The llms.txt is a pointer to substance. If the substance is not there, the pointer does not help.
Here's what advice from Margly looks like
Most analytics dashboards stop at "your number is X". Margly stops at the next sentence — what to do, where, how much it's worth. Recommendations Margly would surface for the patterns described in this article:
-
High priority "Audit your llms.txt against your current top-20 revenue pages and remove any broken or redirected URLs." Search Engine Land's zero-bot-traffic result was partly attributable to stale file content — a clean, accurate file is the baseline before any other optimisation matters. Estimated impact: +$1,500 to +$3,000 / month
-
High priority "Add your three best buying guides and your returns policy to the llms.txt editorial section." GEO research found visibility gains of up to 40% in generative engine responses when AI crawlers can access rich prose content — your llms.txt is the fastest path to surfacing that content. Estimated impact: +$2,000 to +$4,500 / month
-
Medium priority "Reduce your llms.txt to under 30 entries with a one-line description per link." The specification recommends 10–30 key pages; files with 200+ undescribed URLs give AI models no useful signal about content priority. Estimated impact: +$800 to +$1,800 / month
-
Medium priority "Create category-level index pages for your top product groups and link them from llms.txt." 100 products across 50 locations can create 5,000 crawlable AI-optimised entry points — your llms.txt should point to the index layer that makes those discoverable without listing every URL. Estimated impact: +$1,200 to +$2,800 / month
Notice none of those needed a CSV export. That's the difference between raw analytics and concrete advice.
llms.txt what is it
The llms.txt file is a proposed standard, emerging in 2024, designed to help AI crawlers and generative search engines discover and understand your website's content. Place a plain markdown file at yourdomain.com/llms.txt and it acts as a curated reading guide — listing your most important pages with short descriptions so that language models can quickly identify what your site covers and where to find your best content.
Unlike a sitemap, which lists every URL for completeness, an llms.txt is intentionally selective. Jeremy Howard, who published the proposal on September 3, 2024, designed it to give AI models a signal about priority rather than an exhaustive index.
llms.txt best practices
Keep the file short: between 10 and 30 key pages is the recommended range, not a dump of your entire URL structure. Each entry should carry a description of under 15 words so a language model knows what it will find before following the link.
The file requires only 20–50 lines of markdown and can be built in an evening. Group links by content type — products, editorial, trust pages — and update the file whenever you launch or retire major content sections. A stale llms.txt pointing to 404 pages degrades the signal you are sending to AI crawlers.
llms.txt standard
Jeremy Howard proposed the llms.txt standard in September 2024 as a markdown-based convention for helping AI models navigate website content. The specification has been actively maintained and reached version 1.7.0 by May 2026. Google included llms.txt in their Agent2Agent (A2A) protocol in April 2025, which represents the most significant institutional endorsement to date.
As of mid-2026, major AI providers including Anthropic, OpenAI, Google, and Perplexity have not made public commitments to reading llms.txt as part of their retrieval pipelines. The standard is best understood as early infrastructure for agentic AI use cases rather than a confirmed ranking signal for current generative search.
llms.txt for seo
An llms.txt file is not a direct ranking factor for traditional search engines. Google's own guidance states that machine-readable files like llms.txt are not needed for generative AI optimisation. However, the file sits within the broader practice of Generative Engine Optimisation (GEO), and research published on arXiv found that GEO techniques can boost visibility by up to 40% in generative engine responses.
The practical value for e-commerce SEO is in content discoverability: an llms.txt that points AI crawlers to your buying guides, category pages, and trust content increases the likelihood that those pages are indexed and cited when AI assistants answer shopping queries. It is a low-cost, low-risk addition to a broader content strategy rather than a standalone traffic driver.
Sources:
- Semrush: llms.txt Guide
- Ahrefs: What Is llms.txt?
- Practical Ecommerce: llms.txt Could Help AI Find Your Store
- BigCommerce: Ecommerce LLMs.txt
- Search Engine Journal: llms.txt Shows No Clear Effect on AI Citations
- GreenGeeks: What Is llms.txt, Does It Work, and How to Add One
- The Prompt Bench: The llms.txt Standard Explained
- TNGShopper: llms.txt for E-commerce
- EasyLLMsTxt: Ecommerce Template
- SE Ranking: llms.txt Research
- arXiv: GEO Visibility Research