I recently had a client insist “add llms.txt before the competition does,” offering to pay me to make it happen. As a senior SEO specialist, I’ve heard this story before: clients see a shiny new tactic and fear missing out.
In this article on implementing llms.txt in 2026, we’ll cut through the hype. You’ll learn what an llms.txt file actually is, what changed in the August 2026 v2 update, and the honest truth about whether AI crawlers even care. We’ll also walk through a step-by-step llms.txt file format guide and show you how to implement it correctly if you do decide to do it.
(Spoiler: even Google says you don’t need it.) Along the way, we’ll explain the distinction from ai.txt and where this fits into your SEO toolkit. For readers interested in technical SEO, our Refonte Learning SEO & SEA Mastery Program covers foundational skills like these in depth.
The Client Who Wanted This Before the Competition Did
I sat down with the client over a call. He was convinced llms.txt was essential: “We need this for AI search,” he said, worried our rivals might have an edge. In my 10+ years of consulting, I’ve seen this happen many times: a new format emerges, industry buzz builds, and suddenly clients want it done yesterday. I empathized but had to set realistic expectations. Key points I covered:
FOMO vs. Reality: The client feared “everyone’s doing it,” but so far adoption is limited. There’s no evidence all competitors use llms.txt.
Distinguish ai.txt: They’d also heard of an ai.txt file (different purpose) and thought they were the same. I clarified they are two different files (more on this below).
Cost of Implementation: Adding an llms.txt is relatively quick (especially if you have Markdown docs), but someone still has to write it well. It's far from a magic SEO bullet.
In the end, I promised a clear recommendation backed by facts: I’d explain llms.txt exactly, v2 changes, and Google’s stance. If we did it, it would be a neutral “insurance” step, but not at the expense of more important SEO fixes.
What llms.txt Actually Is, in Plain Terms
Put simply, an llms.txt file is a Markdown file that lives at a site’s root (or a subpath) and provides a concise, structured roadmap for AI agents. It was proposed by Jeremy Howard in 2024 and updated in 2026. The goal is to help language models find your site’s most important content efficiently, without wading through all the HTML and navigation. Think of it as a wiki-style table of contents for AI, not for humans.
Important facts about llms.txt:
Purpose: It’s meant as a curated, LLM-friendly summary of your site. It contains the site or project name as an H1 title, a short blockquote summary, and sections (H2 headings) listing key pages or docs.
Format: Written in Markdown (not XML or JSON). Each section has a Markdown link [Title](URL) per item, with optional descriptions. It’s both human- and machine-readable.
Adoption: Thousands of sites publish an llms.txt now. Documentation platforms (Mintlify, GitBook, etc.) can even generate them automatically. The spec expects AI agents (like chatbots) to view or search this file to find where to go next.
Comparison to other files: Unlike robots.txt (which tells bots what not to crawl) or a sitemap (which lists all pages), llms.txt is selective and contextual. It complements those: you can still have robots.txt and sitemap.xml separately.
In a nutshell, llms.txt is a community-driven proposal (not an official internet standard) to give AI models a helpful starting point on your site. It was “Published Sept 3, 2024; Modified Aug 10, 2026” on the llms specification site. Version 2 (v2) just shipped in 2026 to reflect how people actually use it (more on that next).
What Changed in Version 2 (August 2026)
Version 2 of the llms.txt spec (August 10, 2026) incorporated feedback from two years of usage. The key changes include:
Link Relations Added: V2 specifies HTML <link> and HTTP Link: headers to discover llms.txt. Now you can use rel="alternate" type="text/markdown" to point from a page to its Markdown copy, and rel="describedby" to point from a page to the llms.txt that covers it. This fixes a major ask: How does an agent find the right llms.txt file?
Markdown Filename Flexibility: V1 forced a particular .html.md naming convention. V2 allows either appending .md (page.html to page.html.md) or replacing the extension (page.html to page.md). This recognizes how tools actually publish Markdown.
Scope Clarified: Now the spec clearly defines that an llms.txt in a subpath (e.g. /docs/llms.txt) covers all pages under that path, with the most specific file taking precedence. This helps on platforms like GitHub Pages where you only own a subfolder.
Consumption Guidance: V2 explicitly says how agents use the file: they should open or search it, then follow relevant links to detailed content. The old “llms_txt2ctx” tool was dropped, so llms.txt itself is just a pointer, not an expansion tool.
“Optional” Section Semantics: V1 gave a special meaning to an “Optional” section for tools to skip. V2 keeps Optional as a useful heading convention but removes its mechanical significance. It’s now purely by convention (shorter context need) rather than a strict rule.
Examples Updated: The background and examples on the spec site now reflect real agent use (like coding agents fetching docs) instead of speculative scenarios.
Why This Isn't a W3C or IETF Standard
llms.txt is not an official web standard. The llms specification is a community proposal by practitioners (Jeremy Howard et al.), hosted on GitHub. It’s not published by the W3C, IETF, or any formal standards body.
As the site itself titles it, “The /llms.txt file, v2” is a proposal. It’s open for feedback on Discord and GitHub, but there’s no formal Internet Assigned Numbers Authority (IANA) registry for it. In short, it’s a grassroots project, not a mandated protocol.
llms.txt vs. ai.txt: Two Different Files, Often Confused
Two similarly named files can confuse people: llms.txt and ai.txt. Don’t mix them up. In plain terms:
llms.txt: Think “LLM-friendly sitemap.” It’s a guide for AI models to the important pages on your site (as described above). It lives at /llms.txt (or a path like /docs/llms.txt) and uses Markdown lists to link to your site’s content. It’s all about finding and summarizing content.
ai.txt: Think “AI crawler directives.” The ai.txt file (often discussed for data rights) is more like a hybrid of robots.txt and licensing info. It typically lives at /ai.txt and tells crawling systems (human or AI) what content can be used for training, API endpoints, or what to avoid. For example, one recommendation is to include a training license (like CC BY 4.0) or disallow clauses in ai.txt.
In practice, llms.txt and ai.txt serve different purposes. Our Refonte Learning article Generative Engine Optimization for SaaS marketers explains this: ai.txt is about content usage policy, while llms.txt is a content discovery guide. Don’t treat llms.txt as the AI equivalent of robots.txt. Instead, use llms.txt to list your docs, tutorials, and rich resources, and use robots.txt and ai.txt to control access and licensing.
The Claim: AI Labs Publish Their Own llms.txt Files
A fact often cited as proof of llms.txt’s importance is that the major AI companies publish their own llms.txt files on their developer sites. Indeed, the llms spec site notes: “The AI labs themselves publish llms.txt files for their own developer docs: OpenAI, Anthropic, and Gemini.” In other words, if you go to the OpenAI or Anthropic docs homepage, you’ll find an llms.txt detailing their content.
However, be clear what this means: those companies are simply providing a helpful summary for visitors and developers of their own docs. It’s them acting as publishers, not evidence that their crawlers are consuming llms.txt files from the whole web. It is analogous to Google maintaining a robots.txt file on its developer site: useful for organizing that site, but not a signal that Googlebot obeys robots.txt on any random site. In summary:
Published by them: OpenAI’s dev site, Anthropic’s docs, and Google’s Gemini docs each have an llms.txt.
Not confirmation of reading: Their AI models (e.g. GPTBot, ClaudeBot, etc.) have not officially said they fetch other sites’ llms.txt. Companies may publish it for consistency, but that’s not the same as their crawlers using external ones.
Why That's Not the Same as Crawlers Reading Yours
Let’s be crystal clear: just because the AI labs serve llms.txt files for their own documentation doesn’t mean ChatGPT or Googlebot will read your llms.txt. The two facts are unrelated:
Self-documentation: The labs want their own information indexed by AI agents (and maybe checked by tools like Chrome’s Lighthouse). Publishing llms.txt there helps users find their docs.
Crawler behavior unknown: None of the lab’s official docs say “our AI trains on llms.txt from random websites.” In fact, Google’s guidance specifically says Google Search ignores these files, and other AI companies haven’t announced using llms.txt.
In practice, the existence of those official llms.txt files is interesting but not proof. It’s like saying “Google writes blog posts in Markdown, so maybe Googlebot reads all markdown”, different intent.
Google’s Position, Stated Directly
What does Google say about llms.txt? The answer: essentially, “meh.” Google’s search teams have explicitly said you don’t need special AI markup or an llms.txt for Google’s generative search experiences.
In Google’s own words (from Search Central): “It also says you do not need special AI markup or an llms.txt file to gain visibility in Google’s generative Search.” In other words, following SEO best practices is the priority.
Updated Google guidance confirms this stance: Google now clarifies that it won’t harm or help your ranking to have an llms.txt, because “Google Search ignores them.” The relevant changelog notes: “If you create and maintain LLMS.txt files ... doing so won’t harm (nor help) your visibility or rankings in Google Search, as Google Search ignores them.”
To translate: Google will crawl and index your site as usual. It won’t look for llms.txt when deciding which pages to surface in AI Overviews or AI Mode. What matters to Google’s AI search is your content quality, structure, and traditional SEO factors (crawlability, page speed, schema, etc.). So for our purposes, Google’s official position is very clear: llms.txt is completely optional and irrelevant to Google’s algorithms.
As we discuss in our Google AI Mode vs. AI Overviews article, Google emphasizes core SEO fundamentals for generative AI visibility, not special markup. This is a key takeaway: even after the 2026 updates to Google AI Mode, llms.txt does nothing for Google Search. You should prioritize proven tactics over chasing this unconfirmed lever.
What’s Actually Unconfirmed Here
This brings us to the elephant in the room: we simply do not know if any AI crawler actually reads llms.txt files on arbitrary sites. No search engine or AI provider has publicly confirmed using llms.txt from other domains as a ranking or citation factor. In fact:
Google’s documentation only says it ignores llms.txt (for Google Search). It doesn’t address other services.
OpenAI and Anthropic have published llms.txt for their docs, but have not announced that GPTBot or ClaudeBot scans the wider web for llms.txt files.
Some in the SEO community, including experts like Google’s John Mueller, have called the effectiveness of llms.txt “purely speculative” (no one knows for sure). That skepticism is widespread in industry forums.
The documentation platforms and tools (Mintlify, Lighthouse, etc.) checking for llms.txt only prove people are building them, not that they’re being used by agents.
In short, adopting llms.txt is a shot in the dark. It remains an open question whether AI crawlers read llms.txt. Our research found no authoritative source either way; it remains an actively debated issue.
Google at least says there is no benefit for Google Search. Until AI companies clarify, the existence of llms.txt files is mostly a best-practice experiment, not a proven signal.
Why Multiple Sources Couldn’t Settle This Either Way
Even industry blogs and news sites are split. Chrome’s team includes llms.txt in its Lighthouse checks (hinting they think “agents” might care), but Google’s search docs emphatically say it’s unnecessary. Absent a public crawler spec or study, the debate continues. For now, think of llms.txt as a precautionary measure rather than a requirement.
Implementing llms.txt Correctly: The Format, Step by Step
If you do decide to implement llms.txt, here’s a step-by-step breakdown of the file format and placement. Treat this as a mini llms.txt file format guide:
File location: Create a plain text file named llms.txt at the root of your site or in a subpath (/docs/llms.txt if it only covers docs). The file covers all URLs under its path.
H1 Title: The first line must be a top-level heading with your site or project name. Example: # My Awesome Product. This title identifies your site to the agent.
Blockquote summary: Immediately under the H1, include a Markdown blockquote (> Your summary here). Write a concise, high-level description of what the site/project is and its purpose (2-3 sentences). This helps the LLM quickly understand context.
Optional intro paragraphs: You can add a few normal paragraphs (no headings) after the summary to explain any important details or usage notes. Keep it brief.
H2 sections with link lists: Organize main content areas under subheadings. For each H2 section:
Use a Markdown heading (e.g. ## Documentation or ## Blog Posts).
Under that heading, add a bullet list (-) of links. Each item must be in the form [Link Title](URL): Description. Example: - [API Guide](/docs/api.md): Detailed API documentation. The title should be short, and the optional description (after the colon) should give a hint of what’s there.
“Optional” section: It’s common to include an ## Optional H2 at the end for lower-priority links. Agents can skip this if needed. Again, use bullet links here. (This is optional in usage, not required.)
Rel tags (optional but helpful): If possible, add HTML <link> tags on your pages pointing to the Markdown and llms.txt files. For example, in the HTML <head> of a page, you could add <link rel="alternate" type="text/markdown" href="page.html.md"> and <link rel="describedby" href="/llms.txt">. V2 of the spec recommends these for discoverability. You can also use the HTTP Link: header for non-HTML content.
Validation: Make sure the file is valid UTF-8 Markdown and publicly accessible. No XML or special characters outside Markdown syntax. Test it by using it as context for an agent or running a linter.
Update as needed: Whenever your content structure changes, update llms.txt to match. It should always point to existing, valuable resources.
By following the spec details and example format, you’ll have a syntactically correct llms.txt. The key is clarity and conciseness: don’t overload it with every URL, just the most useful ones.
What to Actually Prioritize in the File
Not all pages belong in llms.txt. Focus on quality over quantity. Specifically:
Core summaries & docs: Include your home page, “About” or product overview page, developer docs, or tutorials: these are the places where an agent would first look to understand your site.
Unique content: Prioritize pages with unique expertise (e.g., your original research, whitepapers, key blog posts). LLMs love pages with rich information.
Markdown versions: If you have plain-text or Markdown versions of pages (like GitBook or GitHub-hosted docs), point to those via rel="alternate" links or directly in llms.txt. The spec encourages linking to LLM-friendly formats.
Concise titles and descriptions: Make each link entry clear and informative. Good link text and a brief description help the model grasp the context.
Link to external authoritative sources (sparingly): The spec allows links to other sites if they’re key to understanding (e.g., an industry standard doc). Use this only if it truly aids context.
Limit the size: Keep llms.txt fairly small (a few hundred lines at most) so an agent can read it in one context window. You can always link deeper pages if needed.
Be precise: Use simple, jargon-free language in the summary. The spec even advises testing your file by giving it to an agent and seeing what it infers. Clarity matters.
In summary: treat llms.txt as your curated TL;DR for each part of your site. If a page is only one of a hundred similar blog posts, it might not belong here. Instead, highlight the pages that best represent your site’s value to a user or developer.
Common Implementation Mistakes
Even if you decide to add llms.txt, watch out for these pitfalls:
Treating it like a sitemap: Do NOT list every URL. Unlike a sitemap.xml, llms.txt should be curated. For example, you would not include archive pages or every blog tag. It's meant to guide, not exhaustively enumerate.
Wrong format or location: The file must be Markdown text at the correct path. Don’t accidentally upload an HTML file or put it under /robots.txt. It should be named exactly llms.txt.
Neglecting the structure: Skipping the H1 title or not using headings defeats the purpose. Ensure the first line is an H1 with your site’s name. Without a clear title, agents get disoriented.
Too long or too vague: If your summary or sections ramble, you lose the point. This file is for quick reference; avoid long blocks of text or jargon without explanation.
Over-optimizing: Don’t try to cram keywords or treat it like SEO boilerplate. Agents want helpfulness, not fluff. Irrelevant buzzwords could confuse them.
Ignoring Google’s stance: Remember, this won’t boost you in Google Search (Google ignores it). If you implement it thinking it will help your core traffic, that’s misguided.
Forgetting to update: Just like a sitemap, an outdated llms.txt can mislead agents. Set a reminder to review it when your site structure changes.
Treating It Like an XML Sitemap
One very common error is assuming llms.txt is a drop-in replacement for an XML sitemap. It is not. Key differences:
Purpose: An XML sitemap is for search engines to crawl and index all your pages. llms.txt is for agents to understand and navigate your key content.
Content: Sitemaps contain every indexable URL. llms.txt should contain only a handful of curated links. If you dump every page into llms.txt, an LLM’s context will be wasted.
Format: Sitemaps use XML. llms.txt must use Markdown lists. Feeding an XML sitemap to an LLM likely breaks parsing, since llms.txt expects Markdown syntax.
Permissions: Unlike robots.txt, llms.txt doesn’t grant or restrict access. If you want to block scraping, use robots.txt or proper headers. Putting disallow rules in llms.txt will be ignored by any crawler that reads it.
In short, don’t confuse the two. Use your existing sitemap.xml for comprehensive SEO indexing, and reserve llms.txt for a human-readable guide to your content.
Should You Actually Spend Time on This in 2026
With all this said, should your team invest effort in llms.txt right now? Here’s a balanced view:
Cost is low: In practice, creating an llms.txt (especially on a documentation-heavy site) might take a few hours of a developer or content writer’s time. Maintenance is occasional. It’s a relatively small investment.
Benefit is speculative: There’s no guarantee any AI service prioritizes your llms.txt. Google doesn’t use it, and others haven’t confirmed. The potential upside is that some AI agent might find it useful for structured browsing.
Opportunity cost: Those same few hours could go into proven SEO or content improvements (faster pages, better headlines, schema markup). Those have known ROI.
Competitive landscape: If you know your competitors are experimenting with it (or if your audience is very tech-savvy), having one might help ensure AI agents see the best of your site first. It’s like buying an insurance policy (likely unnecessary, but cheap to add).
Future-proofing: If llms.txt (or a similar standard) gains traction, early adoption means fewer retroactive fixes. Right now it’s optional, but that could change in a few years.
SEO fundamentals unchanged: Regardless of llms.txt, keep doing solid technical SEO, quality content creation, and structured data work. Google’s message is that llms.txt adds nothing to those basics.
In practice, for most sites in 2026 llms.txt is nice to have, not a necessity. It probably won’t hurt anything to add it, but it may not materially change your search visibility. If the client is insistent, it’s fine to implement, but frame it as a minor compliance step rather than a major strategy.
A Realistic Cost-Benefit Framing
Time: A few hours now + occasional updates (e.g., quarterly review).
Potential payoff: If AI agents someday use llms.txt widely, early adopters might see more precise citations. But that’s uncertain.
Downside: Virtually none for SEO because Google ignores it; the only costs are developer time and maintenance.
Recommendation: If you have Slack time or a willing intern, drafting a quality llms.txt can be done. Otherwise, don’t prioritize it above tasks like fixing broken links or improving content authority.
Ultimately, llms.txt is more about good housekeeping for the AI era than about immediate gains. Think of it as optional forward-compatibility insurance.
What to Tell a Client Who Insists on It
When a client insists on llms.txt, honesty and clarity are key. Here’s how to explain it:
Acknowledge the interest: Say that llms.txt is a new best-practice under discussion in the SEO community, and it’s understandable to consider it.
Google’s take: Emphasize that Google itself has said it doesn’t use or require llms.txt. So it won’t change rankings on Google Search. (If their concern is Google traffic, llms.txt won’t help.)
Other AI platforms: Explain that while companies like OpenAI and Google have published llms.txt files for their docs, no one has confirmed their crawlers use external llms.txt. So far it’s just speculation for broader web search.
Implementation plan: Offer to create the file correctly (following our steps above) as a small, paid task. You might say: “We can certainly add it. It’s quick, but let’s treat it like an insurance policy: useful, but not critical.”
Focus on fundamentals: Remind them you will continue prioritizing proven strategies (site speed, structured data, quality content). llms.txt will be a minor checklist item.
Set expectations: Be clear: “I will implement llms.txt for you, but based on Google’s guidelines it’s unlikely to move the needle immediately. We’ll keep monitoring AI trends and adjust as needed.” This shows transparency.
Mitigate FOMO: Sometimes clients worry about “falling behind.” You can reassure them that if llms.txt becomes valuable, the cost of adding it later is still low. Right now it’s more important to get core SEO fundamentals right.
In practice, most clients appreciate the honest expert advice. Pointing them to authoritative statements (like Google’s guidance or our own analysis) helps build trust. If they still want it, proceed, but frame it in the scope of a larger SEO roadmap.
SEO Specialist Salaries in 2026
As context for career outcomes, it’s useful to know how SEO professionals are compensated. According to Indeed (data updated Aug 15, 2026):
Metric | United States (Indeed, Aug 2026) |
Average SEO Specialist salary | $80,500 per year |
Salary range | $41,726 to $155,304 per year |
Data source | Based on 898 postings (past 36 months) |
Indeed reports an average salary of about $80,500/year for SEO Specialists, with top salaries over $155k. This is broadly consistent with Refonte Learning’s cited starting salary of ~$75,000+, suggesting strong demand for skilled SEOs. (The wide range reflects experience and industry.) For a recent grad considering this field, these are encouraging numbers; SEO expertise remains well-compensated.
Building This Skill Set: The Refonte Learning SEO & SEA Mastery Program
If you’re interested in mastering technical SEO tactics like llms.txt and related implementation work, consider the Refonte Learning SEO & SEA Mastery Program. This intensive three-month course (12-14 hours/week) is designed for beginners and covers everything from basics to advanced implementation. Key points:
Curriculum: 7 modules including Introduction to SEO & SEA, SEO Tools and Techniques, Crafting Compelling Content, Google Advertising Fundamentals, Analytics and Performance Tracking, Advanced SEO Strategies, and a Capstone Project (SEO Audit & Campaign). These modules directly teach the foundations needed for tackling technical directives, content strategy, and analytics.
Mentor: The course is led by Ms. Emily Taraji (Digital Marketing Dept.), a seasoned SEO expert with 10+ years of hands-on experience. She guides students through real-world projects and best practices.
Prerequisites: No prior experience is required; you only need a willingness to learn (must be pursuing a bachelor’s or higher). Basic marketing knowledge helps but isn’t mandatory.
Fees: The program costs $300 total (one-time payment), or two installments of $204 and $98. This is a 30% discount off the $387 list price.
Outcomes: Graduates are prepared for roles as SEO Specialists, SEM Specialists, or Digital Marketing Managers. (Refonte’s marketing mentions ~$75K+ starting salaries for these careers, though actual salaries can vary.) The program cites around 100,000 annual job openings in this field.
Relevance to llms.txt: While the curriculum doesn’t name llms.txt explicitly, the SEO Tools and Techniques and Advanced SEO Strategies modules equip you to evaluate and implement new technical SEO files and markups. In short, it builds the precise skill set for discerning whether tactics like llms.txt are worth using, and if so, how to do them correctly.
By linking a practitioner-focused course with industry needs, Refonte Learning aims to prepare students for today’s SEO landscape. Whether or not llms.txt becomes mainstream, understanding file-based directives, markup formats, and crawl behavior is part of advanced SEO know-how, and the program covers these foundations.
