hello@afrazalam.com Kolkata, India
Follow on:
AI Marketing

What Is an llms.txt File? A No-Hype Guide for 2026

An honest look at what an llms.txt file actually does, whether ChatGPT, Claude, or Gemini read it yet, and the 6-step process to build one for your site in under an hour.

Last updated: September 2026 by Afraz Alam, Founder of Afraz Alam Marketing Services. I run AI search visibility work for clients, so I test files like this on real sites before I write about them.

An llms.txt file is a plain Markdown file at yourdomain.com/llms.txt that gives AI agents a short map of your site. This guide covers what a llms.txt file actually is, the real spec, whether ChatGPT, Claude, or Gemini read it, and how to build one in under an hour if you decide it is worth your time.

TL;DR

  • An llms.txt file is a Markdown file at your site root that lists your most important pages for AI agents to read, in a format they can parse fast.
  • Anthropic, OpenAI, and Google all publish one for their own developer docs, but none of them has confirmed their crawlers actually read yours.
  • Ahrefs studied 137,000 domains: 28% publish an llms.txt file, and of the ones with a valid file, 97% got zero requests for it in May 2026.
  • Google's own John Mueller called it a "temporary crutch" for AI coding tools, not a search ranking signal.
  • It takes under an hour to build and carries close to zero risk, so I still recommend it, just not as a replacement for the SEO and AEO work that actually moves the needle.

Who this is for: business owners and marketers who keep hearing "add an llms.txt file" and want a straight answer on whether it is worth the hour, not another generator tool pretending it is the next big ranking factor.

What is an llms.txt file, exactly?

An llms.txt file is a plain Markdown document hosted at yourdomain.com/llms.txt. It gives an AI agent, a coding assistant, a chatbot with browsing, a curated shortlist of your most useful pages instead of making it guess by crawling your whole site.

The idea comes from Jeremy Howard, co-founder of Answer.AI, who first proposed it in September 2024 and shipped a v2 of the spec in August 2026. The file itself is short on purpose. It has one required heading with your site or project name, a one-line summary, and then a set of links grouped under H2 headings, docs, policies, products, whatever matters for your site. The detail lives on the pages it links to, not in the file itself.

At Afraz Alam Marketing Services, I think of it the same way I think of a well-built sitemap for a new hire. It is not the whole company handbook. It is the two-page cheat sheet that gets someone oriented fast.

Why listen to me on this

I run SEO, AEO, and AI search visibility work for clients across e-commerce, coaching, and home services businesses in the US, and I write and publish the technical SEO content on afrazalam.com myself. I already cover the AEO framework in full in my guide to Answer Engine Optimization, and the 9-tool breakdown in ChatGPT SEO tools worth paying for.

I am not going to tell you llms.txt is the next big ranking factor, because the data does not back that up yet. Here is the truth: most of what gets written about llms.txt online is either a generator tool trying to sell you a plugin, or a rewritten version of the original spec with no independent testing behind it. This guide pulls from the actual spec at llmstxt.org and from Ahrefs' own crawl study of 137,000 domains, not guesswork.

Why llms.txt exists in the first place

AI agents now read websites constantly. A coding assistant fetches a library's docs to get an API call right. A chat assistant with browsing reads a page to answer a product question. Regular HTML is built for people, it is wrapped in navigation, ads, and JavaScript, and turning that back into clean text an AI can use is slow and imprecise.

Context windows are bigger than they used to be, but they are still too small to hold most full websites, and every wasted token costs the AI provider time and money. llms.txt exists to hand agents a short, clean map instead of forcing them to parse your whole HTML site to find one answer.

It works mainly for software documentation right now, where coding agents follow it straight to API references. Anthropic, OpenAI, and Google all publish one for their own developer docs, which is the strongest real-world proof the format solves a genuine problem, at least for that use case. When a client at Afraz Alam Marketing Services asks me whether to prioritize this, that developer-docs pattern is exactly what I point to.

llms.txt vs robots.txt vs sitemap.xml

These three files sound similar and do different jobs. Confusing them is the most common mistake I see when I run a technical SEO audit for a new Afraz Alam Marketing Services client.

File What it actually does Who reads it
robots.txt Tells automated crawlers what they are and are not allowed to access. This is the file that actually controls crawl access, including for GPTBot and other AI crawlers. Search engine bots and most AI crawlers that respect the standard
sitemap.xml Lists every indexable page on your site for search engines. Built for completeness, not for a small AI context window. Search engine bots
llms.txt A short, curated map of your most useful pages, in Markdown, meant to be read on demand when an agent needs specific information. AI agents and coding tools, when they choose to look for it, which right now is rarely

The key insight: robots.txt is the file that actually gates AI crawler access today. If you only have time to get one of these three right, make it robots.txt, not llms.txt.

robots.txt vs sitemap.xml vs llms.txt file comparison showing each file's real job

What actually goes inside the file

The spec is specific about the order. A valid llms.txt file contains, in this order: an H1 with your site or project name, a blockquote with a one-line summary, optional plain paragraphs with more context, and then H2 sections that each list links in the format [Link title](url): optional note.

Here is a simplified real-world shape, close to what Anthropic publishes for its own docs:

# Afraz Alam Marketing Services

> Digital marketing consulting for US e-commerce, coaching, and home
> service businesses. SEO, Google Ads, Meta Ads, and AI search
> visibility work.

## Services
- [SEO Strategies](https://afrazalam.com/services/seo/): Local and technical SEO
- [AI Marketing](https://afrazalam.com/services/ai-marketing/): AI-driven visibility and automation

## Guides
- [Answer Engine Optimization](https://afrazalam.com/blog/answer-engine-optimization-aeo/): How to get cited by ChatGPT and Perplexity

## Optional
- [Case studies](https://afrazalam.com/case-study/): Real client results

Notice the "Optional" section at the end. That is a real convention in the spec, not something I invented. It signals content an agent can skip when it only has a short context budget to spend. This is close to the actual structure I would set up for an Afraz Alam Marketing Services client site, service pages first, guides second, case studies filed under Optional.

Does anyone actually read it? What the data says

Ahrefs data showing 97 percent of domains with an llms.txt file received zero requests for it in May 2026

Here is the part most llms.txt content skips. Ahrefs studied 137,000 domains through its own web analytics and found that 28% publish an llms.txt file, more than one in four sites. That sounds like real adoption, until you look at whether anyone actually requests the file. Of roughly 38,000 domains in that study with a valid llms.txt, 97% received zero requests for it in May 2026. No bots, no humans, nothing.

The provider picture is not much better. OpenAI's GPTBot honors robots.txt but has not officially adopted llms.txt. Anthropic publishes its own file but has not confirmed its crawlers use the standard on other sites. Google added llms.txt as a reference inside its Agent2Agent protocol in April 2025 but has not committed to crawling it, and Google's own generative AI optimization guide, published in a section literally titled "mythbusting," told site owners that machine-readable files like llms.txt are not required to appear in AI search results.

Google's John Mueller went further and called llms.txt a "temporary crutch, perhaps to save some tokens" for AI coding tools parsing developer docs, and compared it to the old keywords meta tag: something a site owner claims about their own site, that nobody is obligated to believe or check.

Note: I am not telling you llms.txt is worthless. I am telling you it is unproven for general AI search visibility, useful mainly for developer docs today, and not something Afraz Alam Marketing Services would prioritize over the technical SEO and AEO fundamentals covered in my technical SEO audit checklist.

How to build your own llms.txt file

Since it takes under an hour and carries almost no downside beyond making your page structure slightly easier for a competitor to scan, here is exactly how Afraz Alam Marketing Services would build one for a small business site.

  1. List your 5 to 10 most important pages. Home, your core service pages, your best 2 to 3 guides, your contact or booking page. Do not list everything, that defeats the point.
  2. Write a one-line summary of your business. This becomes the blockquote under your H1. Keep it under 30 words.
  3. Group your links under H2 headings. Services, Guides, About, whatever categories fit. Use the exact link format: [Page title](full URL): one short note.
  4. Add an Optional section for anything secondary. Case studies, press mentions, older content. Anything an agent could skip without losing the core picture.
  5. Save it as plain text named llms.txt and upload it to your site root, so it loads at yourdomain.com/llms.txt with no folder in front of it.
  6. Test it by asking an AI assistant a question about your business, giving it only your llms.txt file as a starting point, and see if the answer holds up.

If your site runs on WordPress, Yoast SEO and AIOSEO both include an llms.txt generator now, so you may already have one without knowing it. Check yoursite.com/llms.txt directly to find out.

Mistakes I see businesses make with this

The biggest one is treating llms.txt as a shortcut around real content work. It is a map. If the pages it points to are thin, outdated, or missing real answers, the map just leads an agent to a dead end faster. This is the same mistake I flag constantly at Afraz Alam Marketing Services when a business wants a quick technical fix instead of fixing the content underneath it.

The second is copying a generic template word for word. An llms.txt file that lists "Home, About, Contact, Blog" tells an agent nothing useful. Every link needs a note that actually says what is on that page.

The third is spending a week on this instead of an hour. I have seen businesses pay a developer to build a "dynamic llms.txt generator" for a 12-page site. That is solving a problem you do not have yet.

Where this fits your real AI visibility strategy

llms.txt is a small, optional piece of a much bigger picture. If you actually want ChatGPT, Perplexity, or Google's AI Overviews citing your business, the work that moves that needle is the same work that has always moved it: clear, well-structured, honestly written content that answers a real question completely, backed by correct schema markup so machines can parse your page without guessing.

I cover that full framework, the ANSWER method, AEO vs traditional SEO, and the schema types that actually matter, in my complete guide to Answer Engine Optimization. Add your llms.txt file after that work is done, not instead of it. That is the order Afraz Alam Marketing Services follows on every client site.


Frequently asked questions

What is an llms.txt file used for?

It gives AI agents and coding assistants a short, curated list of a website's most useful pages, so they can find relevant content without crawling the entire site. It works best for software documentation today.

How do I create an llms.txt file?

Write a plain Markdown file with an H1 for your site name, a one-line blockquote summary, and H2 sections listing your key pages with short notes. Save it as llms.txt and upload it to your site's root folder.

Where do I upload my llms.txt file?

At the root of your domain, so it loads at yourdomain.com/llms.txt with no subfolder. If you run WordPress with Yoast SEO or AIOSEO, check whether one is already being generated for you.

Is llms.txt the same as robots.txt?

No. robots.txt controls what crawlers, including most AI crawlers, are allowed to access. llms.txt is a curated content map that agents may or may not choose to read. robots.txt is the one that actually gates access today.

Do I still need a sitemap.xml if I have an llms.txt file?

Yes. sitemap.xml lists every indexable page for search engines and is built for completeness. llms.txt is a short, hand-picked shortlist for AI agents. They do different jobs and neither replaces the other.

Is llms.txt actually used by ChatGPT, Claude, or Gemini?

Not confirmed. OpenAI's GPTBot has not officially adopted it. Anthropic publishes its own file but has not confirmed its crawlers read the standard on other sites. Google added it to a separate agent protocol but has not committed to crawling it for search.

Is llms.txt worth the effort for a small business?

It takes under an hour and carries very little risk, so I still recommend adding one. Just do not expect it to move your AI search visibility on its own. The real work is the content and schema behind the links it points to.

Does having an llms.txt file help my Google rankings?

No confirmed effect. Google's own guidance has told site owners that machine-readable files like llms.txt are not required to appear in generative AI search results, and Google's John Mueller has described it as a tool mainly useful for AI coding assistants, not a search ranking signal.


Want your AI visibility strategy built properly, not just an llms.txt file bolted on?

I run SEO, AEO, and AI search visibility work for US businesses on a free 30-minute strategy call. No pitch deck, just a straight look at where you actually stand.

Book a free strategy call

Questions, disagreements, or something I missed? Reply on my Instagram @afrazalammarketingservices or drop a comment below. I read everything.

Frequently Asked Questions

It gives AI agents and coding assistants a short, curated list of a website's most useful pages, so they can find relevant content without crawling the entire site. It works best for software documentation today.

Write a plain Markdown file with an H1 for your site name, a one-line blockquote summary, and H2 sections listing your key pages with short notes. Save it as llms.txt and upload it to your site's root folder.

At the root of your domain, so it loads at yourdomain.com/llms.txt with no subfolder. If you run WordPress with Yoast SEO or AIOSEO, check whether one is already being generated for you.

No. robots.txt controls what crawlers, including most AI crawlers, are allowed to access. llms.txt is a curated content map that agents may or may not choose to read. robots.txt is the one that actually gates access today.

Yes. sitemap.xml lists every indexable page for search engines and is built for completeness. llms.txt is a short, hand-picked shortlist for AI agents. They do different jobs and neither replaces the other.

Not confirmed. OpenAI's GPTBot has not officially adopted it. Anthropic publishes its own file but has not confirmed its crawlers read the standard on other sites. Google added it to a separate agent protocol but has not committed to crawling it for search.

It takes under an hour and carries very little risk, so I still recommend adding one. Just do not expect it to move your AI search visibility on its own. The real work is the content and schema behind the links it points to.

No confirmed effect. Google's own guidance has told site owners that machine-readable files like llms.txt are not required to appear in generative AI search results, and Google's John Mueller has described it as a tool mainly useful for AI coding assistants, not a search ranking signal.
Afraz Alam

Afraz Alam

Digital Marketing Consultant

Afraz Alam is the founder of Afraz Alam Marketing Services, a digital marketing consultant helping businesses grow through SEO, Google Ads, Meta Ads, AI-driven marketing, and conversion-focused website development. He works directly on every account himself, no account managers, no handoffs, and reports on leads, calls, and revenue, not likes or impressions. His approach stays data-driven and transparent: real strategy, real reporting, no guaranteed rankings or inflated promises.

Prefer done-for-you over do-it-yourself?

Reading is a great start — but if you'd rather have it handled, book a free call and let's grow your business with SEO, ads, and AI.

Book a Free Strategy Call →