What Is llms.txt? (And Why robots.txt Matters More)
What Is llms.txt? (And Why robots.txt Matters More Right Now)
An llms.txt file cannot block an AI crawler (the automated software AI tools send out to read websites), unblock one, or make ChatGPT recommend your business, no matter how well it's built. It's a short plain-text file that lists your most important pages for an AI tool to read. The file that can actually lock AI out of your website by accident is a much older one: robots.txt, and checking it takes about two minutes.
Key Takeaways
- llms.txt can't block or grant access to anything. It's a reading list, not a gate.
- No major AI provider has confirmed reading it yet. A 2026 study of 137,000 sites found 97% of these files are never even requested.
- robots.txt is the file that can genuinely and accidentally lock AI out of your site, usually without anyone deciding to do it.
- Different AI bots do different jobs. Blocking one doesn't affect another, even from the same company.
- Both files take about two minutes to check, starting today.
What llms.txt Actually Is (and Isn't)
An llms.txt file is a plain text file, written in a simple formatting style called markdown, that sits at the root of a website (the top level, as in yourdomain.com/llms.txt) and lists the site's most important pages in a short, organized format: a heading, a one-line description of the business, then a list of links grouped by what they cover. The idea is to hand an AI tool a clean summary instead of making it dig through every page to figure out what a business does.
Here's the part most guides bury: it does not block anything, does not grant permission for anything, and does not change whether a page gets indexed (listed in search results). Think of it as a table of contents left on the counter, not a lock or a key.
The honest evidence says this file isn't doing much yet, either. A 2026 Ahrefs analysis of 137,000 domains found that 97% of these files are never even requested by any crawler. Google's own John Mueller has said publicly that the same thing shows up in server logs (the record of every visit a website gets, bots included): AI services check a site without ever asking for it.
As of 2026, no major AI provider has confirmed using it to decide rankings or recommendations. If you've read an article promising this file as the new must-have for getting found by AI, that promise is ahead of the evidence and worth a second look before anyone spends real time on it. For anyone who'd rather not go digging alone, GetLocalLeads.AI's free AI Visibility Audit is the simplest first step.
The File That Can Actually Block You: robots.txt
robots.txt has been around for thirty years. It's the file search engines, and now AI crawlers, check first to see what parts of a site they're allowed to read. Unlike llms.txt, robots.txt genuinely controls access. Get it wrong, and you can lock out a real chunk of your visibility without ever meaning to.
That's how it usually goes wrong. A security plugin added years ago to stop scraper bots (programs that copy content off websites) sometimes ships with a default rule that blocks AI crawlers too. Other times a rule written for a staging site (a private test copy of the website) gets copied to the live one and never revisited. Nobody wakes up and chooses to block ChatGPT. It happens by accident, in a file almost nobody looks at twice.
Not every AI bot from the same company does the same job, either. OpenAI's own developer documentation confirms that GPTBot crawls content for model training (teaching the AI), while a separate bot called OAI-SearchBot powers live ChatGPT search results, and the two are controlled independently in robots.txt. Blocking one has no effect on the other.
So a business that says "we already blocked the AI bots" may have blocked the training crawler while leaving itself invisible in ChatGPT's actual search answers, or the reverse. Other AI companies run their own separately named crawlers, each controlled by its own line in the same file.
Say a contractor installed a security plugin two years ago to keep bad bots off the site, checked the box that looked safest at the time, and never opened that settings page again. The plugin's default happens to disallow every AI crawler along with the scrapers it was built to stop. Nobody chose that outcome, and nobody would notice unless they went looking. Worth checking whether that's your website too, or book a call and let GetLocalLeads.AI check it with you.
Why Check Both Anyway, Even Though llms.txt Isn't Proven
In GetLocalLeads.AI's experience, AI visibility (whether AI tools like ChatGPT can find your business and bring it up when someone asks) can shift within 24 hours of a change like an updated robots.txt or llms.txt file, though results are never guaranteed. That's a small window for something that costs almost nothing to check.
The math is lopsided in a useful way. Adding a basic llms.txt costs almost nothing, even unproven, so doing it properly is cheap insurance rather than a priority fight. Checking robots.txt costs nothing either, but getting it wrong silently costs something real: a fast-growing share of customers now ask an AI assistant for local recommendations, and the AI visibility basics are the place to start if that shift is new to you. A wrongly blocked robots.txt shuts that whole channel off without a warning light anywhere on the dashboard.
Take an HVAC company that just spent real money on a redesign, new photos, and a clearer services page, all aimed at looking trustworthy to a customer researching repair options. None of that spending reaches the customer whose first stop was an AI assistant if the site's robots.txt quietly tells every AI crawler to stay out, because the AI never got to read the site. Those are the real stakes: a closed door nobody remembered closing. That check is worth five minutes before spending money on anything else, and GetLocalLeads.AI can walk you through whatever it turns up.
How to Check Both Files on Your Own Site
Checking your robots.txt for AI blocks takes one step: type yourdomain.com/robots.txt into a browser. This is the whole self-diagnostic for "is my website blocking ChatGPT," and it takes less time than reading this sentence.
A clean, healthy file usually looks like a short list of lines starting with "User-agent" (which names the bot a rule applies to) followed by "Disallow" lines (which name the folders that bot must stay out of), aimed at specific spots such as an admin login page, not the whole site. A red flag looks like a bare "Disallow: /" applied to every visitor (the slash on its own means the entire site), or a named AI bot such as GPTBot or OAI-SearchBot singled out and disallowed on its own line. If an AI bot's name shows up next to the word "Disallow," that bot is blocked, full stop.
Checking the other file works the same way: type yourdomain.com/llms.txt. Most small business sites won't have one yet, and that is completely fine. It's not a sign anything is broken, just a low-priority addition still on the list, unlike a robots.txt problem, which is never low-priority once confirmed.
What you do with what you find matters more than the check itself. A confirmed robots.txt block is a same-day fix, usually a quick conversation with whoever manages the site, the hosting provider, or the developer who built it, since the fix is often a single line removed from one file. A missing llms.txt is add-when-convenient, never an emergency, and never worth reshuffling the priority list over.
Don't confuse the two files' urgency just because they sound similar or get mentioned in the same sentence online. If what you find is hard to read, that's a fair thing to bring to a call with GetLocalLeads.AI.
What to Actually Do About It
Fix a confirmed robots.txt block first, always, before spending a minute on the other file. That single fix removes an active barrier; everything else on this page is optimization stacked on top of a door that's already open. If the site is on WordPress or a similar platform, this usually lives in a plugin's settings screen or a single editable text file, and a hosting support team can usually point to it if you're not sure where to look.
Once that's handled, add a basic llms.txt. It costs little, it doesn't hurt anything, and it positions the site for whenever adoption catches up to the hype, without pretending that day has already arrived. Nobody needs to treat it as urgent to still get it done this month.
This is exactly what the AI-visibility guarantee from GetLocalLeads.AI, an AI visibility and digital marketing agency for local and multi-location service brands, covers: making sure a client's robots.txt and llms.txt files are configured so AI can find and crawl the site, since many businesses unknowingly lock AI out of their own website. It's a narrow, specific guarantee: it covers whether AI can get in and read the site, not whether an AI tool decides to name the business afterward, which is its own separate mechanism. If checking your own two files turns up something you're not sure how to fix, that's worth a conversation, and GetLocalLeads.AI's AI visibility work is the place to start.
Frequently Asked Questions
Is llms.txt the same as robots.txt?
No, and confusing the two is the most common mistake in this topic. robots.txt controls what crawlers are allowed to access, and getting it wrong can genuinely hide a site from AI tools entirely. The other file is just a reading list of a site's important pages, with no ability to block or allow anything on its own, no matter how carefully it's written or formatted.
Do I need an llms.txt file right now?
Not urgently. It's cheap to add and doesn't hurt anything, but a large 2026 study found 97% of them are never even requested by a crawler, so it belongs after the robots.txt check, not before it, and definitely not instead of it. Treat it as something to get to this month, not something to drop everything for today.
How do I know if my robots.txt is blocking AI?
Type yourdomain.com/robots.txt into a browser and look for a bare "Disallow: /" or a named AI bot listed as disallowed on its own line. If you want a fuller picture of how AI tools treat your business right now, beyond just this one file, check where you stand next.
Will blocking one AI bot hurt my Google ranking?
Not directly. Google's ranking system and OpenAI's crawlers are separate systems entirely, and even within OpenAI, blocking the training bot doesn't affect the separate bot that powers live ChatGPT search results. Each block only affects what it specifically names, never anything wider than that.
Who usually breaks this by accident?
Most often a security plugin with an overly broad default that was never revisited, or a rule copied from a staging site that shipped to the live one by mistake. Almost nobody blocks AI crawlers on purpose; it's nearly always a setting nobody remembered was there.
The boring file wins this round. robots.txt is unglamorous and thirty years old, and it's still the one worth checking today, while llms.txt waits for the rest of the internet to catch up to the promises already written about it. Neither file will make the phone ring on its own, but one of them can quietly cost you the calls that start with an AI assistant. If you haven't already, work through the self-check guide for the fuller picture. Then get your free AI Visibility Audit and let GetLocalLeads.AI take the second look.