SEO & AI Search

Controlling AI Crawlers: robots.txt, Not llms.txt

IT-Fachkraft prüft mit einem Tablet einen Serverschrank – Sinnbild für die Steuerung von KI-Crawlern über die robots.txt

Images: created using AI

Short answer: In 2026 you control AI crawlers through robots.txt, not llms.txt. Only robots.txt is demonstrably honoured by OpenAI, Google and Anthropic. The critical distinction is between training bots and search bots: block both and you disappear from AI answers.

Ever since AI assistants started citing sources, server logs have filled with names nobody knew three years ago: GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot. Most companies react in one of two reflexive ways – block everything, because “they are training on our content”, or do nothing, because it looks complicated. Both are expensive. This article covers which bots exist, what you can actually control, and why the much-discussed llms.txt file does not solve your problem.

Training, searching, fetching: three different things

The most important distinction gets lost in almost every discussion. AI providers run several bots with different jobs:

  • Training crawlers collect content that future models learn from. Blocking them has no effect on whether you appear in answers.
  • Search crawlers build the index the assistant cites from. Blocking them removes you from the source list – that is the painful case.
  • User-initiated fetches happen when someone pastes a link into a chat or the assistant opens a page on a person’s behalf. These behave differently from classic crawling.

Blocking everything is therefore an unintentional decision against your own visibility. How that visibility is built in the first place is covered in GEO and AEO: getting cited by AI answers.

The bots that matter in 2026

Bot Operator Purpose Controllable via robots.txt Effect of blocking
GPTBot OpenAI Training future models yes no effect on visibility in ChatGPT answers
OAI-SearchBot OpenAI Search index for ChatGPT yes you stop appearing as a source
ChatGPT-User OpenAI Fetch on behalf of a user limited – the documentation notes robots.txt rules may not apply hard to predict
OAI-AdsBot OpenAI Advertising-related fetches yes affects advertising contexts only
ClaudeBot Anthropic Training yes no effect on visibility
Claude-SearchBot Anthropic Search index yes you stop appearing as a source
Claude-User Anthropic Fetch on behalf of a user yes users can no longer pull your page in
Google-Extended Google Training and grounding for Gemini yes (robots.txt token only, no separate user agent) no effect on Google Search
As of: August 2026. Sources: OpenAI bot documentation, Anthropic help article, Google crawler documentation.

Two details deserve attention. First, Google states in its own documentation that Google-Extended has no bearing on whether a site is included in Google Search and is not used as a ranking signal. You can block Gemini training without risking classic visibility. Second, robots.txt changes do not take effect immediately – OpenAI cites roughly 24 hours.

What llms.txt is, and what it cannot do

The llms.txt file goes back to a proposal by Jeremy Howard from September 2024 and has been at version 2 since 10 August 2026. The idea: a Markdown file telling a language model in compact form what a site contains and where machine-readable versions live. Version 2 additionally allows the file at arbitrary paths and links to Markdown variants of individual pages.

Sensible as that sounds, it is no substitute for robots.txt, for a simple reason: llmstxt.org itself calls the format a proposal, not a standard. No major provider has committed to reading it. Google is explicitly against it: in its official guide to optimising for AI features, Google writes that no additional machine-readable AI files are needed and that Google Search ignores them. John Mueller was quoted in 2026 describing the topic as purely speculative for now.

Our recommendation is unspectacular: an llms.txt does no harm and takes twenty minutes. Create one if you like. Just do not expect it to generate visibility. The work that actually pays sits in clean content, clear structure and a robots.txt that lets the right bots through.

A robots.txt that makes sense in 2026

For most mid-sized business websites, the right posture is: allow search bots, decide on training bots according to your business model.

Allow training crawlers if you build reach through content and your text is public anyway. Block them if your content is the product – specialist databases, paid guides, extensive reference catalogues. And block, in every case, the areas that should never reach an index: customer portals, internal search result pages, baskets, print views.

Three rules from practice:

  • Never open with a global disallow. A User-agent: * line with Disallow: / at the top of the file makes everything below it pointless and can cost you your entire visibility.
  • Spell bot names exactly. The tokens are fixed. GPT-Bot or Claude-Bot with a hyphen in the wrong place does nothing.
  • IP blocking is not a substitute. Anthropic states in its own help article that blocking IP ranges is not a workable approach, because addresses change.

If you are unsure what your file currently allows, open your domain with /robots.txt appended and read it line by line. Those five minutes are well spent.

A worked example: what a wrong disallow costs

A trades business receives 40 enquiries a month through its website. Its own analytics attribute 6 of them to referrals from AI answers, or 15 percent. Average order value is EUR 2,800 and the close rate is 20 percent.

  • That channel produces 6 × 0.2 = 1.2 orders a month, or roughly EUR 3,360 in order value.
  • Over twelve months that is EUR 40,320 lost to a blanket disallow for all AI bots.

Against that stands the server load usually cited as the reason for blocking. Run the numbers honestly: 600 URLs, five AI crawlers, two fetches per URL per month gives 6,000 requests. At an average of 1.4 megabytes per page that is around 8.4 gigabytes a month – not a measurable quantity on standard hosting.

The calculation only flips in one case: when a single bot misbehaves and fetches the same URLs a hundred times a day. The targeted remedy for that is a crawl delay for that specific bot, not a general ban. Anthropic explicitly supports this non-standard directive according to its own documentation.

How to measure how many visitors actually arrive from AI answers is covered in getting found as an SME in Google and AI search.

Stay visible in AI answers
We analyse your server logs for AI bots, review your robots.txt and configure it so training access is controlled while your visibility as a cited source stays intact.

See SEO and AI search Have a quick chat

What robots.txt will not achieve

Three honest limits that sit between the lines of vendor documentation:

It does not stop content being taken by third parties. robots.txt is a request, not a fence. The major providers honour it; obscure services may not. Technical protection requires access control.

It does not remove content from models already trained. A block set today works forwards. What sits inside a model trained two years ago stays there.

It controls user-initiated fetches only partially. When somebody pastes your URL into a chat, the assistant fetches the page on that person’s behalf. OpenAI notes in its documentation that robots.txt rules may not apply in that case – an area not yet fully settled, technically or legally.

A four-step approach

  1. Take stock. Search the last 30 days of server logs for the bot tokens. You will see which bots actually affect you – often only three or four.
  2. Make the decision. Allow or block training bots? That call belongs to management, not IT, because it is a question about your business model.
  3. Write the robots.txt. Let search bots through, block sensitive directories, spell the tokens exactly, then check the file in a browser.
  4. Review after four weeks. Back into the logs: are the bots complying? And check in Search Console that classic visibility is unchanged.

For sites we look after, this pass is part of ongoing operation – see managed service. If you want to build visibility in AI answers deliberately, the framework sits under SEO and AI search.

Frequently asked questions

Should I block GPTBot?

Only if your content is the product. GPTBot gathers material for future training and, according to OpenAI, has nothing to do with whether your site appears as a source in answers. That is OAI-SearchBot’s job – and you should normally let it through.

Do I need an llms.txt?

It is not required. Google writes in its own documentation that additional AI files are unnecessary and that Search ignores them. If you do create one, keep it current – an outdated content index is worse than none.

Does blocking Google-Extended cost rankings?

No. Google states plainly in its crawler documentation that Google-Extended neither affects inclusion in Google Search nor serves as a ranking signal.

How do I tell whether a bot is genuine?

Through the providers’ published IP ranges. OpenAI publishes a separate address list for each bot. When in doubt, check the requesting address against that list – spoofed user agents are common.

How long does a change take to work?

OpenAI cites roughly 24 hours after a robots.txt change. Assume similar windows elsewhere and do not schedule changes at the last minute.

Next step

If you want to know which AI bots visit your site and whether your current robots.txt is quietly making you invisible, we will look at it with you. Start with our AI check or write to us via the contact form.

Share
Newsletter

One real-world process, once a month

We only write when we have something substantial: a process we built ourselves, what it cost, what it saves — and where it got stuck. If nothing usable comes together in a given month, you get no email that month. You can unsubscribe from any email with one click.

  • Around once a month
  • Unsubscribe with one click
  • No sharing with third parties

Back to top

Get started

Ready to really put AI to work?

Tell us about your project in a few minutes — you’ll get an honest initial assessment, with no sales pressure.

  • Free & no obligation
  • Reply usually within 1 business day
  • GDPR-compliant