Registered 20 July 2026Tilio Research

Does anything actually read llms.txt?

Pre-registration. No results yet.

Registered 20 July 2026. Results published on or after 30 September 2026, whatever they show.

Research question

In normal operation, do AI systems fetch llms.txt at all?

3
Paths tracked
17 Jul 2026
Window opens
30 Sep 2026
Window closes
Pending
Results

The question

The llms.txt standard proposes that websites publish a plain-text file telling AI models what a site contains and how to read it. Adoption has been fast. Actual use has never been convincingly shown. No major AI provider has confirmed its crawlers fetch the file, and the standard carries no formal endorsement from OpenAI, Anthropic, Google or Perplexity.

We’re testing the simplest version of the question that matters to anyone deciding whether it’s worth having: in normal operation, do AI systems fetch llms.txt at all?

What we predict

Our expectation going in is that AI crawlers fetch llms.txt rarely or never, while fetching established files like robots.txt and sitemap.xml routinely. We’re stating that now so the result can’t be quietly fitted to whatever we find. If the data proves us wrong, that’s the more interesting outcome and we’ll lead with it.

How we're measuring it

We log every request to our web properties at the edge, capturing the requesting user agent, the path, and the timestamp. From that log we isolate requests from identified AI crawlers (GPTBot, ChatGPT-User, ClaudeBot, Claude-User, PerplexityBot, Perplexity-User, Google-Extended and others as they appear) and count how many reached each of three paths:

  • /llms.txt (and /llms-full.txt where served)
  • /robots.txt
  • /sitemap.xml

robots.txt and sitemap.xml are the control. They’re long-established files every serious crawler understands, so they show us what “a crawler paying attention to this site” looks like. llms.txt is the file on trial. The comparison is the whole study: not whether llms.txt gets fetched in isolation, but whether it gets fetched relative to the files we know are fetched.

Every site in the sample serves a valid, discoverable llms.txt for the whole window, so a null result can’t be blamed on a missing or broken file.

Sample

The study runs on Tilio-operated and consenting client web properties where we have edge-level request logging. It opens with tilio.co.uk and expands to further sites as consent and logging allow during the window. The final site list, with a one-line note on each, is published with the results. This is a small, self-assembled sample rather than a representative census, which shapes what it can and can’t claim (see Limitations).

Window

Measurement runs from 17 July 2026 to 30 September 2026. The window was fixed in advance and won’t be extended to reach a better number. Sites added mid-window are counted only from their own logging start date, recorded per site.

Analysis

On close we’ll report, per crawler and in aggregate: total requests to each of the three paths, the ratio of llms.txt fetches to robots.txt and sitemap.xml fetches, and how many distinct AI crawlers touched llms.txt at least once. We’ll show the per-site breakdown alongside the pooled figures, so nobody has to take the aggregate on trust. The only data excluded is non-AI traffic, which is out of scope by definition.

Limitations, stated before we start

  • Edge logs capture crawlers that identify themselves. A bot fetching llms.txt while presenting as an ordinary browser wouldn’t be counted, so this measures declared AI crawler behaviour, not all possible machine access.
  • The sample is small and self-selected. It’s direct evidence of what happens on these sites, not a population estimate for the web.
  • No fetches over this window is evidence that fetching is rare in this setting, not proof that no system anywhere reads the file.
  • A fetch tells us the file was retrieved, not that its contents changed any model’s output. Consumption is a lower bar than usefulness, and we’re only testing the lower bar.

What we'll do with the result

We publish either way. A null result is the finding if that’s what the logs show. A positive result is the finding if it isn’t. Registering this before we look is our commitment to both.

Cite this study

Tilio Research. (2026). Does anything actually read llms.txt? [pre-registration, results pending]. Tilio. https://www.tilio.co.uk/research/does-anything-actually-read-llms-txt