Back to blog
guides11 min read

llms.txt vs WebMCP: What They Are and Whether Your Site Needs Either

Google says Search ignores llms.txt. WebMCP is in a Chrome origin trial. Here is what each one does, who reads it today, and what to fix on your pages first.

Aditi ChaturvediOctober 8, 2026
TL;DR

llms.txt and WebMCP both promise to make a site readable by AI. One is a text file Google Search says it ignores. The other is a browser API in a Chrome origin trial. Neither is a ranking tactic.

  • llms.txt: A Markdown index of your key pages, proposed by Jeremy Howard in September 2024. Google's AI optimization guide lists it under things you can ignore for Google Search. Ahrefs found 97% of the files received zero requests in May 2026.
  • WebMCP: A proposed browser API that lets a page declare its forms and actions as tools an agent can call. It is a Draft Community Group Report, not a W3C standard, and Chrome runs it as an origin trial from Chrome 149.
  • What was said in Barcelona: A community speaker, Carlos Ortega Roldán, put llms.txt and markdown page copies in the not-worth-the-effort pile and pointed at semantic HTML, ARIA, heading order, schema, low CLS and WebMCP instead, according to John Campbell's recap for ROAST.
  • Why that list makes sense: Agents read a page through screenshots, the DOM and the accessibility tree. Every item on the helpful list improves one of those three views. A file at the site root improves none of them.
  • What to do: Fix crawlability, rendering and structure first. Generate an llms.txt if you want one, since it costs minutes. Prototype WebMCP on one form if you have engineers to spare, and do not ship it expecting visibility.

The useful work is the old work: pages a machine can fetch, render and parse without guessing.

Agent readiness is mostly a crawl, rendering and structure question, and those are things you can measure today. CrawlRaven joins Search Console, GA4 and a 200-point crawl into one ranked plan, from $49 at launch. Try CrawlRaven free: 1 site, no credit card →

Two ideas promise to make your site readable by AI. One is a text file that Google Search says it ignores. The other is a browser API that most SEOs had not heard of a year ago, and a Googler says he likes it.

The short version of llms.txt vs WebMCP: llms.txt is a static index nobody has to read, and almost nobody does. WebMCP is a way for a page to tell an agent what its forms do, and it is still an experiment. Neither one is a ranking tactic.

Key Takeaways

  • →llms.txt: cheap, harmless, and unsupported by any evidence that it changes what Google's AI features show.
  • →WebMCP: promising for forms and actions, but a draft in an origin trial. Prototype it. Do not ship it as a visibility play.
  • →The real work: semantic HTML, correct heading order, labelled forms, low CLS and pages that render. It serves people, crawlers and agents at once.

Two ideas, sorted into two piles

At Search Central Live Deep Dive in Barcelona, Carlos Ortega Roldán gave a lightning talk called “Preparing Your Site for Humans and Agents”. He is a freelance SEO consultant and a community speaker, not a Google employee. I was not in the room. This comes from John Campbell's Day 1 recap for ROAST.

According to that recap, he sorted the options into two piles.

  • Not worth the effort: llms.txt, markdown versions of pages, and robots.txt rules written for agents.
  • Helpful: low CLS, schema as a second confirmation of page content, correctly nested headings, semantic landmarks, ARIA markup, and a look at WebMCP for forms.

Reading the recaps, what stood out to me is how ordinary the second pile is. Most of it is standard accessibility practice. The rest of this post checks both piles against primary sources, because one speaker's list is a prompt, not proof.

llms.txt vs WebMCP, side by side

The two are not rivals. They answer different questions. llms.txt asks “which of my pages should a model read?” WebMCP asks “now that an agent is on my page, how does it act?”

Questionllms.txtWebMCP
What it isA Markdown file, usually at /llms.txt, listing key pages with short notes.A browser API that lets a page declare JavaScript functions or annotated HTML forms as tools an agent can call.
Who proposed itJeremy Howard, 3 September 2024.The W3C Web Machine Learning Community Group. Editors from Microsoft and Google.
Who reads it todayNot Google Search, by Google's own statement. Ahrefs: 97% of files got zero requests in May 2026.Agents in Chrome on sites enrolled in the origin trial, or on a machine with the testing flag enabled.
StatusCommunity proposal. No standards body.Draft Community Group Report dated 2 October 2026. Not on the W3C Standards Track. Chrome origin trial from Chrome 149.
EffortMinutes with a generator, plus upkeep when pages change.Engineering work: annotate forms or register tools, handle state, test with an agent.
What it changesNothing in Google's AI features. Possibly helps coding agents on documentation sites.How reliably an agent already on your page completes a task. Not ranking, not citations.

What is llms.txt, and is it actually used?

llms.txt is a Markdown file that lists a site's most important pages with a short note on each. The proposal at llmstxt.org was published by Jeremy Howard on 3 September 2024. Is it actually used? By Google Search, no. By AI crawlers in general, rarely.

llms.txt: in plain English

A table of contents for machines. You write a short Markdown page that says “these are my important URLs and here is what each one covers”, and you hope an AI system reads it before it reads anything else. Nothing obliges it to.

The format is small:

  • One required part. An H1 with the site or project name.
  • Optional parts. A blockquote summary, then H2 sections holding lists of links with notes.
  • A companion idea. The proposal also suggests serving a clean Markdown version of each page at a parallel .md URL. That part matters later.

What Google has said about llms.txt

Google has put this in writing. Its AI optimization guide lists llms.txt under things you can ignore for Google Search. It says you do not need new machine readable files, AI text files, markup or Markdown, “as Google Search itself doesn't use them”.

The same guide says it is fine to keep the file for other services. John Mueller has been blunter on Reddit:

  • April 2025. He compared llms.txt to the keywords meta tag, and said server logs show AI services do not check for it. Our John Mueller page links the comment.
  • 31 May 2026. In an r/SEO thread he called its usefulness “purely speculative for now”.
  • Same comment. He wrote “I like the WebMCP approach”, because it has a clear goal: the agent is already on your site, so how does it do the task properly?

That last line is the hinge of this whole comparison. It is one Googler's opinion on Reddit, not Google policy. It is also the same sorting a community speaker did on stage four months later.

Who fetches the file today

The best public data I found is an Ahrefs study published on 15 June 2026. It covers 137,210 domains in Ahrefs Web Analytics during May 2026.

  • 28% of those domains publish an llms.txt. Ahrefs calls that an upper bound, since its customers skew technical.
  • 97% of the files received zero requests in the month.
  • 1.1% of the requests that did arrive came from AI retrieval bots, the kind that fetch pages to answer a live query.
  • A fetch is not a read. Ahrefs notes every figure is a ceiling on real consumption.

I could not find any major AI provider that documents using llms.txt to decide what its search product retrieves or cites. Where the file does get used is developer documentation. The llmstxt.org page notes that OpenAI, Anthropic and Google Gemini publish one for their own developer docs, which is publishing, not consuming.

What is WebMCP, and where does it stand?

WebMCP is a proposed browser API that lets a web page declare its functions and forms as tools an AI agent can call. Instead of an agent guessing which box is the search field, the page says so. Chrome's documentation calls it a “proposed web standard”.

Be precise about its maturity, because the hype is not:

  • 10 February 2026. Chrome announced an early preview programme for prototyping.
  • Today. Chrome's docs, last updated 7 October 2026, offer an origin trial from Chrome 149 and a local testing flag.
  • The spec. The draft is a Draft Community Group Report dated 2 October 2026, with editors from Microsoft and Google. Its status section says it is “not a W3C Standard nor is it on the W3C Standards Track”.
  • The caveat. Chrome says the API is under active discussion and subject to change.

So: an origin trial in one browser, on a draft that can still change. That is a reason to learn it. It is not a reason to put it on a roadmap as a visibility tactic.

US search demand

WebMCP went from 170 searches a month to 6,600 in under a year

webmcp September 2025170
webmcp August 20266,600
webmcp 12-month average3,600
llms.txt 12-month average5,400
Monthly US searches. Source: DataForSEO, US, October 2026. Search interest measures curiosity, not adoption.

Search interest shows how fast the term arrived. US searches for “webmcp” were 170 in September 2025 and 6,600 in August 2026, with a 12-month average of 3,600. “llms.txt” averages 5,400 a month (DataForSEO, US, October 2026). Interest is curiosity. It is not adoption.

What is the difference between WebMCP and MCP?

The difference between WebMCP and MCP is where the tools live. MCP, the Model Context Protocol, connects an AI application to a server that exposes data and tools. WebMCP puts the tools in the page itself, in the user's browser tab.

  • MCP: backend. You run a server. An assistant such as Claude or ChatGPT connects to it. No web page is involved. Our comparison of SEO MCP servers covers that side.
  • WebMCP: frontend. The page declares tools in script or HTML. The agent uses them inside the tab, with the user's session and the visible interface still in play.
  • Relationship: the community group's repository describes WebMCP as a complement to backend protocols like MCP, not a replacement.

How a site exposes a form

Chrome describes two routes. The imperative API registers a tool in JavaScript, with a name, a description, an input schema and a function to run. The declarative API annotates an ordinary HTML form. The community group's declarative explainer sketches it like this:

<form toolname="search-cars"
      tooldescription="Perform a car make/model search">
  <input type="text" name="make"
         toolparamdescription="The vehicle's make (e.g., BMW, Ford)" required>
  <input type="text" name="model"
         toolparamdescription="The vehicle's model (e.g., 330i, F-150)" required>
  <button type="submit">Search</button>
</form>

Treat that as a sketch. The explainer still marks parts of the declarative design as open, and the draft report's own declarative section is unfinished. Attribute names may change before anything ships widely.

How an agent actually reads your page

This is the part that sorts the piles. Google's AI optimization guide says browser agents gather what they need by analysing visual renderings such as screenshots, inspecting the DOM structure, and interpreting the accessibility tree. Ortega built his list on the same three views, according to the ROAST recap.

How an agent reads a page

Three views of your page, and none of them is llms.txt

01
Screenshot
What the agent gets
A picture of the rendered page, read by a vision model.
What breaks it
Layout that shifts after load, so the button is no longer where the picture said it was.
What helps
Low CLS and controls that look like controls.
02
DOM
What the agent gets
The HTML structure: elements, nesting, attributes.
What breaks it
A clickable div that nothing in the markup identifies as a button.
What helps
Semantic elements, landmarks such as nav, headings nested in order.
03
Accessibility tree
What the agent gets
Roles, names and states of the interactive elements.
What breaks it
Inputs with no label, icons with no accessible name.
What helps
Labels tied to inputs, ARIA only where native HTML falls short.
The three views are named in Google's AI optimization guide. The fixes follow Carlos Ortega Roldán's community lightning talk in Barcelona, as reported in John Campbell's recap for ROAST.
  • Screenshot. A picture of the rendered page. A layout that shifts after load makes the picture wrong, which is why low CLS is on the list.
  • DOM. The HTML structure. A nav element says more than a div, and headings nested in order give the page an outline.
  • Accessibility tree. The browser's summary of roles, names and states. Unlabelled inputs and nameless icon buttons show up here as blanks.

Notice what is missing. None of the three views includes a file at your site root. An agent that is already on a product page has no reason to go and fetch an index of your other pages.

What to do instead, in priority order

This is my ordering, built on the accessibility-tree point. Start at the top and stop when you run out of time.

  1. Make sure machines can fetch and render the page. Mueller's May comment named the most basic step as not blocking agents. Check robots rules and the resources your pages load, including files on third-party hosts.
  2. Use semantic HTML. Real button, nav, main and form elements. Every label tied to its input.
  3. Fix heading nesting. One H1, then H2 and H3 in order. Our free heading analyzer shows the outline of any URL.
  4. Add ARIA where native HTML falls short. Accessible names on icon buttons, states on toggles. Not ARIA sprinkled over everything.
  5. Keep CLS low. A stable layout helps the visitor and the screenshot.
  6. Keep schema, for the right reason. Structured data earns rich results and gives a second confirmation of what the page says. Google's guide says it is not required for generative AI search.
  7. Prototype WebMCP on one form. Only if you have engineering time, and only as a test.
  8. Add llms.txt last, if at all. Our free llms.txt generator builds one in minutes.
Effort versus evidence

Six agent-readiness tasks, sorted by what backs them

Semantic HTML, headings, labelsEffort: Low to mediumEvidence: Strong→ Do first
Feeds the DOM and the accessibility tree, and it is ordinary accessibility work.
Low CLSEffort: MediumEvidence: Strong→ Do first
Already a Core Web Vital. A stable layout also keeps a screenshot accurate.
Schema markupEffort: Low to mediumEvidence: Mixed→ Keep, for the right reason
Google says it is not required for generative AI search. It still earns rich results.
WebMCPEffort: HighEvidence: Early→ Prototype, do not ship
Chrome origin trial. The draft says it can change. Worth a test on one form.
llms.txtEffort: LowEvidence: None for Google→ Optional
Cheap and harmless. Google Search says it does not use the file.
Markdown page copiesEffort: MediumEvidence: None for Google→ Skip
A second URL for every page to keep in sync and canonicalise.
The effort and evidence ratings are the author's reading of the sources cited in this post, not a measurement.

The ratings in that matrix are my reading of the sources above, not a measurement. The pattern is the point: the cheap, well-backed work is the unglamorous work.

Markdown page copies and duplicate content

The idea: serve a clean .md version of every page so a model does not have to parse your HTML. The llms.txt proposal suggests it. Ortega put it in the not-worth-it pile because it creates duplicate content, per the ROAST recap.

The practical problems are easy to list:

  • A second URL per page. Each copy can be discovered and crawled, so it needs a canonical or a noindex decision.
  • Drift. Two versions of the same content go out of sync the first time someone edits one.
  • No stated consumer. Google's guide says Search does not need Markdown. Google may still crawl such files, without treating them in a special way.

If your HTML is so heavy that a model cannot read it, the Markdown copy is a patch over the real fault. Fix the HTML, and every reader benefits.

GEO vs SEO: Google's line and the counterpoint

“geo vs seo” gets 2,900 US searches a month (DataForSEO, US, October 2026), so the question is live. In Barcelona, Google's answer was short. According to Campbell's Day 3 recap, Gary Illyes closed with three messages, the last being that AI on Google is just SEO.

  • Shared crawler. As reported by attendees, Search, AI Overviews and AI Mode all rely on Googlebot. Gemini uses a separate crawler.
  • Shared pipeline. AI features use the same processes as traditional results, which is why Google's indexing timelines apply to both.
  • The counterpoint. Thiago Pojda of SIXT, another community speaker, argued that GEO is harder to measure than SEO, and that the industry got things wrong, such as pushing JSON-LD as a way into AI answers.

Both can be true. The inputs are the same as SEO. The feedback loop is worse. That is exactly the condition under which unproven files spread: when you cannot measure, a checklist item feels like progress.

Opinion· Aditi's take
I have watched the keywords meta tag, authorship markup and a dozen other “tell the engine what you are” ideas come and go. The ones that lasted were the ones a machine could verify against the page. llms.txt is a claim about your site. Semantic HTML is the site. I would put my hours into the second, and keep an eye on WebMCP because it is at least pointed at a real task.

What this means for you

  • Agency. Do not sell llms.txt as an AI visibility deliverable. Sell the structure audit, and include the file free if the client asks.
  • In-house SEO. Put heading order, form labels and CLS into the accessibility backlog you already share with engineering. One ticket, two audiences.
  • Developer or product lead. If your site has a high-value form, WebMCP is worth a spike behind the origin trial. Budget for the API changing.
  • Founder. Generate the file if it buys peace of mind. Then read our guides to LLM SEO and ranking in ChatGPT for what the evidence supports.

What would change my mind on llms.txt

I hold this position loosely. Any one of these would move llms.txt up the list:

  • Provider documentation. A major AI search product stating in its own docs that it reads llms.txt when choosing what to retrieve or cite.
  • Log evidence. A repeat of the Ahrefs study showing retrieval bots, not SEO tools, as the main fetchers of the file.
  • A controlled test. Matched sites with and without the file, published with method and data, showing a difference in citations.
  • A change from Google. The AI optimization guide dropping llms.txt from its ignore list.
  • A request from a platform that sends you customers. Mueller's own test. If one asks, make the file that day.

None of these has happened yet, as far as I can find. When one does, this post gets rewritten.

Where CrawlRaven fits

CrawlRaven does not track prompts or citations. It joins Search Console, GA4 and a 200-point crawl into one ranked plan, and its AI search readiness view is a diagnosis built from those three sources: which pages are hard to fetch, render or parse, ranked by the traffic at stake.

Sources used in this post

Frequently asked questions

What is an llms.txt file?

An llms.txt file is a Markdown document, usually served at /llms.txt, that lists a site's most important pages with a short note on each. Jeremy Howard proposed the format on 3 September 2024. It needs only an H1 with the site name; a summary and sections of links are optional. It is a community proposal with no standards body behind it, and it controls and blocks nothing.

Is llms.txt actually used?

Barely, on the evidence available. Google's AI optimization guide says Google Search does not use llms.txt files. Ahrefs studied 137,210 domains and found 97% of llms.txt files received zero requests in May 2026, with AI retrieval bots making 1.1% of the requests that did arrive. The file does see real use in developer documentation, where coding agents are pointed at it.

Is llms.txt mandatory?

No. llms.txt is not required by Google or by any AI provider that has documented its crawling. Google's guidance says you do not need to create machine readable files, AI text files or Markdown to appear in Google Search, including its generative AI features. John Mueller's rule of thumb on Reddit was to create one if a platform that sends you customers asks for it.

Does llms.txt hurt SEO?

No. Google's guide says it is fine to create and maintain llms.txt files for other services, and that Google Search does not treat the file in a special way. The cost is not a penalty. The cost is attention: an hour spent on a file nothing fetches is an hour not spent on rendering, structure or crawl errors.

What is WebMCP?

WebMCP is a proposed browser API that lets a web page declare its functions and forms as tools an AI agent can call directly, instead of the agent guessing from the interface. It is published as a Draft Community Group Report by the W3C Web Machine Learning Community Group, with editors from Microsoft and Google. It is not a W3C standard. Chrome offers it as an origin trial from Chrome 149.

What is the difference between WebMCP and MCP?

MCP, the Model Context Protocol, connects an AI application to a server that exposes data and tools on the backend. WebMCP runs inside the browser tab: the page itself declares tools in client-side script or HTML, and the agent uses them while the user watches. The WebMCP authors describe it as a complement to MCP, not a replacement.

How do I enable WebMCP in Chrome?

For local testing, open chrome://flags/#enable-webmcp-testing, set the flag to Enabled and relaunch Chrome. To test with real visitors, register for the WebMCP origin trial, which Chrome's documentation says starts from Chrome 149. Chrome describes the API as under active discussion and subject to change, so treat anything you build as a prototype.

Is WebMCP a ranking factor?

No. WebMCP has nothing to do with how Google ranks or cites pages. It only matters once an agent is already on your page and trying to complete a task, such as filling in a form. Chrome's documentation frames it as an API for browser workflows with a human in the loop, and makes no claim about search visibility.

Is GEO different from SEO?

For Google's AI features, not in the technical work. Google's guide says optimising for generative AI search is still SEO, and attendee recaps of Search Central Live in Barcelona report the same closing message. The measurement is different, though. A community speaker at the same event, Thiago Pojda of SIXT, argued that GEO is harder to measure than SEO.

Aditi Chaturvedi
About the Author

Aditi Chaturvedi

15+ years of growing SaaS websites through SEO | Author, 200-Point Audit Checklist

Aditi has spent 15+ years helping SaaS companies scale organic traffic through technical SEO and content strategy. She is the author of the CrawlRaven 200-Point Audit checklist used by agencies and in-house teams to systematically improve search performance.

llms.txtwebmcpis llms.txt actually usedwebmcp vs mcpgeo vs seogenerative engine optimizationai agentsaccessibility treesemantic htmlai search

Most agent-readiness advice asks you to add a new file. Most agent-readiness problems are old ones: blocked resources, broken heading order, unlabelled forms.

Find the structure problems before you add a new file

CrawlRaven joins Search Console, GA4 and a 200-point crawl into one ranked plan, so you see which pages are hard to fetch, render or parse. Start free with 1 site; lifetime licences from a one-time $49 at launch. Lifetime pricing steps up as licenses sell, so check the pricing page for the current batch.

An llms.txt file takes minutes and a WebMCP prototype takes a sprint. Neither helps a page that a machine cannot read cleanly in the first place. Start there.

CrawlRaven connects Google Search Console and GA4, runs 200+ technical SEO checks, and joins all three into one prioritized fix list, so you know what is broken, what it is costing you, and what to fix first.

✓ No credit card required·200+ checks·GSC + GA4 + full-site crawl
Free plan — no credit card

Stop exporting. Start shipping.

Connect Search Console, import your Ahrefs or Semrush lists, and get one ranked plan. Start free with one site, or grab a limited lifetime deal from $39, only 3 licenses left.

3
Data sources joined
200+
Point audit checks
1
Ranked plan out