llms.txt vs WebMCP: What They Are and Whether Your Site Needs Either
Google says Search ignores llms.txt. WebMCP is in a Chrome origin trial. Here is what each one does, who reads it today, and what to fix on your pages first.
llms.txt and WebMCP both promise to make a site readable by AI. One is a text file Google Search says it ignores. The other is a browser API in a Chrome origin trial. Neither is a ranking tactic.
- llms.txt: A Markdown index of your key pages, proposed by Jeremy Howard in September 2024. Google's AI optimization guide lists it under things you can ignore for Google Search. Ahrefs found 97% of the files received zero requests in May 2026.
- WebMCP: A proposed browser API that lets a page declare its forms and actions as tools an agent can call. It is a Draft Community Group Report, not a W3C standard, and Chrome runs it as an origin trial from Chrome 149.
- What was said in Barcelona: A community speaker, Carlos Ortega Roldán, put llms.txt and markdown page copies in the not-worth-the-effort pile and pointed at semantic HTML, ARIA, heading order, schema, low CLS and WebMCP instead, according to John Campbell's recap for ROAST.
- Why that list makes sense: Agents read a page through screenshots, the DOM and the accessibility tree. Every item on the helpful list improves one of those three views. A file at the site root improves none of them.
- What to do: Fix crawlability, rendering and structure first. Generate an llms.txt if you want one, since it costs minutes. Prototype WebMCP on one form if you have engineers to spare, and do not ship it expecting visibility.
The useful work is the old work: pages a machine can fetch, render and parse without guessing.
Agent readiness is mostly a crawl, rendering and structure question, and those are things you can measure today. CrawlRaven joins Search Console, GA4 and a 200-point crawl into one ranked plan, from $49 at launch. Try CrawlRaven free: 1 site, no credit card →
Two ideas promise to make your site readable by AI. One is a text file that Google Search says it ignores. The other is a browser API that most SEOs had not heard of a year ago, and a Googler says he likes it.
The short version of llms.txt vs WebMCP: llms.txt is a static index nobody has to read, and almost nobody does. WebMCP is a way for a page to tell an agent what its forms do, and it is still an experiment. Neither one is a ranking tactic.
Key Takeaways
- →llms.txt: cheap, harmless, and unsupported by any evidence that it changes what Google's AI features show.
- →WebMCP: promising for forms and actions, but a draft in an origin trial. Prototype it. Do not ship it as a visibility play.
- →The real work: semantic HTML, correct heading order, labelled forms, low CLS and pages that render. It serves people, crawlers and agents at once.
Two ideas, sorted into two piles
At Search Central Live Deep Dive in Barcelona, Carlos Ortega Roldán gave a lightning talk called “Preparing Your Site for Humans and Agents”. He is a freelance SEO consultant and a community speaker, not a Google employee. I was not in the room. This comes from John Campbell's Day 1 recap for ROAST.
According to that recap, he sorted the options into two piles.
- Not worth the effort: llms.txt, markdown versions of pages, and robots.txt rules written for agents.
- Helpful: low CLS, schema as a second confirmation of page content, correctly nested headings, semantic landmarks, ARIA markup, and a look at WebMCP for forms.
Reading the recaps, what stood out to me is how ordinary the second pile is. Most of it is standard accessibility practice. The rest of this post checks both piles against primary sources, because one speaker's list is a prompt, not proof.
llms.txt vs WebMCP, side by side
The two are not rivals. They answer different questions. llms.txt asks “which of my pages should a model read?” WebMCP asks “now that an agent is on my page, how does it act?”
| Question | llms.txt | WebMCP |
|---|---|---|
| What it is | A Markdown file, usually at /llms.txt, listing key pages with short notes. | A browser API that lets a page declare JavaScript functions or annotated HTML forms as tools an agent can call. |
| Who proposed it | Jeremy Howard, 3 September 2024. | The W3C Web Machine Learning Community Group. Editors from Microsoft and Google. |
| Who reads it today | Not Google Search, by Google's own statement. Ahrefs: 97% of files got zero requests in May 2026. | Agents in Chrome on sites enrolled in the origin trial, or on a machine with the testing flag enabled. |
| Status | Community proposal. No standards body. | Draft Community Group Report dated 2 October 2026. Not on the W3C Standards Track. Chrome origin trial from Chrome 149. |
| Effort | Minutes with a generator, plus upkeep when pages change. | Engineering work: annotate forms or register tools, handle state, test with an agent. |
| What it changes | Nothing in Google's AI features. Possibly helps coding agents on documentation sites. | How reliably an agent already on your page completes a task. Not ranking, not citations. |
What is llms.txt, and is it actually used?
llms.txt is a Markdown file that lists a site's most important pages with a short note on each. The proposal at llmstxt.org was published by Jeremy Howard on 3 September 2024. Is it actually used? By Google Search, no. By AI crawlers in general, rarely.
A table of contents for machines. You write a short Markdown page that says “these are my important URLs and here is what each one covers”, and you hope an AI system reads it before it reads anything else. Nothing obliges it to.
The format is small:
- One required part. An H1 with the site or project name.
- Optional parts. A blockquote summary, then H2 sections holding lists of links with notes.
- A companion idea. The proposal also suggests serving a clean Markdown version of each page at a parallel
.mdURL. That part matters later.
What Google has said about llms.txt
Google has put this in writing. Its AI optimization guide lists llms.txt under things you can ignore for Google Search. It says you do not need new machine readable files, AI text files, markup or Markdown, “as Google Search itself doesn't use them”.
The same guide says it is fine to keep the file for other services. John Mueller has been blunter on Reddit:
- April 2025. He compared llms.txt to the keywords meta tag, and said server logs show AI services do not check for it. Our John Mueller page links the comment.
- 31 May 2026. In an r/SEO thread he called its usefulness “purely speculative for now”.
- Same comment. He wrote “I like the WebMCP approach”, because it has a clear goal: the agent is already on your site, so how does it do the task properly?
That last line is the hinge of this whole comparison. It is one Googler's opinion on Reddit, not Google policy. It is also the same sorting a community speaker did on stage four months later.
Who fetches the file today
The best public data I found is an Ahrefs study published on 15 June 2026. It covers 137,210 domains in Ahrefs Web Analytics during May 2026.
- 28% of those domains publish an llms.txt. Ahrefs calls that an upper bound, since its customers skew technical.
- 97% of the files received zero requests in the month.
- 1.1% of the requests that did arrive came from AI retrieval bots, the kind that fetch pages to answer a live query.
- A fetch is not a read. Ahrefs notes every figure is a ceiling on real consumption.
I could not find any major AI provider that documents using llms.txt to decide what its search product retrieves or cites. Where the file does get used is developer documentation. The llmstxt.org page notes that OpenAI, Anthropic and Google Gemini publish one for their own developer docs, which is publishing, not consuming.
What is WebMCP, and where does it stand?
WebMCP is a proposed browser API that lets a web page declare its functions and forms as tools an AI agent can call. Instead of an agent guessing which box is the search field, the page says so. Chrome's documentation calls it a “proposed web standard”.
Be precise about its maturity, because the hype is not:
- 10 February 2026. Chrome announced an early preview programme for prototyping.
- Today. Chrome's docs, last updated 7 October 2026, offer an origin trial from Chrome 149 and a local testing flag.
- The spec. The draft is a Draft Community Group Report dated 2 October 2026, with editors from Microsoft and Google. Its status section says it is “not a W3C Standard nor is it on the W3C Standards Track”.
- The caveat. Chrome says the API is under active discussion and subject to change.
So: an origin trial in one browser, on a draft that can still change. That is a reason to learn it. It is not a reason to put it on a roadmap as a visibility tactic.
WebMCP went from 170 searches a month to 6,600 in under a year
Search interest shows how fast the term arrived. US searches for “webmcp” were 170 in September 2025 and 6,600 in August 2026, with a 12-month average of 3,600. “llms.txt” averages 5,400 a month (DataForSEO, US, October 2026). Interest is curiosity. It is not adoption.
What is the difference between WebMCP and MCP?
The difference between WebMCP and MCP is where the tools live. MCP, the Model Context Protocol, connects an AI application to a server that exposes data and tools. WebMCP puts the tools in the page itself, in the user's browser tab.
- MCP: backend. You run a server. An assistant such as Claude or ChatGPT connects to it. No web page is involved. Our comparison of SEO MCP servers covers that side.
- WebMCP: frontend. The page declares tools in script or HTML. The agent uses them inside the tab, with the user's session and the visible interface still in play.
- Relationship: the community group's repository describes WebMCP as a complement to backend protocols like MCP, not a replacement.
How a site exposes a form
Chrome describes two routes. The imperative API registers a tool in JavaScript, with a name, a description, an input schema and a function to run. The declarative API annotates an ordinary HTML form. The community group's declarative explainer sketches it like this:
<form toolname="search-cars"
tooldescription="Perform a car make/model search">
<input type="text" name="make"
toolparamdescription="The vehicle's make (e.g., BMW, Ford)" required>
<input type="text" name="model"
toolparamdescription="The vehicle's model (e.g., 330i, F-150)" required>
<button type="submit">Search</button>
</form>Treat that as a sketch. The explainer still marks parts of the declarative design as open, and the draft report's own declarative section is unfinished. Attribute names may change before anything ships widely.
How an agent actually reads your page
This is the part that sorts the piles. Google's AI optimization guide says browser agents gather what they need by analysing visual renderings such as screenshots, inspecting the DOM structure, and interpreting the accessibility tree. Ortega built his list on the same three views, according to the ROAST recap.
Three views of your page, and none of them is llms.txt
- Screenshot. A picture of the rendered page. A layout that shifts after load makes the picture wrong, which is why low CLS is on the list.
- DOM. The HTML structure. A
navelement says more than adiv, and headings nested in order give the page an outline. - Accessibility tree. The browser's summary of roles, names and states. Unlabelled inputs and nameless icon buttons show up here as blanks.
Notice what is missing. None of the three views includes a file at your site root. An agent that is already on a product page has no reason to go and fetch an index of your other pages.
What to do instead, in priority order
This is my ordering, built on the accessibility-tree point. Start at the top and stop when you run out of time.
- Make sure machines can fetch and render the page. Mueller's May comment named the most basic step as not blocking agents. Check robots rules and the resources your pages load, including files on third-party hosts.
- Use semantic HTML. Real
button,nav,mainandformelements. Every label tied to its input. - Fix heading nesting. One H1, then H2 and H3 in order. Our free heading analyzer shows the outline of any URL.
- Add ARIA where native HTML falls short. Accessible names on icon buttons, states on toggles. Not ARIA sprinkled over everything.
- Keep CLS low. A stable layout helps the visitor and the screenshot.
- Keep schema, for the right reason. Structured data earns rich results and gives a second confirmation of what the page says. Google's guide says it is not required for generative AI search.
- Prototype WebMCP on one form. Only if you have engineering time, and only as a test.
- Add llms.txt last, if at all. Our free llms.txt generator builds one in minutes.
Six agent-readiness tasks, sorted by what backs them
The ratings in that matrix are my reading of the sources above, not a measurement. The pattern is the point: the cheap, well-backed work is the unglamorous work.
Markdown page copies and duplicate content
The idea: serve a clean .md version of every page so a model does not have to parse your HTML. The llms.txt proposal suggests it. Ortega put it in the not-worth-it pile because it creates duplicate content, per the ROAST recap.
The practical problems are easy to list:
- A second URL per page. Each copy can be discovered and crawled, so it needs a canonical or a noindex decision.
- Drift. Two versions of the same content go out of sync the first time someone edits one.
- No stated consumer. Google's guide says Search does not need Markdown. Google may still crawl such files, without treating them in a special way.
If your HTML is so heavy that a model cannot read it, the Markdown copy is a patch over the real fault. Fix the HTML, and every reader benefits.
GEO vs SEO: Google's line and the counterpoint
“geo vs seo” gets 2,900 US searches a month (DataForSEO, US, October 2026), so the question is live. In Barcelona, Google's answer was short. According to Campbell's Day 3 recap, Gary Illyes closed with three messages, the last being that AI on Google is just SEO.
- Shared crawler. As reported by attendees, Search, AI Overviews and AI Mode all rely on Googlebot. Gemini uses a separate crawler.
- Shared pipeline. AI features use the same processes as traditional results, which is why Google's indexing timelines apply to both.
- The counterpoint. Thiago Pojda of SIXT, another community speaker, argued that GEO is harder to measure than SEO, and that the industry got things wrong, such as pushing JSON-LD as a way into AI answers.
Both can be true. The inputs are the same as SEO. The feedback loop is worse. That is exactly the condition under which unproven files spread: when you cannot measure, a checklist item feels like progress.
What this means for you
- Agency. Do not sell llms.txt as an AI visibility deliverable. Sell the structure audit, and include the file free if the client asks.
- In-house SEO. Put heading order, form labels and CLS into the accessibility backlog you already share with engineering. One ticket, two audiences.
- Developer or product lead. If your site has a high-value form, WebMCP is worth a spike behind the origin trial. Budget for the API changing.
- Founder. Generate the file if it buys peace of mind. Then read our guides to LLM SEO and ranking in ChatGPT for what the evidence supports.
What would change my mind on llms.txt
I hold this position loosely. Any one of these would move llms.txt up the list:
- Provider documentation. A major AI search product stating in its own docs that it reads llms.txt when choosing what to retrieve or cite.
- Log evidence. A repeat of the Ahrefs study showing retrieval bots, not SEO tools, as the main fetchers of the file.
- A controlled test. Matched sites with and without the file, published with method and data, showing a difference in citations.
- A change from Google. The AI optimization guide dropping llms.txt from its ignore list.
- A request from a platform that sends you customers. Mueller's own test. If one asks, make the file that day.
None of these has happened yet, as far as I can find. When one does, this post gets rewritten.
Where CrawlRaven fits
CrawlRaven does not track prompts or citations. It joins Search Console, GA4 and a 200-point crawl into one ranked plan, and its AI search readiness view is a diagnosis built from those three sources: which pages are hard to fetch, render or parse, ranked by the traffic at stake.
Sources used in this post
- llmstxt.org: the llms.txt proposal by Jeremy Howard, 3 September 2024.
- Google Search Central, AI optimization guide: last updated 10 July 2026.
- John Mueller on Reddit, r/SEO: 31 May 2026.
- Ahrefs, llms.txt study of 137,210 domains: 15 June 2026.
- Chrome for Developers, WebMCP: last updated 7 October 2026. The early preview announcement is dated 10 February 2026.
- W3C Web Machine Learning Community Group, WebMCP draft report: 2 October 2026, with the group's repository and explainers.
- ROAST, Barcelona Day 1 recap by John Campbell: 30 September 2026.
- ROAST, Barcelona Day 3 recap by John Campbell: 2 October 2026.
Related reading on CrawlRaven
- How long Google takes to index a page: the timing tables from the same event.
- robots.txt on third-party hosts: why a file on another domain can block your rendering.
- LLM SEO: how AI systems pick the pages they cite.
- Best SEO MCP servers: the backend side of MCP.
- llms.txt generator: free, if you want the file anyway.
Frequently asked questions
What is an llms.txt file?
An llms.txt file is a Markdown document, usually served at /llms.txt, that lists a site's most important pages with a short note on each. Jeremy Howard proposed the format on 3 September 2024. It needs only an H1 with the site name; a summary and sections of links are optional. It is a community proposal with no standards body behind it, and it controls and blocks nothing.
Is llms.txt actually used?
Barely, on the evidence available. Google's AI optimization guide says Google Search does not use llms.txt files. Ahrefs studied 137,210 domains and found 97% of llms.txt files received zero requests in May 2026, with AI retrieval bots making 1.1% of the requests that did arrive. The file does see real use in developer documentation, where coding agents are pointed at it.
Is llms.txt mandatory?
No. llms.txt is not required by Google or by any AI provider that has documented its crawling. Google's guidance says you do not need to create machine readable files, AI text files or Markdown to appear in Google Search, including its generative AI features. John Mueller's rule of thumb on Reddit was to create one if a platform that sends you customers asks for it.
Does llms.txt hurt SEO?
No. Google's guide says it is fine to create and maintain llms.txt files for other services, and that Google Search does not treat the file in a special way. The cost is not a penalty. The cost is attention: an hour spent on a file nothing fetches is an hour not spent on rendering, structure or crawl errors.
What is WebMCP?
WebMCP is a proposed browser API that lets a web page declare its functions and forms as tools an AI agent can call directly, instead of the agent guessing from the interface. It is published as a Draft Community Group Report by the W3C Web Machine Learning Community Group, with editors from Microsoft and Google. It is not a W3C standard. Chrome offers it as an origin trial from Chrome 149.
What is the difference between WebMCP and MCP?
MCP, the Model Context Protocol, connects an AI application to a server that exposes data and tools on the backend. WebMCP runs inside the browser tab: the page itself declares tools in client-side script or HTML, and the agent uses them while the user watches. The WebMCP authors describe it as a complement to MCP, not a replacement.
How do I enable WebMCP in Chrome?
For local testing, open chrome://flags/#enable-webmcp-testing, set the flag to Enabled and relaunch Chrome. To test with real visitors, register for the WebMCP origin trial, which Chrome's documentation says starts from Chrome 149. Chrome describes the API as under active discussion and subject to change, so treat anything you build as a prototype.
Is WebMCP a ranking factor?
No. WebMCP has nothing to do with how Google ranks or cites pages. It only matters once an agent is already on your page and trying to complete a task, such as filling in a form. Chrome's documentation frames it as an API for browser workflows with a human in the loop, and makes no claim about search visibility.
Is GEO different from SEO?
For Google's AI features, not in the technical work. Google's guide says optimising for generative AI search is still SEO, and attendee recaps of Search Central Live in Barcelona report the same closing message. The measurement is different, though. A community speaker at the same event, Thiago Pojda of SIXT, argued that GEO is harder to measure than SEO.
15+ years of growing SaaS websites through SEO | Author, 200-Point Audit Checklist
Aditi has spent 15+ years helping SaaS companies scale organic traffic through technical SEO and content strategy. She is the author of the CrawlRaven 200-Point Audit checklist used by agencies and in-house teams to systematically improve search performance.