Back to blog
technical seo10 min read

How to Check the robots.txt Files You Don't Control

Scripts and CSS served from a CDN or subdomain follow that host's robots.txt, not yours. A six-step audit to find what Google cannot load and fix the render.

Aditi ChaturvediOctober 8, 2026
TL;DR

Every file a page loads is checked against the robots.txt of the host that serves it. To find out whether someone else's file is breaking how Google renders your pages:

  1. List the hosts: Open one URL per template in DevTools, Network tab, and note every domain that serves CSS, JavaScript or data.
  2. Read each host's robots.txt: Fetch /robots.txt on every host and match the exact resource paths against its rules. A 404 means nothing is blocked.
  3. Compare the render: Run a URL Inspection live test and compare Google's screenshot and rendered HTML with your browser.
  4. Fix or work around: Edit the hosts you own. For the rest, self-host or inline the resource, or change provider.
  5. Set robots meta, then wait a day: Add max-image-preview:large and max-snippet:-1, then re-test after about 24 hours, which is how long Google caches robots.txt.

"robots.txt" gets 6,600 US searches a month (DataForSEO, US, October 2026). Almost none of the answers mention that the file in question may belong to someone else.

Your robots.txt can be spotless while Google still fails to build the page, because the scripts and styles come from hosts with robots.txt files of their own. This guide walks the six-step audit across every host, using Google's documentation and the attendee recaps from Search Central Live in Barcelona. Try CrawlRaven free: 1 site, no credit card →

The short answer: the robots.txt that blocks your render may not be yours

Search for "robots.txt blocking javascript and css" and the top results are about ten years old. They all tell you to check your own file. Reading the recaps from Google's Search Central Live Deep Dive in Barcelona (30 September to 2 October 2026), what stood out to me was a lightning talk that points somewhere else.

Dave Smart of Tame the Bots said that when a page loads files from another domain, such as a CDN or a subdomain, "it's that domain's robots.txt that applies, not yours", in the words of John Campbell's Day 1 recap for ROAST. Your file can pass every check while one you have never opened breaks the render.

So the how-to is this: list every host your templates load from, read each host's robots.txt, and compare what Google renders with what a visitor sees. Set aside an hour or two the first time. You need DevTools and Search Console access.

Key Takeaways

  • →robots.txt is per host: Google's documentation scopes a file to the host, protocol and port it is served from. A subdomain or CDN hostname has its own file, and your root file does not cover it.
  • →Blocked resources are the top rendering problem: That was Rebecca Yu's message in Barcelona, as reported by ROAST. If the JavaScript is disallowed, Google cannot build a page that looks fine to users.
  • →robots.txt and robots meta do different jobs: One controls crawling per host. The other controls indexing and snippets per URL, and Google only sees it on pages it is allowed to crawl.

Why Google checks a separate robots.txt for every host

Google published no official recap of the Barcelona event, so Dave Smart's point reaches us through attendees. It does not need to rest on them. Google's own page on how it interprets the robots.txt specification says the same thing in three rules, which I re-read on 8 October 2026:

  • Host, protocol and port. The rules in a robots.txt file apply only to the host, protocol and port number where the file is hosted.
  • Subdomains are separate. A robots.txt on a subdomain is only valid for that subdomain. The file on www.example.com is not valid for example.com.
  • Protocol matters too. The file at https://example.com/robots.txt is not valid for http://example.com/.

Apply that to a normal page. The HTML comes from your main host, the bundles from an asset subdomain, a library from a public CDN and a widget from a vendor. That is one page, four hosts and four robots.txt files, and you can edit two of them at most.

Per-host scope

One page, four hosts, four robots.txt files

The page Google renders
https://www.example.com/pricing
HostWhat it servesFile Google checksWho can edit it
www.example.comThe HTML documentwww.example.com/robots.txtYou
static.example.comYour CSS and JavaScript bundlesstatic.example.com/robots.txtYou, but it is a separate file
cdn.example.netA framework or library from a public CDNcdn.example.net/robots.txtThe CDN provider
api.example.orgA reviews or pricing widget and its dataapi.example.org/robots.txtThe widget vendor
Illustrative hosts on reserved example domains. Google's specification scopes each robots.txt to the host, protocol and port it is served from, so a clean file on the first row says nothing about the other three.
Host (illustrative)What it servesrobots.txt Google checksWho can edit it
www.example.comThe HTML documentwww.example.com/robots.txtYou
static.example.comYour CSS and JavaScript bundlesstatic.example.com/robots.txtYou, but it is a separate file
cdn.example.netA framework or library from a public CDNcdn.example.net/robots.txtThe CDN provider
api.example.orgA reviews or pricing widget and its dataapi.example.org/robots.txtThe widget vendor

What Google's renderer does with the files it can fetch

Two more reported points from Day 2 explain why a blocked file matters so much. Both come from John Campbell's Day 2 recap for ROAST:

  • Rebecca Yu on the top rendering problem. It is robots.txt blocking the resources the JavaScript needs to run. If the JS files are disallowed, Google cannot build the page even though it looks fine in a browser.
  • Erin Sparling on how rendering works. Google loads pages in a viewport about 10,000 pixels tall. It does not scroll and it does not click, so JavaScript waiting for those events never runs and scroll-triggered content is not seen.

Google's introduction to robots.txt draws the same line from the other side: blocking unimportant image, script or style files is fine, but not when losing them makes the page harder for the crawler to understand.

robots.txt vs robots meta tag vs X-Robots-Tag

Three controls share similar names. "robots meta tag" gets 1,000 US searches a month against 6,600 for "robots.txt" (DataForSEO, US, October 2026), and they are not interchangeable. This table is built from Google's robots.txt pages and its robots meta tag specifications.

Questionrobots.txtrobots meta tagX-Robots-Tag header
What it controlsCrawling: whether a URL may be fetchedIndexing and how the result is shownIndexing and how the result is shown
ScopeOne host, protocol and portOne HTML pageOne URL of any file type
Where it lives/robots.txt at the root of the hostA meta tag in the page's headAn HTTP response header
When Google reads itBefore fetching anything on that hostOnly after crawling the pageOnly after crawling the URL
Removes a page from search?No. A blocked URL can still be indexed from linksYes, with noindexYes, with noindex
Use it forKeeping crawlers out of URL spaces that should never be fetchednoindex, max-snippet, max-image-preview on pagesThe same rules on PDFs, images and other non-HTML files

The classic trap sits in the fourth row. Google's specifications page says robots meta tags and X-Robots-Tag headers are discovered when a URL is crawled. Block a page in robots.txt and its noindex is never found, so the URL can stay in the index.

Crawling vs indexing: in plain English

robots.txt is a sign on the door that says whether a crawler may come in. The robots meta tag is a note inside the room that says what may be shown in search. If the sign keeps the crawler out, it never reads the note.

How to audit robots.txt across every host in six steps

Work through the steps in order, one template at a time. Start with the templates that carry revenue or leads: a blocked bundle on a product template affects every product page at once.

The audit

Six steps: map, compare, fix, verify

Map
1List every hostDevTools Network tab, Domain column, per template
2Fetch each robots.txtOne file per host; match the resource paths
Compare
3See what Google rendersURL Inspection live test: screenshot, HTML, More info
Fix
4Fix or work aroundEdit your own hosts; self-host or switch for the rest
5Set robots metamax-image-preview:large, max-snippet:-1
Verify
6Re-check after a dayGoogle caches robots.txt for up to 24 hours
Run it once per template, not once per URL. A product template, an article template and the homepage usually cover every host a site loads from.

Step 1: List every host your key templates load from

  1. Pick one live URL for each template: homepage, a category or listing page, a product or article page.
  2. Open it in Chrome, open DevTools, go to the Network tab and reload with the cache disabled.
  3. Right-click the column headers, turn on the Domain column, and sort by it.
  4. Write down every domain that serves a stylesheet, a script, a font or a fetch/XHR response the page uses to draw its main content.

View source is a quicker first pass, but it misses anything JavaScript requests after load. That is often the product data or reviews you most need Google to see.

Step 2: Fetch each host's robots.txt and test the resource URLs

  1. For every host on your list, load https://that-host/robots.txt in a browser. Use the same protocol the resource uses.
  2. Paste each host into the free robots.txt tester, which lays out every Allow and Disallow rule per user agent.
  3. Find the group Googlebot follows: its own group if the file has one, otherwise User-agent: *. Google does not combine the two.
  4. Match the exact resource path against the rules. Google uses the most specific rule by path length, and the least restrictive one on a conflict.

The pattern to look for is a host that was blocked wholesale to keep its files out of search results, like this:

# https://static.example.com/robots.txt
User-agent: *
Disallow: /

That one line makes every bundle on the host unavailable to the renderer. Narrower versions do the same damage: a Disallow on /assets/, /js/ or /api/ when your content depends on files under that path.

Step 3: Compare what Google renders with what users see

  1. In Search Console, paste the URL into URL Inspection and click Test live URL.
  2. Click View tested page and open the Screenshot tab. Compare it with your browser.
  3. Open the HTML tab and search for a sentence from the main content, a product price, or a review.
  4. Open More info to see the page resources and JavaScript console output from the test, and note any resource that did not load.

Google's URL Inspection help page notes that differences in the screenshot can come from resources blocked to its inspection crawler. Match each failed resource against your host list from step 1. A sentence missing from the rendered HTML is the finding that matters.

Step 4: Fix the hosts you control and work around the rest

Sort the blocked resources by who owns the host and how much the page depends on them.

  • Your own subdomains. Edit that subdomain's robots.txt. Remove the blanket Disallow or add Allow lines for the asset paths. The robots.txt generator writes a clean file if the host has none.
  • Your CDN configuration. If the asset hostname is yours but the CDN serves it, find out which file answers at /robots.txt on that hostname and whether your CDN lets you set it.
  • A third party, critical resource. Self-host the file on a host you control, inline the critical CSS, or render the content on the server so it does not depend on the blocked request.
  • A third party, no workaround. Ask the vendor, or switch to a provider whose robots.txt allows the paths you load.
  • Leave alone. A blocked analytics tag, ad script or chat widget changes nothing about the content Google indexes. Spend the effort on resources that produce main content, layout or links.

Step 5: Set the robots meta directives for previews and snippets

Once Google can render the page, tell it how much it may show. John Mueller walked through every robots meta directive on Day 2 and, according to John Campbell's recap for ROAST, ended by recommending two for maximum visibility:

<meta name="robots" content="max-image-preview:large, max-snippet:-1">
  • max-image-preview:large allows a larger image preview, up to the width of the viewport. Without the rule, Google's specification says it may show a preview of the default size.
  • max-snippet:-1 lets Google choose the snippet length it believes is most effective. The same page says a max-snippet limit also caps how much content can be used as direct input for AI Overviews and AI Mode.
  • Non-HTML files take the same rules through an X-Robots-Tag header. Confirm what a URL returns with the free HTTP header checker.
Opinion· Aditi's take
Reading Google's specification next to the recap, the image directive is the one that changes behaviour, because the default preview is the smaller size. Per the same page, Google already chooses snippet length when max-snippet is absent, so -1 mostly makes your intent explicit. I would still ship both. Setting the pair on purpose forces you to look at what your theme or plugins already output, and a low max-snippet left behind by one of them is easy to miss.

Step 6: Re-check after about a day

  1. Reload each changed robots.txt in a browser and confirm the new rules are live.
  2. Wait about a day. Google's documentation says it generally caches robots.txt for up to 24 hours.
  3. Repeat the live test from step 3 and check that the screenshot, the HTML and the resource list now match your browser.

The Barcelona timing table agrees with the documentation. Gary Illyes's Day 3 slides, as reported in ROAST's Day 3 recap and reproduced by Search Engine Roundtable, put a robots.txt update at about 24 hours typical and 25 hours at the slowest. The full table is in how long Google takes to index a page.

What Google does when a host's robots.txt returns an error

A host does not need a Disallow line to block you. The status code of its /robots.txt response matters as much as the rules inside it. Google's specification page sets out four cases:

  • 2xx. Google processes the file as served.
  • 3xx. Google follows at least five redirect hops, then treats the file as a 404.
  • 4xx, except 429. Treated as if no valid robots.txt exists, which means no crawl restrictions. A resource host with no robots.txt is fine.
  • 5xx. For the first 12 hours Google stops crawling the site and keeps retrying the file. For up to 30 days after that it uses the last good copy.
Status codes

What a host's /robots.txt response means for the files it serves

2xx
Success
What Google does
Google processes the file as served.
For your page resources
The rules decide. Match your resource paths against them.
3xx
Redirect
What Google does
Google follows at least five hops, then treats the file as a 404.
For your page resources
Read the file at the end of the chain, not the first URL.
4xx
Client error (except 429)
What Google does
Treated as if no valid robots.txt exists: no crawl restrictions.
For your page resources
A host with no robots.txt blocks nothing. This one is fine.
5xx
Server error
What Google does
Crawling stops for the first 12 hours. Then the last good copy is used for up to 30 days.
For your page resources
Treat it as a block: Google stops crawling that host while the error is fresh.
Source: Google Search Central, How Google interprets the robots.txt specification, read 8 October 2026. Google generally caches a robots.txt file for up to 24 hours.

That 5xx behaviour is what John Mueller reached for in a Day 1 whiteboard exercise. The scenario: 500,000 product pages with updated metadata, a 30-day CDN edge cache, and Googlebot crawling stale copies. As reported by ROAST, Mueller suggested serving a 503 on robots.txt to pause crawling. Gary Illyes's answer was to do nothing.

The same mechanism can hurt you by accident. The Day 1 crawl errors session reportedly flagged CDNs blocking more bot traffic, with the blocks showing up as HTTP errors. A bot rule that answers Googlebot's robots.txt request with a server error is a block. Our guide to the 429 status code and SEO covers the rate-limit side.

Tips that save you a second audit

  • Audit templates, not URLs. Three or four templates usually cover every host a site loads from.
  • Test the exact file path. A host can allow /v2/ and disallow /v1/. Checking the homepage of the CDN tells you nothing.
  • Look for scroll and click triggers while you are in the rendered HTML. Content that waits for a scroll event is missing for a different reason, and no robots.txt edit will bring it back.
  • Repeat after a vendor change. A new tag manager container, review widget or font host adds a robots.txt you have not read.
  • Do not assume AI crawlers behave the same way. Google's documentation describes Google's crawlers. I have not verified how each AI crawler treats cross-host resources, so treat that as untested.

Why the page still renders wrong, and what to check next

The live test reports a blocked resource, but my robots.txt allows it. Why? Because your file is not the one being checked. Read the hostname in the resource URL and open /robots.txt on that host. Check the protocol and the www variant too, since each combination has its own file.

Every host allows the files, but content is still missing from the rendered HTML. What now? Rule out robots.txt and look at how the content loads:

  • It waits for a scroll or click. As reported from Barcelona, Google's renderer does neither. Load the content on page load or render it on the server.
  • The host returns an error to bots. Check the resource's status code in the live test. A firewall rule can serve crawlers a different response from the one your browser gets.
  • The page is still in the queue. The reported timing table lists rendering as seconds to render and hours in the queue, with a slowest case of days to weeks.

I unblocked the file yesterday and Search Console still says blocked. Is the fix wrong? Probably not. Google generally caches robots.txt for up to 24 hours. Confirm the live file, wait the full day, then run the live test again before changing anything else.

The page renders correctly now but sits in "Crawled – currently not indexed". Is that the same problem? Not necessarily. Rendering is one cause worth ruling out, and this audit rules it out. The rest are covered in Crawled vs Discovered – currently not indexed.

Tools that make this audit faster

  • Chrome DevTools for the host list. Free, and already on your machine.
  • The robots.txt tester for reading each host's rules per user agent. "robots.txt tester" gets 1,900 US searches a month (DataForSEO, US, October 2026). Point it at every host on your list, not only your own domain.
  • URL Inspection for Google's own view of the render. It is the only check here that uses Google's fetcher.
  • The HTTP header checker for X-Robots-Tag and for the status code a robots.txt URL really returns.

What none of those tell you is which template to audit first. CrawlRaven joins Search Console, GA4 and a 200-point crawl into one ranked plan, so the pages Google reports as "Blocked by robots.txt" sit next to their clicks and the crawl's view of each URL. We built it to answer the order question.

Lifetime licences launched at $49 for 3 sites. Lifetime pricing steps up as licenses sell, so check the pricing page for the current batch. See pricing.

What you have at the end of the audit

You finish with a list of every host your templates depend on, a verdict for each host's robots.txt, and a live test showing Google's render matches the browser. Keep the host list. It is the part that goes stale when a vendor or CDN changes.

On Shopify, the first-party file has its own rules, covered in our Shopify robots.txt guide. Everywhere else, add the six steps to the checklist for every new third-party script.

Sources used in this post

Frequently asked questions

Does my robots.txt apply to files loaded from a CDN or another domain?

No. Google's robots.txt documentation says the rules in a file apply only to the host, protocol and port number where that file is hosted, and that a file on a subdomain is valid only for that subdomain. A script served from a CDN hostname is checked against the robots.txt on that CDN hostname. Your own file has no say over it.

Does blocking JavaScript and CSS in robots.txt affect SEO?

Yes, when the page depends on those files. Google's introduction to robots.txt says not to block resources if their absence makes the page harder for its crawler to understand. At Search Central Live in Barcelona, Rebecca Yu named robots.txt blocking the resources JavaScript needs as the top rendering problem, according to John Campbell's recap for ROAST.

How do I test whether a specific script URL is blocked by robots.txt?

Fetch /robots.txt on the host that serves the script, find the group that applies to Googlebot (its own group if one exists, otherwise User-agent: *), and match the script's path against the Disallow and Allow lines. The longest matching rule wins. Then confirm with a URL Inspection live test of a page that loads the script.

How long does Google take to pick up a robots.txt change?

About a day. Google's documentation says it generally caches robots.txt for up to 24 hours. The timing table Gary Illyes showed in Barcelona on 2 October 2026 put a robots.txt update at about 24 hours typical and 25 hours at the slowest, as reported by ROAST and Search Engine Roundtable. Re-test the day after you change a file.

What happens if a robots.txt file returns a 404 or a 503?

A 404 means no restrictions and a 503 means a pause. Google's documentation says 4xx responses other than 429 are treated as if no valid robots.txt exists. For 5xx responses Google stops crawling the site for the first 12 hours while retrying the file, then uses the last good copy for up to 30 days.

What is the difference between robots.txt and the robots meta tag?

robots.txt controls crawling for a whole host, and the robots meta tag controls indexing and snippets for one page. The meta tag sits in the page's HTML, so Google has to crawl the page to read it. X-Robots-Tag does the same job as the meta tag from an HTTP header, which makes it the option for PDFs and images.

What do max-image-preview:large and max-snippet:-1 do?

They tell Google it may show a large image preview and choose the text snippet length itself. Google's robots meta documentation defines large as a preview up to the width of the viewport and -1 as Google choosing the length it believes is most effective. John Mueller recommended the pair in Barcelona for maximum visibility, according to John Campbell's recap for ROAST.

Why can't Google see the noindex tag on my page?

Usually because robots.txt blocks the page. Google's documentation says robots meta tags and X-Robots-Tag headers are discovered when a URL is crawled, so on a disallowed URL the rule is never found and is ignored. Remove the Disallow, let Google crawl the page and read the noindex, and add the block back only after the URL has dropped out.

Does Googlebot scroll or click when it renders a page?

No, as reported from Search Central Live in Barcelona. According to John Campbell's recap for ROAST, Erin Sparling said Google loads pages in a viewport about 10,000 pixels tall and does not scroll or click. JavaScript that waits for a scroll or a click does not run, so content that only loads on those events is not seen.

Aditi Chaturvedi
About the Author

Aditi Chaturvedi

15+ years of growing SaaS websites through SEO | Author, 200-Point Audit Checklist

Aditi has spent 15+ years helping SaaS companies scale organic traffic through technical SEO and content strategy. She is the author of the CrawlRaven 200-Point Audit checklist used by agencies and in-house teams to systematically improve search performance.

robots.txtrobots meta tagrobots.txt testerrobots.txt blocking javascript and cssmax-image-preview largemax-snippetjavascript seox-robots-taggoogle renderingsearch central live

Reader has pages that look fine in a browser but render thin or broken for Google, and wants to know which templates are affected and what the affected pages are worth.

Find the templates worth auditing first

Search Console and GA4 data read against a full crawl of the site

CrawlRaven joins Search Console, GA4 and a 200-point crawl into one ranked plan. It lists the URLs Google reports as blocked by robots.txt next to the crawl's view of each one, ordered by the traffic at stake. Free plan for 1 site, no credit card required.

CrawlRaven connects Google Search Console and GA4, runs 200+ technical SEO checks, and joins all three into one prioritized fix list, so you know what is broken, what it is costing you, and what to fix first.

✓ No credit card required·200+ checks·GSC + GA4 + full-site crawl
Free plan — no credit card

Stop exporting. Start shipping.

Connect Search Console, import your Ahrefs or Semrush lists, and get one ranked plan. Start free with one site, or grab a limited lifetime deal from $39, only 3 licenses left.

3
Data sources joined
200+
Point audit checks
1
Ranked plan out