How to Check the robots.txt Files You Don't Control
Scripts and CSS served from a CDN or subdomain follow that host's robots.txt, not yours. A six-step audit to find what Google cannot load and fix the render.
Every file a page loads is checked against the robots.txt of the host that serves it. To find out whether someone else's file is breaking how Google renders your pages:
- List the hosts: Open one URL per template in DevTools, Network tab, and note every domain that serves CSS, JavaScript or data.
- Read each host's robots.txt: Fetch /robots.txt on every host and match the exact resource paths against its rules. A 404 means nothing is blocked.
- Compare the render: Run a URL Inspection live test and compare Google's screenshot and rendered HTML with your browser.
- Fix or work around: Edit the hosts you own. For the rest, self-host or inline the resource, or change provider.
- Set robots meta, then wait a day: Add max-image-preview:large and max-snippet:-1, then re-test after about 24 hours, which is how long Google caches robots.txt.
"robots.txt" gets 6,600 US searches a month (DataForSEO, US, October 2026). Almost none of the answers mention that the file in question may belong to someone else.
Your robots.txt can be spotless while Google still fails to build the page, because the scripts and styles come from hosts with robots.txt files of their own. This guide walks the six-step audit across every host, using Google's documentation and the attendee recaps from Search Central Live in Barcelona. Try CrawlRaven free: 1 site, no credit card →
The short answer: the robots.txt that blocks your render may not be yours
Search for "robots.txt blocking javascript and css" and the top results are about ten years old. They all tell you to check your own file. Reading the recaps from Google's Search Central Live Deep Dive in Barcelona (30 September to 2 October 2026), what stood out to me was a lightning talk that points somewhere else.
Dave Smart of Tame the Bots said that when a page loads files from another domain, such as a CDN or a subdomain, "it's that domain's robots.txt that applies, not yours", in the words of John Campbell's Day 1 recap for ROAST. Your file can pass every check while one you have never opened breaks the render.
So the how-to is this: list every host your templates load from, read each host's robots.txt, and compare what Google renders with what a visitor sees. Set aside an hour or two the first time. You need DevTools and Search Console access.
Key Takeaways
- →robots.txt is per host: Google's documentation scopes a file to the host, protocol and port it is served from. A subdomain or CDN hostname has its own file, and your root file does not cover it.
- →Blocked resources are the top rendering problem: That was Rebecca Yu's message in Barcelona, as reported by ROAST. If the JavaScript is disallowed, Google cannot build a page that looks fine to users.
- →robots.txt and robots meta do different jobs: One controls crawling per host. The other controls indexing and snippets per URL, and Google only sees it on pages it is allowed to crawl.
Why Google checks a separate robots.txt for every host
Google published no official recap of the Barcelona event, so Dave Smart's point reaches us through attendees. It does not need to rest on them. Google's own page on how it interprets the robots.txt specification says the same thing in three rules, which I re-read on 8 October 2026:
- Host, protocol and port. The rules in a robots.txt file apply only to the host, protocol and port number where the file is hosted.
- Subdomains are separate. A robots.txt on a subdomain is only valid for that subdomain. The file on www.example.com is not valid for example.com.
- Protocol matters too. The file at https://example.com/robots.txt is not valid for http://example.com/.
Apply that to a normal page. The HTML comes from your main host, the bundles from an asset subdomain, a library from a public CDN and a widget from a vendor. That is one page, four hosts and four robots.txt files, and you can edit two of them at most.
One page, four hosts, four robots.txt files
| Host (illustrative) | What it serves | robots.txt Google checks | Who can edit it |
|---|---|---|---|
| www.example.com | The HTML document | www.example.com/robots.txt | You |
| static.example.com | Your CSS and JavaScript bundles | static.example.com/robots.txt | You, but it is a separate file |
| cdn.example.net | A framework or library from a public CDN | cdn.example.net/robots.txt | The CDN provider |
| api.example.org | A reviews or pricing widget and its data | api.example.org/robots.txt | The widget vendor |
What Google's renderer does with the files it can fetch
Two more reported points from Day 2 explain why a blocked file matters so much. Both come from John Campbell's Day 2 recap for ROAST:
- Rebecca Yu on the top rendering problem. It is robots.txt blocking the resources the JavaScript needs to run. If the JS files are disallowed, Google cannot build the page even though it looks fine in a browser.
- Erin Sparling on how rendering works. Google loads pages in a viewport about 10,000 pixels tall. It does not scroll and it does not click, so JavaScript waiting for those events never runs and scroll-triggered content is not seen.
Google's introduction to robots.txt draws the same line from the other side: blocking unimportant image, script or style files is fine, but not when losing them makes the page harder for the crawler to understand.
robots.txt vs robots meta tag vs X-Robots-Tag
Three controls share similar names. "robots meta tag" gets 1,000 US searches a month against 6,600 for "robots.txt" (DataForSEO, US, October 2026), and they are not interchangeable. This table is built from Google's robots.txt pages and its robots meta tag specifications.
| Question | robots.txt | robots meta tag | X-Robots-Tag header |
|---|---|---|---|
| What it controls | Crawling: whether a URL may be fetched | Indexing and how the result is shown | Indexing and how the result is shown |
| Scope | One host, protocol and port | One HTML page | One URL of any file type |
| Where it lives | /robots.txt at the root of the host | A meta tag in the page's head | An HTTP response header |
| When Google reads it | Before fetching anything on that host | Only after crawling the page | Only after crawling the URL |
| Removes a page from search? | No. A blocked URL can still be indexed from links | Yes, with noindex | Yes, with noindex |
| Use it for | Keeping crawlers out of URL spaces that should never be fetched | noindex, max-snippet, max-image-preview on pages | The same rules on PDFs, images and other non-HTML files |
The classic trap sits in the fourth row. Google's specifications page says robots meta tags and X-Robots-Tag headers are discovered when a URL is crawled. Block a page in robots.txt and its noindex is never found, so the URL can stay in the index.
robots.txt is a sign on the door that says whether a crawler may come in. The robots meta tag is a note inside the room that says what may be shown in search. If the sign keeps the crawler out, it never reads the note.
How to audit robots.txt across every host in six steps
Work through the steps in order, one template at a time. Start with the templates that carry revenue or leads: a blocked bundle on a product template affects every product page at once.
Six steps: map, compare, fix, verify
Step 1: List every host your key templates load from
- Pick one live URL for each template: homepage, a category or listing page, a product or article page.
- Open it in Chrome, open DevTools, go to the Network tab and reload with the cache disabled.
- Right-click the column headers, turn on the Domain column, and sort by it.
- Write down every domain that serves a stylesheet, a script, a font or a fetch/XHR response the page uses to draw its main content.
View source is a quicker first pass, but it misses anything JavaScript requests after load. That is often the product data or reviews you most need Google to see.
Step 2: Fetch each host's robots.txt and test the resource URLs
- For every host on your list, load https://that-host/robots.txt in a browser. Use the same protocol the resource uses.
- Paste each host into the free robots.txt tester, which lays out every Allow and Disallow rule per user agent.
- Find the group Googlebot follows: its own group if the file has one, otherwise User-agent: *. Google does not combine the two.
- Match the exact resource path against the rules. Google uses the most specific rule by path length, and the least restrictive one on a conflict.
The pattern to look for is a host that was blocked wholesale to keep its files out of search results, like this:
# https://static.example.com/robots.txt
User-agent: *
Disallow: /That one line makes every bundle on the host unavailable to the renderer. Narrower versions do the same damage: a Disallow on /assets/, /js/ or /api/ when your content depends on files under that path.
Step 3: Compare what Google renders with what users see
- In Search Console, paste the URL into URL Inspection and click Test live URL.
- Click View tested page and open the Screenshot tab. Compare it with your browser.
- Open the HTML tab and search for a sentence from the main content, a product price, or a review.
- Open More info to see the page resources and JavaScript console output from the test, and note any resource that did not load.
Google's URL Inspection help page notes that differences in the screenshot can come from resources blocked to its inspection crawler. Match each failed resource against your host list from step 1. A sentence missing from the rendered HTML is the finding that matters.
Step 4: Fix the hosts you control and work around the rest
Sort the blocked resources by who owns the host and how much the page depends on them.
- Your own subdomains. Edit that subdomain's robots.txt. Remove the blanket Disallow or add Allow lines for the asset paths. The robots.txt generator writes a clean file if the host has none.
- Your CDN configuration. If the asset hostname is yours but the CDN serves it, find out which file answers at /robots.txt on that hostname and whether your CDN lets you set it.
- A third party, critical resource. Self-host the file on a host you control, inline the critical CSS, or render the content on the server so it does not depend on the blocked request.
- A third party, no workaround. Ask the vendor, or switch to a provider whose robots.txt allows the paths you load.
- Leave alone. A blocked analytics tag, ad script or chat widget changes nothing about the content Google indexes. Spend the effort on resources that produce main content, layout or links.
Step 5: Set the robots meta directives for previews and snippets
Once Google can render the page, tell it how much it may show. John Mueller walked through every robots meta directive on Day 2 and, according to John Campbell's recap for ROAST, ended by recommending two for maximum visibility:
<meta name="robots" content="max-image-preview:large, max-snippet:-1">- max-image-preview:large allows a larger image preview, up to the width of the viewport. Without the rule, Google's specification says it may show a preview of the default size.
- max-snippet:-1 lets Google choose the snippet length it believes is most effective. The same page says a max-snippet limit also caps how much content can be used as direct input for AI Overviews and AI Mode.
- Non-HTML files take the same rules through an X-Robots-Tag header. Confirm what a URL returns with the free HTTP header checker.
Step 6: Re-check after about a day
- Reload each changed robots.txt in a browser and confirm the new rules are live.
- Wait about a day. Google's documentation says it generally caches robots.txt for up to 24 hours.
- Repeat the live test from step 3 and check that the screenshot, the HTML and the resource list now match your browser.
The Barcelona timing table agrees with the documentation. Gary Illyes's Day 3 slides, as reported in ROAST's Day 3 recap and reproduced by Search Engine Roundtable, put a robots.txt update at about 24 hours typical and 25 hours at the slowest. The full table is in how long Google takes to index a page.
What Google does when a host's robots.txt returns an error
A host does not need a Disallow line to block you. The status code of its /robots.txt response matters as much as the rules inside it. Google's specification page sets out four cases:
- 2xx. Google processes the file as served.
- 3xx. Google follows at least five redirect hops, then treats the file as a 404.
- 4xx, except 429. Treated as if no valid robots.txt exists, which means no crawl restrictions. A resource host with no robots.txt is fine.
- 5xx. For the first 12 hours Google stops crawling the site and keeps retrying the file. For up to 30 days after that it uses the last good copy.
What a host's /robots.txt response means for the files it serves
That 5xx behaviour is what John Mueller reached for in a Day 1 whiteboard exercise. The scenario: 500,000 product pages with updated metadata, a 30-day CDN edge cache, and Googlebot crawling stale copies. As reported by ROAST, Mueller suggested serving a 503 on robots.txt to pause crawling. Gary Illyes's answer was to do nothing.
The same mechanism can hurt you by accident. The Day 1 crawl errors session reportedly flagged CDNs blocking more bot traffic, with the blocks showing up as HTTP errors. A bot rule that answers Googlebot's robots.txt request with a server error is a block. Our guide to the 429 status code and SEO covers the rate-limit side.
Tips that save you a second audit
- Audit templates, not URLs. Three or four templates usually cover every host a site loads from.
- Test the exact file path. A host can allow /v2/ and disallow /v1/. Checking the homepage of the CDN tells you nothing.
- Look for scroll and click triggers while you are in the rendered HTML. Content that waits for a scroll event is missing for a different reason, and no robots.txt edit will bring it back.
- Repeat after a vendor change. A new tag manager container, review widget or font host adds a robots.txt you have not read.
- Do not assume AI crawlers behave the same way. Google's documentation describes Google's crawlers. I have not verified how each AI crawler treats cross-host resources, so treat that as untested.
Why the page still renders wrong, and what to check next
The live test reports a blocked resource, but my robots.txt allows it. Why? Because your file is not the one being checked. Read the hostname in the resource URL and open /robots.txt on that host. Check the protocol and the www variant too, since each combination has its own file.
Every host allows the files, but content is still missing from the rendered HTML. What now? Rule out robots.txt and look at how the content loads:
- It waits for a scroll or click. As reported from Barcelona, Google's renderer does neither. Load the content on page load or render it on the server.
- The host returns an error to bots. Check the resource's status code in the live test. A firewall rule can serve crawlers a different response from the one your browser gets.
- The page is still in the queue. The reported timing table lists rendering as seconds to render and hours in the queue, with a slowest case of days to weeks.
I unblocked the file yesterday and Search Console still says blocked. Is the fix wrong? Probably not. Google generally caches robots.txt for up to 24 hours. Confirm the live file, wait the full day, then run the live test again before changing anything else.
The page renders correctly now but sits in "Crawled – currently not indexed". Is that the same problem? Not necessarily. Rendering is one cause worth ruling out, and this audit rules it out. The rest are covered in Crawled vs Discovered – currently not indexed.
Tools that make this audit faster
- Chrome DevTools for the host list. Free, and already on your machine.
- The robots.txt tester for reading each host's rules per user agent. "robots.txt tester" gets 1,900 US searches a month (DataForSEO, US, October 2026). Point it at every host on your list, not only your own domain.
- URL Inspection for Google's own view of the render. It is the only check here that uses Google's fetcher.
- The HTTP header checker for X-Robots-Tag and for the status code a robots.txt URL really returns.
What none of those tell you is which template to audit first. CrawlRaven joins Search Console, GA4 and a 200-point crawl into one ranked plan, so the pages Google reports as "Blocked by robots.txt" sit next to their clicks and the crawl's view of each URL. We built it to answer the order question.
Lifetime licences launched at $49 for 3 sites. Lifetime pricing steps up as licenses sell, so check the pricing page for the current batch. See pricing.
What you have at the end of the audit
You finish with a list of every host your templates depend on, a verdict for each host's robots.txt, and a live test showing Google's render matches the browser. Keep the host list. It is the part that goes stale when a vendor or CDN changes.
On Shopify, the first-party file has its own rules, covered in our Shopify robots.txt guide. Everywhere else, add the six steps to the checklist for every new third-party script.
Sources used in this post
- Google Search Central: How Google interprets the robots.txt specification (scope, status codes, caching), read 8 October 2026
- Google Search Central: Robots meta tag, data-nosnippet and X-Robots-Tag specifications
- Google Search Central: Introduction to robots.txt
- Search Console Help: URL Inspection tool
- John Campbell for ROAST: Day 1 recap (30 September 2026), Day 2 recap (1 October 2026) and Day 3 recap (2 October 2026)
- Search Engine Roundtable: Google's crawling, indexing and serving timing data (5 October 2026)
- Search volumes: DataForSEO, US, October 2026. Google published no official recap of the Barcelona event; every session detail above is attendee-reported.
Related reading on CrawlRaven
- How long Google takes to index a page: the full timing tables from Barcelona.
- The 429 status code and SEO: what rate limiting does to crawl rate.
- Crawled vs Discovered – currently not indexed: what each status means and what to fix.
- Shopify robots.txt: the default file and when to customise it.
- robots.txt, noindex and URL Inspection in the glossary.
Frequently asked questions
Does my robots.txt apply to files loaded from a CDN or another domain?
No. Google's robots.txt documentation says the rules in a file apply only to the host, protocol and port number where that file is hosted, and that a file on a subdomain is valid only for that subdomain. A script served from a CDN hostname is checked against the robots.txt on that CDN hostname. Your own file has no say over it.
Does blocking JavaScript and CSS in robots.txt affect SEO?
Yes, when the page depends on those files. Google's introduction to robots.txt says not to block resources if their absence makes the page harder for its crawler to understand. At Search Central Live in Barcelona, Rebecca Yu named robots.txt blocking the resources JavaScript needs as the top rendering problem, according to John Campbell's recap for ROAST.
How do I test whether a specific script URL is blocked by robots.txt?
Fetch /robots.txt on the host that serves the script, find the group that applies to Googlebot (its own group if one exists, otherwise User-agent: *), and match the script's path against the Disallow and Allow lines. The longest matching rule wins. Then confirm with a URL Inspection live test of a page that loads the script.
How long does Google take to pick up a robots.txt change?
About a day. Google's documentation says it generally caches robots.txt for up to 24 hours. The timing table Gary Illyes showed in Barcelona on 2 October 2026 put a robots.txt update at about 24 hours typical and 25 hours at the slowest, as reported by ROAST and Search Engine Roundtable. Re-test the day after you change a file.
What happens if a robots.txt file returns a 404 or a 503?
A 404 means no restrictions and a 503 means a pause. Google's documentation says 4xx responses other than 429 are treated as if no valid robots.txt exists. For 5xx responses Google stops crawling the site for the first 12 hours while retrying the file, then uses the last good copy for up to 30 days.
What is the difference between robots.txt and the robots meta tag?
robots.txt controls crawling for a whole host, and the robots meta tag controls indexing and snippets for one page. The meta tag sits in the page's HTML, so Google has to crawl the page to read it. X-Robots-Tag does the same job as the meta tag from an HTTP header, which makes it the option for PDFs and images.
What do max-image-preview:large and max-snippet:-1 do?
They tell Google it may show a large image preview and choose the text snippet length itself. Google's robots meta documentation defines large as a preview up to the width of the viewport and -1 as Google choosing the length it believes is most effective. John Mueller recommended the pair in Barcelona for maximum visibility, according to John Campbell's recap for ROAST.
Why can't Google see the noindex tag on my page?
Usually because robots.txt blocks the page. Google's documentation says robots meta tags and X-Robots-Tag headers are discovered when a URL is crawled, so on a disallowed URL the rule is never found and is ignored. Remove the Disallow, let Google crawl the page and read the noindex, and add the block back only after the URL has dropped out.
Does Googlebot scroll or click when it renders a page?
No, as reported from Search Central Live in Barcelona. According to John Campbell's recap for ROAST, Erin Sparling said Google loads pages in a viewport about 10,000 pixels tall and does not scroll or click. JavaScript that waits for a scroll or a click does not run, so content that only loads on those events is not seen.
15+ years of growing SaaS websites through SEO | Author, 200-Point Audit Checklist
Aditi has spent 15+ years helping SaaS companies scale organic traffic through technical SEO and content strategy. She is the author of the CrawlRaven 200-Point Audit checklist used by agencies and in-house teams to systematically improve search performance.