Robots.txt vs Noindex: Keep the Right SaaS Pages Out
Robots.txt vs noindex for SaaS: which to use on preview, account, utility, and filter pages, why they conflict, and why private data needs a login instead.
TL;DR Robots.txt controls crawling. Noindex controls indexing. If you want a page out of Google, let Google crawl it and serve noindex, because a URL blocked in robots.txt can never have its noindex read and can still show up in results if something links to it. And neither one protects private data. That's what authentication is for.
Here's a pattern I keep running into on small SaaS sites. A founder ships a shareable report feature. Every report gets a public preview URL so users can send it to their boss. A few weeks later they search their brand name and find one of those previews in Google, with a customer's company name in the title. They panic, add Disallow: /r/ to robots.txt and redeploy.
Next week the URL is still there. Now the snippet just reads "No information is available for this page."
They made it worse. Google can no longer crawl the page, so it can never see the noindex they're about to add, and the URL is still known because somebody posted the link in a public Slack archive. The real fix was never going to be robots.txt anyway. If a preview shows customer data, it shouldn't be reachable by anyone without a link token or a login.
I see some version of this on nearly every small SaaS site I look at. So here's how I'd think about robots.txt vs noindex, page type by page type.
Robots txt vs noindex: crawling is not indexing
These two tools answer different questions.
Robots.txt answers "may a crawler request this URL?" Google's robots.txt introduction says it is "used primarily to manage crawler traffic to your site." Then, in plain words: "it is not a mechanism for keeping a web page out of Google."
Noindex answers "may this page appear in search results?" Google's robots meta tag reference defines it as "Do not show this page, media, or resource in search results." You set it in the HTML head:
<meta name="robots" content="noindex">
Or as an HTTP response header, which also works for PDFs and other non-HTML files:
X-Robots-Tag: noindex
The trap is that the second tool depends on the first. Google has to fetch a page to read its noindex. From the block indexing docs: "If the page is blocked by a robots.txt file or the crawler can't access the page, the crawler will never see the noindex rule, and the page can still appear in search results, for example if other pages link to it."
The robots meta tag page says it even more bluntly: if a page is disallowed in robots.txt, "any information about indexing or serving rules will not be found and will therefore be ignored."
So the robots txt vs noindex question isn't "which one is stronger." Stack them on the same URL and the disallow cancels the noindex.
What a disallowed URL looks like in Google
Disallowing a URL doesn't erase it. Google's robots.txt intro explains that if a blocked page is linked from elsewhere, the URL "and, potentially, other publicly available information such as anchor text in links to the page can still appear in Google Search results." The result just won't have a description.
That's the "No information is available for this page" snippet Google describes in its help page on missing page information. Google knows the URL exists, it knows what people call it in their anchor text, and it has no idea what's on it. For a founder trying to hide a page, that's the worst of both worlds.
Noindex in robots.txt doesn't work
The other common mistake is putting a Noindex: line inside robots.txt itself. Some old tutorials still recommend it. Google never documented the rule, and whatever unofficial support existed is gone. In July 2019 Google announced it was "retiring all code that handles unsupported and unpublished rules (such as noindex) on September 1, 2019." The same post found these rules were "contradicted by other rules in all but 0.001% of all robots.txt files on the internet."
If you search "robots txt noindex" and land on a guide that shows this:
User-agent: *
Noindex: /preview/
close the tab. Google ignores that line. The same goes for a Nofollow: line. If you want the robots txt noindex nofollow combination, it lives in the meta tag or header, not in robots.txt:
<meta name="robots" content="noindex, nofollow">
Google's 2019 post lists the supported alternatives: noindex in meta tags or headers, 404 and 410 status codes, password protection, disallow in robots.txt (with the linked-URL caveat), and the Search Console removals tool for temporary removal.
Robots.txt is not access control
This deserves its own section because it's the one with real consequences.
Robots.txt is a public file. Anyone can read yours at /robots.txt. A line like Disallow: /admin-exports/ is a signpost telling every curious visitor where the interesting stuff is. Well-behaved crawlers follow it voluntarily. Nothing forces anyone else to.
Google says this directly in the robots.txt intro: "if you want to keep information secure from web crawlers, it's better to use other blocking methods, such as password-protecting private files on your server." And the 2019 announcement notes that "hiding a page behind a login will generally remove it from Google's index."
Noindex isn't access control either. A noindexed page still loads for anyone who has the URL. It just won't appear in search.
My rule: if the page would be a problem in a screenshot on Twitter, SEO directives are the wrong layer. Put it behind authentication, or at minimum behind an unguessable signed token that expires. Then handle crawling and indexing as a separate, much smaller question.
The decision table for SaaS page types
Here's how I'd handle the page types that show up on almost every SaaS site. "Crawl" means allowed in robots.txt. "Index" means no noindex.
| Page type | Example | Crawl | Index | Why |
|---|---|---|---|---|
| Public utility page | Free calculator, generator, checker | Yes | Yes | These can rank. Treat them as content, not plumbing |
| Thin utility result | Output URL for one user's calculator input | Yes | Noindex | Infinite variations, no search demand for each one |
| Shared preview, no private data | Public demo report, template preview | Yes | Noindex, unless it's a real landing page | Lets Google see the noindex and drop it |
| Shared preview with customer data | Report or dashboard share link | No access without token or login | Not applicable | This is an auth problem, not an SEO one |
| Account and app pages | Dashboard, settings, billing | Behind login | Not applicable | Login already keeps them out. A disallow is optional, to save crawl requests |
| Login and signup | Sign-in, register, password reset | Yes | Usually noindex on reset and verification flows | People do search "yourbrand login", so the main login page can stay indexable |
| Internal site search results | Search-query URLs on your blog or docs | Disallow | Not applicable | Unbounded URL space, no unique value |
| Filter and sort combinations | Integration directory filtered by category and sort | Disallow the parameters | Not applicable | Google's guidance for faceted URLs |
| Staging or preview deploys | Branch preview URLs | Password protect | Not applicable | Never rely on either directive here |
A few rows need more explanation.
Search and filter pages: this is where disallow wins
For faceted and filtered URLs, Google's faceted navigation guidance says to "use robots.txt to disallow crawling of faceted navigation URLs," because "there's no good reason to allow crawling of filtered items, as it consumes server resources for no or negligible benefit."
This is the case where robots txt disallow vs noindex tips toward disallow. Noindex requires Google to crawl every combination first to discover it shouldn't index them. With five filters and a sort order, that's a lot of fetches for a site Google already doesn't crawl much.
The one exception: if a filtered view has real search demand, like an integration category people actually search for, give it a clean static URL, write something useful on it, and let it index. That's a real page, not a filter state. My post on programmatic SEO for SaaS and when it turns into spam covers where that line is.
Pages already in the index: noindex first, disallow later (maybe)
If a page is already indexed and you want it gone, the order matters:
- Make sure robots.txt allows the URL.
- Add noindex via meta tag or
X-Robots-Tag. - Wait for Google to recrawl it and drop it. The URL Inspection tool in Search Console shows when it was last crawled.
- Only then, if you still want to save crawl requests, add a disallow.
Most founders do step 4 first. That's how the report preview in my opening story got stuck.
If the page is gone for good, skip all of this and return a 404 or 410. Google's 2019 post says both codes "will drop such URLs from Google's index once they're crawled and processed." That's also why a soft 404 is worth fixing: a page that says "not found" while the server returns 200 sends Google mixed signals.
Duplicates are a different tool
If two URLs show the same content and you want one of them to rank, you don't need robots.txt or noindex. That's canonicalization, and it has its own rules. I covered it in duplicate without user-selected canonical. Don't stack noindex on top of a canonical tag for the same URL: one says drop this page, the other says treat it as a copy of another page. Pick one job per URL.
Why this matters more on a small site
A big site can leak a few thousand junk URLs into the index and nobody notices. On a site with 60 real pages, 400 indexed filter combinations and preview links change what Google thinks the site is.
The effect shows up in Search Console. Low-value URLs pile up in the Page indexing report, and your real posts sit in Discovered, currently not indexed or Crawled, currently not indexed while Google works through the rest. I can't give you a clean number for how much that hurts, because Google doesn't publish one. But I've never seen a small site where cleaning up indexable junk made things worse.
This is also why I think about indexing hygiene as part of the content plan, not a separate chore. The pages you want indexed and the pages you publish should be the same list. When I built the calendar logic for Boomranq, the whole premise was that a low-authority site gets a small number of shots. Wasting some of Google's attention on filter URLs and orphaned previews is the same mistake as writing posts for keywords you can't win.
A 20-minute audit
Here's what I'd do this week:
- Read your own robots.txt. Look for any
Noindex:lines (delete them, they do nothing) and any disallow that covers a page you're also trying to noindex. - Run a site search in Google for your domain and scroll. Look for previews, share links, filter URLs, and staging hosts.
- Check the Page indexing report in Search Console for "Indexed, though blocked by robots.txt." That status is the exact conflict described above: Google indexed a URL it isn't allowed to read.
- Search your robots.txt for anything sensitive. If a path in there would embarrass you, it needs a login, not a disallow.
- Pick one rule per page type from the table and apply it at the template level, not URL by URL.
If you're earlier than this and still setting up the basics, this fits into the SaaS SEO checklist for your first six months. And if pages you do want are missing from Google, start with the diagnostic for a website not showing up on Google before you touch any directive.
The short version of robots.txt vs noindex: robots.txt manages crawler traffic, noindex manages search results, and a login manages who sees what. Use each for its own job and they stop fighting each other.