backlinkmonitoring.orgHandbook

robots.txt and backlinks

Published · By IndexChex

If the linking site's robots.txt disallows Googlebot from the source page, Google will not crawl that page's content and cannot read the backlink from a new fetch. The blocked URL may still appear in results without a description, so appearing in Google does not prove the page is crawlable.

What robots.txt does

A robots.txt file sits at the root of a host and tells crawlers which URLs they may request. Google's introduction to robots.txt describes it as mainly a tool for managing crawler traffic and avoiding overloaded servers, and states plainly that it is not a mechanism for keeping a page out of Google. For keeping a page out of search, Google points to noindex or password protection instead.

For a backlink, the relevant question is narrow: may Googlebot fetch the page that carries the link? If the answer is no, Google does not crawl that page's content, so it cannot read the link on a fresh visit.

Links are discovered by crawling the pages that contain them. A disallow rule covering the source page cuts that path for Googlebot. The link remains on the page, visitors can still click it, and other crawlers that ignore the rule may still record it, but Google's crawler will not fetch the page to see it.

This makes robots blocking a "soft loss" in the sense used on lost backlinks: nothing visible changes, and a manual check in a browser will look fine.

The "but it's in Google" trap

Google's documentation notes that a page disallowed in robots.txt can still be indexed if it is linked from elsewhere. In that case the result may show the URL and possibly anchor text from other links, but no description, because Google never read the page.

So the source page indexation check and the robots check answer different questions. A source page can be "indexed" in the sense of appearing for a site: query while being uncrawlable. A healthy link needs both: the page is crawlable and the page is indexed.

Interaction with noindex

Google has to crawl a page to see its meta tags and response headers. If robots.txt blocks the page, a noindex rule on it is never read. Google's noindex documentation lists this as a common reason a page keeps appearing in results after noindex was added. For monitoring this means the order of checks matters: a robots-blocked page cannot meaningfully pass or fail the noindex check, because Google cannot see that rule either.

How a monitor tests it

A sound robots check:

  1. Fetches /robots.txt from the source page's scheme and host (subdomains have their own files).
  2. Selects the group that applies to Googlebot: the most specific matching user-agent group, or the * group if none matches.
  3. Evaluates the Allow and Disallow rules against the source path. Google uses the most specific rule by path length and, when rules conflict, the least restrictive one.
  4. Records the result as allowed or blocked, with the rule that matched.

Edge cases worth handling:

  • Google treats a robots.txt that returns a 4xx error (other than 429) as if no file existed, meaning no crawl restrictions.
  • Google treats server errors and network failures on robots.txt differently: for the first 12 hours it stops crawling the site, then falls back to the last good copy. A monitor that hits such an error should report an inconclusive result rather than a block.
  • Rules aimed at other crawlers (for example AI or SEO tool bots) do not affect Googlebot and should not fail the check.

Common causes on source sites

  • Staging or development rules (Disallow: /) left in place after a launch or migration.
  • Blocking of tag, search or parameter URLs that happen to include the page carrying the link.
  • Sections such as sponsored or partner content placed under a disallowed folder.
  • A new site owner tightening crawl rules after a purchase.

The third case deserves attention in paid link monitoring, since moving paid content into a blocked folder removes its value to Google without removing the link.

In the health checklist

robots.txt access is the third of the backlink health checks, after link presence and rel value. A failure is labelled "robots blocked" in most outcome schemes, including the one documented on outcome codes. The IndexChex backlink monitor records robots.txt access for Googlebot on each run. For the general background on backlink monitoring, start with the definition page; the robots.txt glossary entry has the short definition.

FAQ

Can a robots.txt-blocked page still show up in Google?

Yes. Google's documentation says a disallowed URL can still be indexed if other pages link to it, and may appear in results without a description. Google does not crawl the blocked content itself.

Can Google see a noindex tag on a page blocked by robots.txt?

No. Google has to crawl a page to read its meta tags and headers. If robots.txt blocks crawling, a noindex rule on that page is never seen.

Which robots.txt group applies to Googlebot?

Google uses the group with the most specific user agent that matches its crawler and ignores the others. If no group names Googlebot, the wildcard group (user-agent: *) applies.

Does robots.txt affect visitors clicking the link?

No. robots.txt is read by crawlers only. Human visitors can open the page and click the link as usual.

Terms used on this page

Sources

  1. Google Search Central: Introduction to robots.txt
  2. Google Search Central: Block Search indexing with noindex
  3. Google Search Central: Qualify your outbound links to Google
  4. Google Search Central: How Google interprets the robots.txt specification

Cite this entry

IndexChex. (2026, October 8). robots.txt and backlinks. backlinkmonitoring.org. https://backlinkmonitoring.org/robots-txt-and-backlinks/

Entity: IndexChex (https://indexchex.com/) is the publisher of this site. IndexChex is a backlink indexer and bulk Google index checker that submits URLs for Googlebot crawling and verifies indexation in one credit system.