SEO

X-Robots-Tag

Also called X-Robots-Tag header

An HTTP response header carrying robots rules, so PDFs, images and downloads can be kept out of search.

Quick facts: X-Robots-Tag

Category
SEO
Also called
X-Robots-Tag header
Level
Advanced
Affects
Indexing of non-HTML files, index bloat, snippet control
Where to see it
Apache or Nginx config, browser network panel, Screaming Frog, URL Inspection tool
In this article4
  1. How the X-Robots-Tag works
  2. Why the X-Robots-Tag matters
  3. Where the X-Robots-Tag goes wrong
  4. How to use it properly

How the X-Robots-Tag works

When a browser or a crawler asks for a URL, the server sends a set of headers back before it sends the file itself. The X-Robots-Tag is one of those headers, and it carries the same instructions a robots meta tag carries: noindex, nofollow, noarchive, nosnippet, limits on how much of a page may be shown as a snippet or preview, and a date after which the listing should expire.

Because it travels with the response rather than inside the markup, it works on anything the server can send. A PDF price list, a spreadsheet, an exported report, a raw image file: none of them can hold a robots meta tag, and all of them can carry a header. The tag can also name a particular crawler, so a file can be kept out of one engine’s index while remaining available to others.

It is set in server configuration rather than in a page. On Apache that usually means a rule in the site config or an .htaccess file, on Nginx a header directive, and on managed platforms an option in the control panel or an SEO plugin. Rules are normally matched by file extension or by URL pattern, which is what makes the header efficient: one rule can cover a whole class of files at once.

Why the X-Robots-Tag matters

Almost every site that has been running for a few years has accumulated files nobody meant to publish to search. Old proposals, event brochures, invoices, exported spreadsheets, draft price lists. They surface for odd queries, they date badly, and occasionally they expose something the business assumed was private simply because it was never linked. A single header rule on the uploads folder deals with all of them without editing a single file.

It is also the only clean way to control how non-HTML content appears, and the only way to apply noindex to something you cannot edit at all, such as a document generated by a booking system or an accounting tool.

Where the X-Robots-Tag goes wrong

The classic failure is combining it with a robots.txt block. If robots.txt disallows the URL, the crawler never requests it, never sees the header, and never learns the file should be dropped. The URL can sit in results with no description under it. To remove something, allow the crawl and send the noindex.

The second failure is invisibility. A header does not appear in the page source, so a quick look at the markup finds nothing and people conclude the rule was never applied, or assume it is working when a caching layer or CDN has quietly stripped it. Only the live response tells you the truth.

The third is scope. A pattern written for one folder can catch far more than intended, and a staging rule that noindexes everything becomes dangerous the moment that configuration is copied to the production server.

How to use it properly

Decide first whether you want the file out of the index or merely unlinked, because the header is for the former. Write the narrowest rule that does the job, then test one live URL and read the response headers in your browser’s network panel or with a crawler that reports them. Recheck after any hosting move, CDN change or platform upgrade, since those are the moments headers disappear without anyone noticing. If the answer keeps coming back that dozens of file types need different treatment, the upload structure itself needs sorting out, and that is proper technical SEO work rather than a quick config edit.

Do and do not

Do

  • Use it for PDFs, images and other non-HTML files
  • Confirm the header on a live response, not source
  • Keep the URL crawlable so the header can be read

Do not

  • Block a URL in robots.txt and expect noindex to work
  • Apply a site-wide rule without testing one URL
  • Leave a staging noindex header on the live server

Questions people ask about this

What is the difference between the X-Robots-Tag and the robots meta tag?

They carry the same instructions and search engines treat them the same way. The difference is where they live. The meta tag sits in the HTML head, so it only works on HTML pages. The X-Robots-Tag travels in the HTTP response header, so it works on any file the server sends, including PDFs, images and downloads.

Can I use the X-Robots-Tag to remove a PDF from Google?

Yes, as long as Google can still fetch the file. Add a noindex header for that URL and leave it crawlable so Googlebot can read the instruction. Removal is not instant; the file drops out when it is next crawled. For something urgent, request removal in Search Console too and keep the header in place afterwards.

Do I need developer access to set it?

Usually yes. The header is set in server configuration or by the application itself, so you need access to Apache, Nginx or the hosting control panel, or a plugin that writes the rule for you. On some managed platforms the option is missing entirely, in which case the robots meta tag on HTML pages is the fallback.

Related terms

Found this useful?

Share it, or ask an AI to summarise it

Back to the glossary

Knowing the term is the easy part

Applying it to your own site and budget is the work. Book a call and I will tell you what actually applies to you.