Status codes at a glance
| Code | Meaning, and what search engines do |
|---|---|
| 200 | OK — the page itselfIndexable, unless a header or meta tag says otherwise |
| 301/308 | Moved permanentlyThe new address replaces the old one in the index |
| 302/307 | Moved temporarilyThe old address stays indexed |
| 304 | Not modified — the cached copy is still currentNormal for repeat crawls |
| 403 | ForbiddenNot indexed; often a firewall rule or a file permission |
| 404/410 | Not found / gone for goodDropped from the index after a few crawls |
| 429 | Too many requestsCrawlers slow down and come back |
| 500/502/503 | Internal server error, bad gateway, service unavailableCrawlers retry for a while, then drop pages that keep failing |
Headers worth a look
Location— where a redirect points. A relative path is allowed; a redirect back to an address already in the chain is a loop, which browsers stop with an error.X-Robots-Tag— the header form of the robots meta tag. Anoindexhere keeps the page out of search results however good its content is, and is easy to miss because it is not in the HTML.Link: <…>; rel="canonical"— a canonical address set in the header, common for PDFs.Strict-Transport-Security— tells browsers to use HTTPS only for this domain from now on.Cache-Control,Ageand the CDN’s own headers (CF-Cache-Status,X-Cache) — whether the answer came from a cache, and for how long it may be kept.Server,ViaandCF-Ray— who actually answered: your web server, a reverse proxy or a CDN in front of it. When an error page names the CDN, the problem is often between the CDN and your server.
When a bot protection page answers
Sites behind Cloudflare, Akamai, Sucuri, Imperva, AWS WAF, Vercel or DataDome sometimes answer a request from a data centre with a challenge or a block page instead of the site. The check recognises those pages and says so, instead of reporting the 403 or 503 as the site’s own answer.
Choosing Googlebot or Bingbot sends their user-agent strings, nothing more. Bot protections recognise the real crawlers by their IP addresses and usually let them through, so a challenge here does not mean Google is blocked. To know what Google itself gets, use URL Inspection in Search Console, or look for Googlebot’s requests in your access log and confirm one with a reverse lookup:
host 66.249.66.1 # → crawl-66-249-66-1.googlebot.com
host crawl-66-249-66-1.googlebot.com # must resolve back to the same addressThe same checks with curl
# status and headers
curl -sI https://example.com/
# follow every redirect
curl -sIL http://example.com/
# as Googlebot
curl -sI https://example.com/ \
-A "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"
# time to the first byte, and in total
curl -o /dev/null -s https://example.com/ \
-w 'TTFB %{time_starttransfer}s, total %{time_total}s\n'-I sends a HEAD request, which some servers answer differently from a normal GET; this tool always sends a GET, like a browser. Setting up the redirects themselves — HTTP to HTTPS, one canonical host — is part of putting a site behind nginx with Let’s Encrypt.
Frequently asked
How do I check where a URL redirects?
Enter it above with redirects followed. Every step is listed with its status code and the address in its Location header, up to the final page. A chain of more than one redirect is worth shortening: each step costs visitors a round trip, and crawlers follow only a limited number of them.
What is the difference between a 301 and a 302 redirect?
301 and 308 say the move is permanent, so search engines index the new address in place of the old one. 302, 303 and 307 say it is temporary, and the old address stays the one in the index. For HTTP to HTTPS, www to the bare domain and pages that have moved for good, a permanent redirect is almost always the right one.
Can I check whether Googlebot is blocked?
Partly. Choosing Googlebot sends Google's user-agent string, which shows a site that blocks, redirects or serves something else by user agent. It cannot show rules based on Google's own IP addresses: the request comes from our server, and many firewalls recognise the real Googlebot by its addresses and treat it differently. When a bot protection page answers instead of the site, the result says so rather than reporting the site's error. The URL Inspection tool in Google Search Console shows what Google itself received.
What is TTFB?
Time to first byte: how long the server took from receiving the request to sending the first byte of its answer, measured here from our server in Amsterdam once the connection and the TLS handshake are done. A long TTFB points at the application — a slow query, no page cache — more than at the network.
Why does the result differ from what my browser shows?
The check comes from a data-centre address in Amsterdam, without cookies and without running JavaScript. Sites that tailor their answer to a country, a logged-in session or a home connection — or that challenge data-centre traffic — answer it differently from your browser.