Canonical Tags, Redirect Loops and Indexing: The Pre-Launch Technical SEO Checklist
Launch-day SEO mistakes are uniquely expensive because the symptom arrives weeks later, long after anyone connects it to the deploy. Here is the pre-flight list.

The worst technical SEO problems are not the ones that break the site. They are the ones that work perfectly for humans and quietly tell Google the wrong thing. Nobody notices at launch. Six weeks later traffic is down 40% and the deploy is no longer an obvious suspect.
This is the pre-flight list for a new site, a redesign, or a domain migration — ordered by how expensive the mistake is to discover late.
1. Staging is not indexable, production is
The classic pair of mistakes, and they are mirror images: staging gets indexed because nobody blocked it, or production ships with the noindex that was protecting staging.
<!-- On staging. Must NOT be on production. --> <meta name="robots" content="noindex, nofollow">
Protect staging with HTTP authentication rather than noindex. A password cannot be forgotten in a merge, and it also stops the staging URL leaking into a client email chain. Then, immediately after the production deploy, fetch the live HTML and confirm the tag is gone:
curl -s https://example.com | grep -i 'name="robots"' curl -sI https://example.com | grep -i 'x-robots-tag'
Check the header as well as the meta tag. An X-Robots-Tag: noindex set at the server or CDN is invisible in the page source and blocks indexing just as effectively.
2. Canonicals point at the live, final URL
A canonical tag tells Google which URL is authoritative. Get it wrong and you can deindex a site while every page loads perfectly.
The failures, in order of severity:
| Mistake | Effect |
|---|---|
| Canonical points at the staging domain | Google indexes nothing on production |
| Every page canonicalises to the homepage | Entire site collapses to one indexed URL |
Canonical uses http:// or the wrong host | Signals split between duplicates |
| Canonical points at a redirecting URL | Signal diluted, canonical often ignored |
| Missing on parameterised URLs | Filter and sort URLs compete with the real page |
Relative canonical on a page with a <base> | Resolves somewhere unintended |
Generate it from the request, not from a config string someone has to remember to change. In Next.js:
// app/products/[slug]/page.jsx
export async function generateMetadata({ params }) {
const base = process.env.NEXT_PUBLIC_SITE_URL // https://www.example.com
return {
alternates: { canonical: `${base}/products/${params.slug}` },
}
}And in WordPress, where the usual cause is two SEO plugins each emitting a canonical:
# Should return exactly one line curl -s https://example.com/product/x | grep -c 'rel="canonical"'
3. One redirect, never a chain, never a loop
Pick a canonical host — apex or www, HTTPS — and enforce it once, at the edge. Problems arise when redirect rules accumulate in three places: an .htaccess rule, a CDN page rule, and application middleware. Each was correct alone.
# Every hop, with status codes curl -sIL http://example.com/old-page | grep -iE "^(HTTP/|location:)" # Expect at most: # HTTP/1.1 301 # location: https://www.example.com/new-page # HTTP/2 200
On a migration, map old URLs to new ones directly. A chain where /old → /interim → /new works for visitors but loses signal at every hop, and a loop between apex and www takes the site down entirely for crawlers while often looking fine in a browser with a cached redirect.
Use 301 for permanent moves and 302 only when the move genuinely is temporary. A 302 left on a permanent migration keeps ranking signals attached to the old URL indefinitely.
4. The page renders without JavaScript
This is the failure mode that has grown fastest, and it is invisible in every browser test. A React or Vue app that ships an empty <div id="root"> and builds the page client-side serves identical, contentless HTML on every URL.
Google will usually render it eventually, on a second pass, with a delay. But social crawlers — Facebook, LinkedIn, WhatsApp, Slack — and most AI answer engines do not execute JavaScript at all. They see nothing.
# What a non-JS crawler actually receives curl -s https://example.com/pricing | grep -o '<title>[^<]*' curl -s https://example.com/pricing | wc -c # If every route returns the same bytes and the same title, # your content is invisible to anything that does not run JS.
The fix is server rendering or build-time prerendering for anything that should rank or be shared. Marketing pages should ship as HTML; the application behind a login can stay client-rendered, since it should not be indexed anyway.
5. Sitemap and robots.txt agree with reality
- Every URL in the sitemap returns 200 — no redirects, no 404s, no
noindexpages. lastmodreflects real changes. Stamping every URL with the build date teaches Google to ignore the field.- The sitemap is referenced in
robots.txtand submitted in Search Console. - Generate the sitemap from your route list at build time, so a new page cannot be live and absent from it.
robots.txtdoes not block CSS or JS — Google needs them to render the page.- Pages disallowed in
robots.txtare not also in the sitemap. That contradiction is a common Search Console warning.
noindex. To remove a page from the index, allow the crawl and serve noindex. Block it in robots.txt only after it has dropped out.6. Structured data and the basics
- Exactly one
<h1>per page, describing that page. - Unique
<title>and meta description per URL — not templated to identical text. - Organization and WebSite JSON-LD on the homepage; Product, Article or LocalBusiness where relevant.
- JSON-LD is in the served HTML, not injected by a script after load.
hreflangonly if you genuinely have localised versions, and reciprocal on both sides.- An
og:imageat 1200×630 and under about 300KB, or link previews stay blank on WhatsApp.
The SEO Fundamentals module checks canonicals, robots directives, redirect chains, heading structure, structured data and indexability — before you point the domain at it.
Audit my siteThe launch-day sequence
- Deploy. Immediately curl the homepage and one deep page; check for
noindexin the HTML and the headers. - Verify canonicals resolve to the live host and to themselves.
- Trace the redirect from
http://, from the apex, and from three old URLs. - Fetch a page with JavaScript disabled and confirm the content is there.
- Submit the sitemap in Search Console and request indexing for the top ten URLs.
- Re-run the whole check 24 hours later, after CDN caches have turned over.
Step 6 matters more than it looks. A CDN serving a cached copy of the old robots.txt or an old HTML shell will pass every test at deploy time and fail quietly for a day afterwards, which is exactly long enough for a crawl to pick up the wrong version.
- Why Did Google Ads Disapprove My Landing Page? Fixing "Destination Not Working" and "Circumventing Systems"
- Server Response Time (TTFB), DNS Propagation and SSL Handshakes: The Infrastructure QA Guide
- How to Fix Content Security Policy (CSP) Headers and Mixed Content Errors That Break Third-Party Scripts
