Skip to content

SEO for PWAs

Search engine optimization for a Progressive Web App is ordinary technical SEO with a few extra traps. The manifest, the service worker and installability are invisible to search engines. The architecture that PWAs tend to use (client-side rendering, an app shell, client-side routing and an offline fallback) is fully visible, and it's where rankings get lost. This page explains how Googlebot and other crawlers fetch and render JavaScript, why they never run your service worker, and how to design URLs, rendering, metadata, structured data, sitemaps and international variants so every route of your PWA can be indexed. It ends with a repeatable testing workflow.

Key takeaways

  • Search engines index the HTML they get from your server, plus, for Google, the DOM after one rendering pass. Your service worker, Cache Storage, IndexedDB and install state play no part: crawlers are stateless first-time visitors.
  • Server-side rendering or static generation is the only rendering strategy that works for every crawler. Google renders JavaScript, but social link-preview bots and most AI crawlers don't. Google calls dynamic rendering a workaround, not a long-term solution.
  • Give every view a real URL with the History API and a real <a href>, return real HTTP status codes (404 for unknown routes), and never put content behind fragments (#/route).
  • The manifest has no documented ranking effect. Core Web Vitals are used by Google's ranking systems, but relevance always comes first.
  • Strip launch-tracking parameters such as ?source=pwa with a canonical URL, keep the offline fallback page out of the index, and keep content-hashed assets available after a deploy because Google's renderer caches aggressively.
  • Test with the URL Inspection tool's live test (rendered HTML, screenshot, console, resources), not with your own browser, which has the service worker and your cookies.

What changes when a website becomes a PWA

The components that make something a PWA barely affect search engines. The architectural choices that often come with it matter a great deal. The table sorts the common PWA ingredients by their SEO impact.

PWA ingredient Seen by crawlers? SEO impact
Web app manifest Fetched at most as a plain file None documented. name, description and icons aren't used for search snippets.
Service worker No: not installed or run during rendering None directly. Indirectly through real-user performance (Core Web Vitals).
Installability, install prompts No None, unless a promotion covers content like an intrusive interstitial.
Client-side rendering (CSR) Google: after a deferred render. Most other bots: no High. Empty HTML means nothing to index for non-rendering crawlers and a delay for Google.
App shell served for every URL Yes: the shell HTML is what the server returns High. Identical <title>, missing content, soft 404s.
Client-side routing Only if links are <a href> and routes are real URLs High. Fragment routes and onclick navigation hide pages.
Offline fallback page Only if your server ever returns it Medium. A server that answers errors with the offline page creates soft 404s.
start_url tracking parameters If the URL is linked or shared Low to medium. Duplicate URLs, unless canonicalized.
Precaching and runtime caching No None for crawlers. Stale HTML for users can carry stale metadata.

In short: everything a crawler knows about your PWA comes from HTTP responses your server sends to a first-time, stateless visitor, plus for Google, one rendering pass over that response.

How Googlebot crawls, renders and indexes JavaScript

The three phases

Google processes a JavaScript page in three phases: crawling, rendering and indexing. Its JavaScript SEO basics documentation describes the pipeline:

flowchart LR
    A["Crawl queue"] --> B["Fetch URL<br/>(robots.txt check)"]
    B --> C["Parse raw HTML<br/>for links"]
    C --> A
    B --> D{"HTTP 200 and<br/>no noindex?"}
    D -- yes --> E["Render queue"]
    D -- "no" --> G["Rendering may be skipped"]
    E --> F["Headless Chromium<br/>(Web Rendering Service)"]
    F --> H["Parse rendered HTML<br/>for links"]
    H --> A
    F --> I["Index rendered HTML"]

The details that matter for a PWA:

  • Every 200 response is queued for rendering, whether or not it contains JavaScript, unless a robots meta tag or header says not to index it. Google says a page "may stay on this queue for a few seconds, but it can take longer than that".
  • Non-200 responses might not be rendered. Google's documentation says that "if the HTTP status code is non-200 (for example, on error pages with 404 status code), rendering might be skipped". A client-rendered error page that relies on JavaScript to explain itself shows nothing to Google.
  • noindex can prevent rendering. When Google sees noindex in the initial HTML, it may skip rendering and JavaScript execution. Removing noindex with JavaScript therefore may not work. Setting it with JavaScript, the reverse, does work, and Google recommends it for client-side error views.
  • Links are extracted twice: from the raw HTML and again from the rendered HTML. Server-rendered links are discovered earlier.

The Web Rendering Service is evergreen but stateless

Since May 2019, Googlebot renders pages with an evergreen Chromium that Google updates regularly (announcement, Chromium 74 at the time). Modern JavaScript, IntersectionObserver and Web Components work, so polyfills and transpilation that exist only for Googlebot are unnecessary. What differs from a user's browser is state and interaction. Google's JavaScript troubleshooting guide lists the constraints:

Constraint What Google documents Consequence for a PWA
No persistent state Local Storage, Session Storage and HTTP cookies are cleared across page loads Content gated on a cookie, a stored locale or IndexedDB data is invisible.
Permissions declined "Expect Googlebot to decline user permission requests" Never gate content on geolocation, notifications or camera permission.
HTTP only No WebSockets or WebRTC connections Content that arrives over a WebSocket is missing. Provide an HTTP fallback.
Feature detection Use feature detection with fallbacks for all critical APIs; Google's example is that feature detection shows Googlebot doesn't support WebGL A rendering path that throws on a missing API renders a blank page.
Aggressive caching "WRS may ignore caching headers" Googlebot can render with old JavaScript or CSS. Use content-hashed file names.
No interaction "Google Search does not interact with your page" (lazy-loading guide) Content loaded on scroll, click or hover is invisible. Load on viewport visibility instead.
File size cutoff Googlebot fetches the first 2 MB of a supported file, measured uncompressed; each CSS and JavaScript resource is fetched separately with the same limit (Googlebot docs) A monolithic bundle over 2 MB uncompressed is truncated, and a truncated script usually fails to parse.

The 2 MB limit is a recent tightening worth auditing. A single-page app's main bundle can exceed 2 MB before compression even when it is 400 KB over the wire. If the truncated file is the one that renders your content, Google indexes an empty shell. Code-splitting per route (which you want for loading performance anyway) keeps each file far below the limit.

Googlebot and service workers

Google's crawling and rendering documentation doesn't mention service workers at all, but the mechanism makes the outcome predictable. Googlebot fetches each URL itself over HTTP, and the Web Rendering Service clears Local Storage, Session Storage and cookies across page loads, so every render is a first visit. On a first visit, the page's own navigation request is never handled by a service worker: a worker can only control later navigations (or the current page after clients.claim(), by which time the HTML has already arrived from the network). Even if your registration code runs during rendering, the HTML Google indexes came from your server. Treat that as a design rule: the service worker doesn't exist for crawlers. The consequences:

  • navigator.serviceWorker.register() may exist in the renderer, but your page must never wait for the registration, for controller or for a message from the worker before rendering content.
  • Anything your service worker synthesizes, such as HTML assembled from cached partials, streamed composite responses or content from IndexedDB, must also be producible by the server. If only the worker can build the page, only returning users can see it.
  • Offline fallbacks, precaching and runtime caching don't affect indexing. They affect users, and therefore field metrics.
  • Your first-visit experience is what gets indexed. Test with the service worker bypassed (DevTools Application panel → Service workers → Bypass for network, or a fresh profile).

The URL Inspection live test is the authoritative check: if the rendered HTML contains your content with no service worker involved, you're fine.

Google is the most capable JavaScript renderer among crawlers. Plan for the least capable one that matters to you.

  • Bing. Bing's crawler can render JavaScript but has described limits on doing so at scale. Its 2018 bingbot post on JavaScript and dynamic rendering recommended serving prerendered HTML to bingbot for JavaScript-heavy sites. Server-rendered HTML removes the question entirely.
  • AI crawlers. Vercel's December 2024 analysis of crawler traffic, The rise of the AI crawler, found that none of the major AI crawlers it analyzed rendered JavaScript: OpenAI's, Anthropic's, Meta's (Meta-ExternalAgent), ByteDance's (Bytespider) and Perplexity's (PerplexityBot). OpenAI's and Anthropic's crawlers did download JavaScript files (11.50% and 23.84% of their requests) but didn't execute them. Only Google's Gemini (through Googlebot's infrastructure) and Applebot rendered JavaScript, and Common Crawl's CCBot, a common source of LLM training data, didn't. A client-rendered PWA is effectively empty to most AI assistants.
  • Link-preview bots. Messaging apps and social networks build previews from Open Graph and Twitter Card meta tags in the raw HTML. They generally don't run JavaScript, so tags inserted by a client-side head manager are ignored.
  • IndexNow. Bing, Yandex and several other engines accept push notifications of changed URLs through the IndexNow protocol: a POST to an IndexNow endpoint with up to 10,000 URLs and a key that you prove by hosting a key file. Google doesn't participate. It's covered in sitemaps and discovery.

Choosing a rendering strategy

The rendering strategy decides what a stateless crawler receives. It's also the biggest lever for Largest Contentful Paint on first visits, which are the visits search engines send you.

Strategy Raw HTML contains content Google Non-rendering crawlers First-visit LCP Fit for a PWA
Static site generation (SSG) Yes ✅ ✅ Best Content that changes on deploy: docs, marketing, catalogs
Server-side rendering (SSR) Yes ✅ ✅ Good (depends on TTFB) Dynamic, personalized or frequently changing content
SSR or SSG + hydration / islands Yes ✅ ✅ Good Most PWAs: indexable HTML plus an interactive app
Incremental or on-demand static regeneration Yes ✅ ✅ Best to good Large catalogs
Client-side rendering (CSR) with an app shell No ⚠️ after deferred rendering ❌ Worst Only for routes that shouldn't be indexed (behind login, noindex)
Dynamic rendering (prerender for bots only) For bots only ⚠️ workaround ✅ Unchanged for users Legacy; plan a migration away

⚠️ Google renders CSR pages, but later than it crawls them, and any rendering failure (a JavaScript error, a timeout, a truncated bundle, an API that blocks bots) produces an empty index entry.

The app shell trap

The app shell model serves a minimal HTML skeleton and fills it with JavaScript. It gives excellent repeat-visit performance from the cache, but a naive deployment answers every URL with the same shell:

What the server returns for /products/42 in a naive app-shell PWA
<!doctype html>
<html lang="en">
  <head>
    <title>Acme Store</title>
    <link rel="manifest" href="/manifest.webmanifest">
    <script type="module" src="/assets/app.3f9c1e.js"></script>
  </head>
  <body>
    <div id="root"></div>
  </body>
</html>

Every product page has the same title, no description, no canonical, no structured data and no content until JavaScript runs, and /products/does-not-exist returns the same 200. The fix isn't to drop the app shell but to separate two concerns:

  1. What the server returns for a navigation: fully rendered HTML for that URL, with its own <title>, meta, canonical, structured data and status code.
  2. What the service worker serves on later navigations: whatever is fastest for returning users. That can be the cached shell plus client rendering, or streamed HTML assembled from cached header and footer partials plus network content.

Because crawlers never have the worker, (1) is what gets indexed and (2) only affects returning users. Both need to produce the same visible content.

Keeping server HTML authoritative in the service worker

The worker should never serve an old server-rendered page in place of the current one for a long time, or users see stale titles and content that differ from what's indexed. A network-first navigation strategy with navigation preload keeps HTML fresh and still works offline:

sw.js
const HTML_CACHE = "html-v1";
const OFFLINE_URL = "/offline.html";

self.addEventListener("install", (event) => {
  event.waitUntil(
    caches.open(HTML_CACHE).then((cache) => cache.add(OFFLINE_URL))
  );
});

self.addEventListener("activate", (event) => {
  event.waitUntil(
    (async () => {
      // Navigation preload starts the network request while the worker boots.
      if (self.registration.navigationPreload) {
        await self.registration.navigationPreload.enable();
      }
      await self.clients.claim();
    })()
  );
});

self.addEventListener("fetch", (event) => {
  if (event.request.mode !== "navigate") return;

  event.respondWith(
    (async () => {
      const cache = await caches.open(HTML_CACHE);
      try {
        // Prefer the preloaded response; fall back to a normal fetch.
        const response = (await event.preloadResponse) || (await fetch(event.request));
        // Only cache complete, successful HTML. Never cache 404s or redirects
        // under the requested URL, or the offline copy would disagree with
        // what search engines see.
        if (response.ok && response.type === "basic") {
          event.waitUntil(cache.put(event.request, response.clone()));
        }
        return response;
      } catch {
        // Offline: the last good copy of this exact URL, else the offline page.
        return (
          (await cache.match(event.request, { ignoreSearch: false })) ||
          (await cache.match(OFFLINE_URL)) ||
          Response.error()
        );
      }
    })()
  );
});

The strategy trade-offs are covered in Caching Strategies. For SEO, the only requirement is that the user-facing HTML and the crawler-facing HTML describe the same page.

Dynamic rendering is deprecated

Dynamic rendering serves prerendered HTML to bots (detected by User-Agent) and the client-rendered app to users. Google's dynamic rendering page now opens with the statement that it "was a workaround and not a long-term solution for problems with JavaScript-generated content in search engines", and recommends server-side rendering, static rendering or hydration instead. It isn't cloaking as long as both versions show the same content, but it doubles your rendering paths, needs a headless browser farm, and fails silently when the prerender cache goes stale. If you run it today, treat it as a migration step toward SSR.

URLs, routing and the History API

Search engines index URLs. A view that has no URL of its own, or that can only be reached by a script event, doesn't exist for them. Google's rules:

  • Use the History API, not fragments. Google's documentation says to "use the History API to implement routing between different views of your web app" and not to use fragments to load different page content. /#/products/42 is the same URL as / to a crawler.
  • Use <a href>. Google's crawlable links documentation says that "Google can only crawl your link if it's an <a> HTML element with an href attribute". It lists <a routerLink="…">, <span href="…">, <a onclick="goto(…)"> and javascript: URLs as not recommended; <button> navigation isn't a link at all. An <a href="/shoes" onclick="…"> is fine: the crawler reads the href and your router handles the click.

A minimal router that keeps crawlable links and intercepts clicks for client-side navigation:

router.js
const routes = new Map(); // pathname pattern -> async render(params, url)

export function route(pattern, render) {
  routes.set(new URLPattern({ pathname: pattern }), render);
}

async function renderURL(url, { replace = false, push = false } = {}) {
  for (const [pattern, render] of routes) {
    const match = pattern.exec(url.href);
    if (match) {
      if (push) history.pushState({}, "", url);
      if (replace) history.replaceState({}, "", url);
      await render(match.pathname.groups, url);
      return true;
    }
  }
  // Unknown route on the client: a full navigation lets the server answer
  // with its real status code (404, 301, ...), which crawlers need.
  location.assign(url);
  return false;
}

document.addEventListener("click", (event) => {
  // Respect modifier keys, middle clicks and links that opt out.
  if (event.defaultPrevented || event.button !== 0) return;
  if (event.metaKey || event.ctrlKey || event.shiftKey || event.altKey) return;
  const link = event.target.closest("a[href]");
  if (!link || link.target || link.hasAttribute("download")) return;
  const url = new URL(link.href);
  if (url.origin !== location.origin) return;
  event.preventDefault();
  renderURL(url, { push: true });
});

addEventListener("popstate", () => renderURL(new URL(location.href)));

URLPattern isn't available in every browser. Feature-detect it or use your framework's router; the principle (real href, History API, unknown routes go to the server) is what matters. If you use the Navigation API's navigate event instead of click handlers, the same rules hold: the links in the DOM must be real <a href> elements.

Real status codes, even in a single-page app

A soft 404 is a page that says "not found" but returns 200. Google detects many of them and excludes them, but it's better not to create them. For a server-rendered PWA, the server knows the route table and returns 404 (or 410 for permanently removed content) directly. For a client-rendered PWA, Google's JavaScript SEO basics documentation allows two strategies:

not-found.js
// The server must answer /not-found with HTTP 404.
export function showNotFound() {
  window.location.replace("/not-found");
}
not-found.js
export function showNotFound() {
  let robots = document.querySelector('meta[name="robots"]');
  if (!robots) {
    robots = document.createElement("meta");
    robots.name = "robots";
    document.head.append(robots);
  }
  // Adding noindex from JavaScript works; removing it doesn't.
  robots.content = "noindex";
  document.title = "Page not found";
  renderNotFoundView();
}

Status codes and what they do to indexing:

Response Search engine effect PWA-specific note
200 Rendered and considered for indexing Only for URLs that really exist
301 / 308 Strong canonical signal; target replaces source Use for moved routes. A client-side history.replaceState() is not a redirect.
302 / 307 Temporary; source usually stays indexed Fine for sign-in flows
404 / 410 Removed from the index over time Server must know the route table; don't answer unknown routes with the shell
401 / 403 Not indexed Correct for authenticated app routes
429 / 5xx Crawling slows down; persistent errors drop URLs Never answer server errors with your offline page and a 200

URL hygiene for PWAs

  • Launch parameters. A marker such as start_url: "/?source=pwa" (used for launch attribution) creates a second URL for your home page. Keep a canonical without the parameter, set an explicit manifest id so the marker isn't part of the app's identity (App Identity & Updates), and remove the parameter from the address bar after reading it with history.replaceState().
  • Shortcuts, share targets and file handlers point at URLs too. Give them noindex targets or canonicalize them when they render normal content with extra parameters.
  • Normalize. Choose trailing slash or not, lowercase paths, and one host (www or apex), and redirect the other variants with 301 on the server. The service worker doesn't see crawler requests, so it can't normalize for them.
  • Faceted and stateful URLs. Filters that live only in memory or localStorage can't be indexed. Put indexable states in the path or query string, and leave the rest out of links.
  • Infinite scroll. Google recommends giving each chunk a persistent unique URL (for example ?page=12), linking pages sequentially, and updating the address with the History API as chunks come into view (lazy-loading guide).

Canonical URLs in a PWA

The canonical URL tells search engines which of several duplicate URLs to index. Google's canonicalization documentation ranks the signals: redirects and rel="canonical" are strong signals, sitemap inclusion is a weak one. The rules that matter for PWAs:

  • Put <link rel="canonical" href="https://www.example.com/products/42"> in the server-rendered <head>, with an absolute URL, on every indexable page, including a self-referencing one on the canonical page itself.
  • If you must set it with JavaScript, Google's JavaScript SEO basics says that you "shouldn't use JavaScript to change the canonical URL to something else than the URL you specified as the canonical URL in the original HTML". Set the same value in both places, and make sure there is only one rel="canonical" element. Conflicting canonicals lead to unexpected results.
  • Don't use different canonicals through different methods (<link>, Link: header, sitemap) for the same page.
  • A canonical is a hint. Google can pick another URL if other signals disagree.

A client-side head manager must replace, not append, when the route changes:

head.js
/**
 * Keep per-route metadata in sync after client-side navigations.
 * The server-rendered HTML must already contain the same values, because
 * non-rendering crawlers and link-preview bots only see that.
 */
export function setRouteMetadata({ title, description, canonical, robots, jsonLd }) {
  document.title = title;
  upsertMeta("name", "description", description);
  upsertMeta("property", "og:title", title);
  upsertMeta("property", "og:description", description);
  upsertMeta("property", "og:url", canonical);

  let link = document.head.querySelector('link[rel="canonical"]');
  if (!link) {
    link = document.createElement("link");
    link.rel = "canonical";
    document.head.append(link);
  }
  link.href = new URL(canonical, location.origin).href; // always absolute

  if (robots) upsertMeta("name", "robots", robots);

  // One JSON-LD block owned by the router; other blocks are left alone.
  let script = document.head.querySelector('script[type="application/ld+json"][data-route]');
  if (jsonLd) {
    if (!script) {
      script = document.createElement("script");
      script.type = "application/ld+json";
      script.dataset.route = "";
      document.head.append(script);
    }
    script.textContent = JSON.stringify(jsonLd);
  } else {
    script?.remove();
  }
}

function upsertMeta(attr, key, content) {
  if (content == null) return;
  let el = document.head.querySelector(`meta[${attr}="${CSS.escape(key)}"]`);
  if (!el) {
    el = document.createElement("meta");
    el.setAttribute(attr, key);
    document.head.append(el);
  }
  el.setAttribute("content", content);
}

Common canonical mistakes in PWAs:

  • Every route canonicalizes to / because the shell template hardcodes it. This asks Google to drop every page but the home page.
  • The offline page or a cached shell carries a canonical of the page the user was on. Give /offline.html a noindex and no canonical.
  • The canonical includes the launch parameter (/?source=pwa) because it was built from location.href. Build it from your route data.

Titles, descriptions and per-route metadata

Search snippets come from the page, never from the manifest. Each indexable route needs:

  • A unique, descriptive <title> in the server HTML. The manifest's name and short_name are for the launcher and task switcher only.
  • A <meta name="description"> summarizing that page. Google may generate its own snippet, but a good description is often used.
  • <meta name="robots"> only where needed (noindex for search results pages, account areas, offline and error pages).
  • Open Graph and Twitter Card tags in the server HTML, because preview bots don't run JavaScript.
  • <html lang> matching the content language.

theme-color, apple-mobile-web-app-* tags and the manifest link are harmless to SEO and not ranking factors. They're covered in Splash Screens & Theming.

Structured data

Structured data describes page content to search engines in a machine-readable form and makes pages eligible for rich results. Use JSON-LD, render it on the server, and keep it consistent with visible content.

  • JavaScript-generated JSON-LD works for Google. Google can "understand and process structured data that's available in the DOM when it renders the page" (generate structured data with JavaScript). One caveat is specific to commerce: dynamically generated markup "can make Shopping crawls less frequent and less reliable", which matters for fast-changing price and availability. Render Product markup on the server.
  • Match the page. Markup must describe what the user sees on that URL. Updating visible content client-side without updating JSON-LD (or the reverse) creates mismatches.
  • Rich result types change. In June 2025 Google announced that it was phasing out seven structured data features (Book Actions, Course Info, Claim Review, Estimated Salary, Learning Video, Special Announcement and Vehicle Listing), stating that this doesn't affect ranking. Check the search gallery before investing in a type.
  • Describing the app itself. Google's Software app structured data supports SoftwareApplication and the subtypes WebApplication and MobileApplication. Required properties are name, offers.price (use 0 for free apps) and either aggregateRating or review. Only mark up ratings that are real and shown on the page.
Server-rendered JSON-LD for the app's landing page
<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "WebApplication",
  "name": "Acme Tasks",
  "url": "https://tasks.example.com/",
  "applicationCategory": "BusinessApplication",
  "operatingSystem": "Any (web browser); installable on Android, ChromeOS, iOS, macOS, Windows",
  "browserRequirements": "Requires JavaScript. Works offline after the first visit.",
  "offers": { "@type": "Offer", "price": 0, "priceCurrency": "USD" },
  "aggregateRating": { "@type": "AggregateRating", "ratingValue": 4.6, "ratingCount": 312 }
}
</script>

Validate with the Rich Results Test, which renders the page like Googlebot and shows the structured data it extracted.

Sitemaps and discovery

Client-side routers often leave pages that are hard to discover through links alone. An XML sitemap lists them. Google's sitemap rules:

  • A sitemap holds at most 50,000 URLs or 50 MB uncompressed. Beyond that, split it and list the parts in a sitemap index.
  • Use absolute, canonical URLs, UTF-8 encoding, and place the sitemap at the root to cover the whole site (a sitemap only covers descendants of its directory unless submitted in Search Console).
  • Google ignores <priority> and <changefreq>. It uses <lastmod> only when it's "consistently and verifiably accurate". Set it to the date of a meaningful change, not the build date.
  • Reference the sitemap from robots.txt with Sitemap: https://www.example.com/sitemap.xml and submit it in Search Console (or with the Search Console API). Google also accepts sitemap updates through WebSub for feeds.
  • The old sitemap "ping" endpoint is gone: Google announced its deprecation in June 2023. Accurate lastmod plus Search Console submission replaces it.

A build-time generator that emits a sitemap index and sitemaps with hreflang alternates:

scripts/build-sitemaps.mjs
// Usage: node scripts/build-sitemaps.mjs  (Node 18+)
// Reads route records from your CMS/API and writes dist/sitemap*.xml.
import { mkdir, writeFile } from "node:fs/promises";

const ORIGIN = "https://www.example.com";
const LOCALES = ["en", "de", "fr"];
const MAX_URLS = 50_000;

const escapeXml = (s) =>
  s.replace(/[<>&'"]/g, (c) => ({ "<": "&lt;", ">": "&gt;", "&": "&amp;", "'": "&apos;", '"': "&quot;" })[c]);

async function loadRoutes() {
  // Replace with your data source. Each record: { path, updatedAt, indexable }.
  const res = await fetch(`${ORIGIN}/api/routes?fields=path,updatedAt,indexable`);
  if (!res.ok) throw new Error(`Route API failed: ${res.status}`);
  const routes = await res.json();
  // Exclude noindex routes, the offline page, app-only and authenticated views.
  return routes.filter((r) => r.indexable && !r.path.startsWith("/app/"));
}

function urlEntry({ path, updatedAt }) {
  const alternates = LOCALES.map(
    (l) => `    <xhtml:link rel="alternate" hreflang="${l}" href="${escapeXml(`${ORIGIN}/${l}${path}`)}"/>`
  ).join("\n");
  // x-default points at the language selector / auto-detect page.
  const xDefault = `    <xhtml:link rel="alternate" hreflang="x-default" href="${escapeXml(`${ORIGIN}${path}`)}"/>`;
  return LOCALES.map(
    (l) => `  <url>
    <loc>${escapeXml(`${ORIGIN}/${l}${path}`)}</loc>
    <lastmod>${new Date(updatedAt).toISOString()}</lastmod>
${alternates}
${xDefault}
  </url>`
  ).join("\n");
}

async function main() {
  const routes = await loadRoutes();
  const entries = routes.flatMap((r) => urlEntry(r).split(/\n(?=  <url>)/));
  await mkdir("dist", { recursive: true });

  const files = [];
  for (let i = 0; i * MAX_URLS < entries.length; i++) {
    const chunk = entries.slice(i * MAX_URLS, (i + 1) * MAX_URLS);
    const name = `sitemap-${i + 1}.xml`;
    const xml = `<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9"
        xmlns:xhtml="http://www.w3.org/1999/xhtml">
${chunk.join("\n")}
</urlset>
`;
    if (Buffer.byteLength(xml) > 50 * 1024 * 1024) {
      throw new Error(`${name} exceeds 50 MB; lower MAX_URLS`);
    }
    await writeFile(`dist/${name}`, xml);
    files.push(name);
  }

  const index = `<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
${files.map((f) => `  <sitemap><loc>${ORIGIN}/${f}</loc></sitemap>`).join("\n")}
</sitemapindex>
`;
  await writeFile("dist/sitemap.xml", index);
  console.log(`Wrote ${files.length} sitemap(s), ${entries.length} URLs`);
}

main().catch((err) => {
  console.error(err);
  process.exitCode = 1;
});

Each <url> repeats the full set of alternates, including itself. That reciprocity is required for hreflang, as the next section explains.

For engines that support IndexNow, you can also push changed URLs as they're published:

scripts/indexnow.mjs
// POST changed URLs to IndexNow (Bing, Yandex and other participants).
// The key file must be served at https://www.example.com/<KEY>.txt containing the key.
const KEY = process.env.INDEXNOW_KEY; // 8-128 chars: a-z, A-Z, 0-9, dashes

export async function submitToIndexNow(urls) {
  if (!urls.length) return;
  const res = await fetch("https://api.indexnow.org/indexnow", {
    method: "POST",
    headers: { "Content-Type": "application/json; charset=utf-8" },
    body: JSON.stringify({
      host: "www.example.com",
      key: KEY,
      keyLocation: `https://www.example.com/${KEY}.txt`,
      urlList: urls.slice(0, 10_000), // protocol maximum per request
    }),
  });
  // 200/202: accepted. 400: bad format. 403: key not valid.
  // 422: URLs don't belong to host or key mismatch. 429: too many requests.
  if (res.status !== 200 && res.status !== 202) {
    throw new Error(`IndexNow rejected submission: HTTP ${res.status}`);
  }
}

robots.txt for PWAs

  • Don't disallow the JavaScript, CSS, JSON API endpoints or images needed to render a page. If Googlebot can't fetch /api/products/42, the rendered product page is empty.
  • Disallowing /sw.js or the manifest is pointless: they have no search effect either way.
  • robots.txt controls crawling, not indexing. Use noindex (which requires the page to be crawlable) to keep pages out of the index.

hreflang and international PWAs

International PWAs usually share one origin, one service worker and one manifest (or one per locale). Search engines need each language version at its own URL, annotated with hreflang. Google's localized versions documentation:

  • Annotate with <link rel="alternate" hreflang="…"> in HTML, Link HTTP headers, or the sitemap. The three methods are equivalent; pick one.
  • Annotations must be reciprocal: "If two pages don't both point to each other, the tags will be ignored." Each version lists every version, including itself.
  • Codes are ISO 639-1 language plus optional ISO 3166-1 alpha-2 region (en, en-GB, de-AT). Region-only codes aren't valid, and neither are UK or EU.
  • Use x-default for the fallback or language-picker page.
  • Use fully qualified URLs.
In the server-rendered head of /de/products/42
<link rel="alternate" hreflang="en" href="https://www.example.com/en/products/42">
<link rel="alternate" hreflang="de" href="https://www.example.com/de/products/42">
<link rel="alternate" hreflang="fr" href="https://www.example.com/fr/products/42">
<link rel="alternate" hreflang="x-default" href="https://www.example.com/products/42">
<link rel="canonical" href="https://www.example.com/de/products/42">

PWA-specific rules:

  • Don't choose the language on the client only. A locale stored in localStorage or IndexedDB doesn't exist for crawlers, whose storage is cleared between page loads. Google's locale-adaptive pages documentation adds that Googlebot's default IP addresses appear to be in the USA and that it sends requests without an Accept-Language header. A PWA that picks the language from the header or IP and serves it at one URL gets one language indexed.
  • Canonicalize each language to itself, not to the English version. Cross-language canonicals tell Google that the translations are duplicates.
  • One service worker per origin serves all locales. Precache locale-specific bundles per locale, or cache them at runtime, as Pinterest did with per-locale service worker builds (see Case Studies).
  • Per-locale manifests can localize name and description, but those don't affect search. start_url should point to the locale's path.

Core Web Vitals and page experience

Google's documentation is explicit: "Core Web Vitals are used by our ranking systems" (page experience). It's equally explicit about the limits:

  • "There is no single signal." Core ranking systems look at a variety of signals that align with page experience.
  • Relevance wins: Google "always seeks to show the most relevant content, even if the page experience is sub-par".
  • Good results in the Core Web Vitals report "doesn't guarantee that your pages will rank at the top".
  • Evaluation is generally page-specific, with some site-wide assessments.

The thresholds are the Core Web Vitals "good" values at the 75th percentile of real-user data: LCP ≤ 2.5 s, INP ≤ 200 ms, CLS ≤ 0.1. INP replaced FID as a Core Web Vital on March 12, 2024. The data comes from the Chrome UX Report (CrUX) field data, not from Lighthouse.

How PWA architecture shows up in these numbers is covered in Core Web Vitals. The SEO-relevant points:

  • Search traffic is mostly first visits. Your service worker can't help a visitor who has never been to your site, and precaching during that first visit competes for bandwidth. Server-rendered HTML and a small critical path matter more for search landing pages than any caching strategy. See Loading Performance.
  • Repeat visits count too. CrUX aggregates all eligible Chrome page loads, so a service worker that serves assets from Cache Storage improves the field distribution. A worker that adds startup latency to every navigation makes it worse. Navigation preload or the Static Routing API remove that cost.
  • Soft navigations. Chrome measures Core Web Vitals per hard navigation. A single-page PWA's route changes are "soft navigations". Chrome 151 enables soft-navigation measurement by default in the web platform, but Chrome's own soft navigations documentation says that how CrUX will report them "is also still to be determined". Until then, the landing page's hard navigation is what CrUX sees, and one slow interaction late in a long SPA session counts against that page's INP.
  • Install prompts and interstitials. A full-screen custom install promotion shown on arrival is the kind of intrusive interstitial that Google's page experience guidance tells you to avoid, and it can cause layout shifts. Follow the patterns in Install Prompts & Custom UI.

The manifest, installability and ranking

Google documents no ranking signal tied to the manifest, service worker registration or installability, and none of the page experience documentation mentions them. Being a PWA helps SEO only through what it does to users: faster repeat visits, better engagement and better Core Web Vitals. A few related facts:

  • Lighthouse removed its PWA category in Lighthouse 12. See Lighthouse & Auditing. The removal has no search implications, because the category never influenced ranking.
  • A PWA published to Google Play as a Trusted Web Activity is still indexed as a website. The store listing is a separate Play Store entity.
  • related_applications and prefer_related_applications in the manifest influence the browser's install UI, not search results.

Offline pages, stale caches and other PWA pitfalls

  • Offline pages must never be served by the server. Your service worker serves /offline.html locally; crawlers never see it. The danger is a server or CDN configured to answer errors or unknown routes with the offline page and a 200. Give /offline.html a <meta name="robots" content="noindex">, keep it out of the sitemap, and answer real errors with real status codes.
  • Keep old hashed assets deployable. Google's renderer "may ignore caching headers" and render your latest HTML with JavaScript it cached earlier. If app.3f9c1e.js disappears after a deploy, a render that references it fails. Keep previous builds' assets for at least a few days, which also helps users whose service worker hasn't updated yet.
  • Don't block rendering on APIs that treat bots differently. Rate limiters, bot-protection challenges and geo-blocks that return 403 or a challenge page to Googlebot's API calls produce empty rendered pages, even though the HTML request succeeded.
  • Check how your Web Components expose content. Google's renderer flattens the light DOM and shadow DOM; for components that don't project light DOM content through <slot>, Google tells you to consult its documentation, and content that appears only after interaction is never rendered. Verify component-heavy pages in the URL Inspection rendered HTML.
  • Don't hide content behind "open in app" walls. A page that requires installation or sign-in to show its content has no content to index.
  • Watch authenticated app areas. Routes behind login should return 401/403 or redirect to sign-in, and ideally live under a separate path (such as /app/) that isn't linked from public pages.

Testing and debugging

URL Inspection in Search Console

The URL Inspection tool shows what Google fetched and rendered:

  • Indexed result: what Google has in the index from the last crawl.
  • Live test: fetches and renders the URL now. View tested page shows the rendered HTML, a screenshot, the HTTP response and headers, the page resources (including the ones that failed to load) and JavaScript console messages.
  • The live test doesn't check everything indexing does: manual actions, canonical selection and duplicate detection aren't part of it.
  • Request indexing has a daily quota. Use sitemaps for bulk changes.

Check the rendered HTML for your main content, the <title>, the canonical and the structured data. Check the resources list for blocked API calls and the console for errors thrown by code that assumed a service worker, storage or permissions.

A local raw-versus-rendered comparison

Before deploying, compare what a non-rendering crawler gets (raw HTML) with what a renderer gets, with no service worker, no storage and no cookies. This Puppeteer script flags routes whose indexable content exists only after JavaScript runs:

scripts/render-check.mjs
// Usage: node scripts/render-check.mjs https://staging.example.com/ /products/42 /about
import puppeteer from "puppeteer";

const [origin, ...paths] = process.argv.slice(2);
if (!origin || !paths.length) {
  console.error("Usage: render-check.mjs <origin> <path> [path...]");
  process.exit(2);
}

const UA =
  "Mozilla/5.0 (Linux; Android 10; K) AppleWebKit/537.36 (KHTML, like Gecko) " +
  "Chrome/140.0.0.0 Mobile Safari/537.36 (compatible; render-check)";

function summarize(html) {
  const pick = (re) => (html.match(re)?.[1] ?? "").trim();
  const text = html
    .replace(/<script[\s\S]*?<\/script>|<style[\s\S]*?<\/style>/gi, " ")
    .replace(/<[^>]+>/g, " ")
    .replace(/\s+/g, " ")
    .trim();
  return {
    title: pick(/<title[^>]*>([\s\S]*?)<\/title>/i),
    canonical: pick(/<link[^>]+rel=["']canonical["'][^>]*href=["']([^"']+)/i),
    robots: pick(/<meta[^>]+name=["']robots["'][^>]*content=["']([^"']+)/i),
    words: text ? text.split(" ").length : 0,
  };
}

const browser = await puppeteer.launch();
let failures = 0;
try {
  for (const path of paths) {
    const url = new URL(path, origin).href;

    // 1. Raw HTML, as a non-rendering crawler sees it.
    const rawRes = await fetch(url, { headers: { "User-Agent": UA }, redirect: "manual" });
    const raw = summarize(await rawRes.text());

    // 2. Rendered DOM in a fresh incognito context: no SW, no storage, no cookies.
    const context = await browser.createBrowserContext();
    const page = await context.newPage();
    await page.setUserAgent(UA);
    await page.setViewport({ width: 412, height: 915, isMobile: true });
    // Block service worker registration to mimic crawlers.
    await page.evaluateOnNewDocument(() => {
      if (navigator.serviceWorker) {
        navigator.serviceWorker.register = () => Promise.reject(new Error("blocked"));
      }
    });
    const errors = [];
    page.on("pageerror", (err) => errors.push(err.message));
    const res = await page.goto(url, { waitUntil: "networkidle0", timeout: 30_000 });
    const rendered = summarize(await page.content());
    await context.close();

    const problems = [];
    // page.goto() resolves to null for same-document navigations; treat as unknown.
    const renderedStatus = res ? res.status() : 0;
    if (rawRes.status !== renderedStatus) problems.push(`status raw ${rawRes.status} vs rendered ${renderedStatus}`);
    if (!raw.title || raw.title !== rendered.title) problems.push(`title "${raw.title}" -> "${rendered.title}"`);
    if (raw.canonical !== rendered.canonical) problems.push(`canonical "${raw.canonical}" -> "${rendered.canonical}"`);
    if (raw.words < rendered.words * 0.5) problems.push(`raw HTML has ${raw.words} words, rendered ${rendered.words}`);
    if (errors.length) problems.push(`JS errors: ${errors.join(" | ")}`);

    console.log(`${problems.length ? "FAIL" : "ok  "} ${url}`);
    for (const p of problems) console.log(`     - ${p}`);
    failures += problems.length ? 1 : 0;
  }
} finally {
  await browser.close();
}
process.exitCode = failures ? 1 : 0;

Run it in CI against a preview deployment, alongside the checks in Automated Testing.

Quick manual checks

  • curl -s https://www.example.com/products/42 | grep -i "<title>\|canonical\|robots" shows what non-rendering crawlers get.
  • curl -sI https://www.example.com/does-not-exist must return 404, not 200.
  • In Chrome DevTools, disable JavaScript (Command Menu, Ctrl+Shift+P or Cmd+Shift+P on macOS, then "Disable JavaScript") and reload in a fresh profile to see the no-JS view. See Browser DevTools.
  • Search Console's Page indexing report lists soft 404s, "Crawled – currently not indexed" and canonical conflicts. The Core Web Vitals report shows field data grouped by similar URLs.

Common pitfalls

  • Serving the same app shell HTML for every URL, with one title and a 200 for nonexistent routes.
  • Hash-based routing (/#/products) inherited from an old SPA template.
  • Navigation implemented with onclick handlers on <div> or <button> elements instead of <a href>.
  • Content that waits for navigator.serviceWorker.ready, a cookie, localStorage data or a permission before rendering.
  • A canonical built from location.href that includes ?source=pwa or other launch parameters.
  • JavaScript that removes a server-set noindex, expecting Google to index the page.
  • Open Graph tags added only on the client.
  • Main bundles larger than 2 MB uncompressed.
  • Deleting the previous build's hashed assets on deploy.
  • Language chosen from Accept-Language, IP address or storage at a single URL.
  • A CDN rule that answers origin errors with the cached offline page and a 200.

Further reading

On this site

External references