Loading Performance¶
Loading performance is everything that happens between a user tapping a link or your PWA's home-screen icon and the page being useful: the network requests, their order and priority, how many bytes cross the wire, and how much parsing and rendering the browser must do before the first meaningful paint. A service worker changes the rules for repeat visits, but it cannot help the first visit, and it introduces its own costs, so a fast PWA still needs a lean critical path, correct resource priorities, good compression and well-built images and fonts. This page goes through each stage of the loading pipeline at the mechanism level, with the browser-specific details and production code you need to make both first and repeat loads fast.
Key takeaways
- The critical rendering path is gated by render-blocking CSS and parser-blocking scripts. Everything else is a question of discovery (can the preload scanner see it?) and priority (does it compete with the LCP resource?).
fetchpriorityis supported in all three engines (Chrome 102, Safari 17.2, Firefox 132). Usefetchpriority="high"on the LCP image andlowon everything that can wait; do not use it to paper over late discovery.- Speculation Rules (Chromium only) prefetch or prerender whole navigations. Since Chrome 138, prefetches to service-worker-controlled URLs go through the worker's
fetchhandler instead of being cancelled, so PWAs can finally use them. - zstd
Content-Encodingis supported in Chrome 123, Firefox 126 and Safari 26.3 (only on macOS Tahoe 26.3, iOS 26.3 and later). Brotli remains the best choice for precompressed static assets; zstd shines for on-the-fly compression. - 103 Early Hints work only on top-level navigations over HTTP/2 or HTTP/3. Chrome and Firefox honor
preloadandpreconnect; Safari 17+ honors onlypreconnect. A navigation answered from Cache Storage never reaches the server, so it never sees the hints. - A PWA has two loading problems: the first visit (no worker, cold caches, precaching competes with the page) and the repeat visit (worker start-up, cache reads, stale code). Optimize both, and measure them separately.
How the browser turns bytes into pixels¶
Before you can make a page load faster you need an accurate model of what the browser waits for. The critical rendering path is the minimum sequence of work needed to paint the first frame: fetch the HTML, parse it into the DOM, fetch and parse CSS into the CSSOM, run any scripts that block parsing, compute style, lay out and paint.
flowchart LR
A["Navigation request"] --> B["HTML bytes arrive"]
B --> C["HTML parser builds DOM"]
B --> S["Preload scanner discovers URLs"]
S --> N["Subresource requests"]
C -->|"sync script"| J["Fetch + execute script"]
J --> C
N --> CSS["CSS parsed to CSSOM"]
C --> R["Style + layout"]
CSS --> R
R --> P["First paint"] Render-blocking and parser-blocking resources¶
Two kinds of resources stop the first frame:
- Render-blocking stylesheets. A
<link rel="stylesheet">in the<head>whosemediaquery matches blocks rendering until it has been downloaded and parsed. The browser keeps parsing HTML and fetching other resources, but it will not paint. Stylesheets whosemediadoes not match (for examplemedia="print", ormedia="(min-width: 1200px)"on a phone) are still downloaded, at the lowest priority, but they do not block rendering. - Parser-blocking scripts. A classic
<script src>withoutasyncordeferstops the HTML parser until the script has been fetched and executed. Because scripts can read computed styles, a parser-blocking script also waits for every stylesheet that precedes it, which chains CSS latency into script latency.
Chromium (since Chrome 105) and Safari (since 18.2) also support an explicit opt-in, blocking="render", on <script>, <style> and <link> elements (Firefox does not support it), which makes an element render-blocking even though it would not be otherwise. It is useful for small, essential scripts that must run before first paint (for example a theme selector that prevents a flash of the wrong color scheme), and dangerous for anything large.
The table below summarizes how each way of loading JavaScript interacts with the parser:
| Markup | Blocks parser | Executes | Order guaranteed | Typical use |
|---|---|---|---|---|
<script src> | Yes, fetch and execution | Immediately when fetched | Yes, in document order | Avoid in <head> |
<script defer src> | No | After parsing, before DOMContentLoaded | Yes | Classic app bundles |
<script async src> | No (execution interrupts parsing) | As soon as fetched | No | Independent third-party scripts |
<script type="module" src> | No (deferred by default) | After parsing, before DOMContentLoaded | Yes | Modern app entry points |
<script type="module" async src> | No | As soon as it and its imports are fetched | No | Independent module widgets |
import() | No | When the promise resolves | N/A | Route and component code splitting |
Inline <script> | Yes (execution only) | Immediately | Yes | Tiny bootstrap code |
A module script's dependency graph must be fully fetched before it executes. A static import three levels deep costs three sequential round trips unless the browser learns about the deeper modules earlier, which is exactly what modulepreload (below) and bundlers are for.
The preload scanner¶
Browsers run a lightweight preload scanner (also called the speculative parser) ahead of the main HTML parser. When the main parser is blocked on a script, the scanner keeps reading the raw markup and starts fetching the URLs it finds in <link rel=stylesheet>, <script src>, <img src/srcset>, <link rel=preload> and similar elements. It is responsible for much of the parallelism you see in a waterfall.
The scanner only sees what is literally in the HTML. It cannot discover:
- resources referenced from CSS (
background-image,@font-face,@import), - resources injected by JavaScript (client-rendered images, dynamically created
<script>elements), - images that use
data-srclazy-loading libraries instead of nativeloading="lazy", - anything in a client-side-rendered app shell that is filled by an API response.
In an app-shell PWA this is the single most common cause of a slow LCP on the first visit: the LCP image is only known after the shell's JavaScript has run and fetched data. Either render the LCP element's markup on the server, or make it discoverable with <link rel="preload"> (see App Shell Model for the architectural trade-off).
Critical CSS¶
Because every byte of render-blocking CSS delays the first paint, a common technique is to inline the CSS needed for the first viewport in a <style> element and load the rest without blocking:
<style>
/* Critical CSS: layout of the shell, header, above-the-fold typography. Keep it
under roughly 14 KB compressed together with the rest of the head so the first
round trip carries everything needed for the first paint. */
:root { color-scheme: light dark; }
body { margin: 0; font: 16px/1.5 system-ui, sans-serif; }
.app-header { position: sticky; top: 0; height: 56px; }
</style>
<link rel="preload" href="/css/app.4f9a1c.css" as="style">
<link rel="stylesheet" href="/css/app.4f9a1c.css" media="print" onload="this.media='all'">
<noscript><link rel="stylesheet" href="/css/app.4f9a1c.css"></noscript>
The media="print" trick downloads the stylesheet without blocking rendering and switches it on when it has loaded. It has costs: inline handlers need 'unsafe-hashes' or a hash under a strict Content Security Policy, and a late-applied stylesheet can cause layout shifts if the critical CSS is incomplete.
For a PWA the trade-off shifts on repeat visits. Inlined CSS cannot be cached separately: it is part of every HTML response. If your service worker serves the HTML from cache, the inlined CSS costs nothing extra on repeat visits; if HTML always comes from the network (a network-first strategy for fresh content), every navigation re-downloads it. Keep inline CSS small and let the full stylesheet be a long-lived, precached, fingerprinted file.
Where the service worker sits in the pipeline¶
Once a service worker controls a page, every request in the diagram above, the navigation and each subresource, is first dispatched to the worker's fetch event (unless a static route sends it elsewhere). For the navigation request this happens before the browser has any HTML, so the worker's start-up time lands directly on the critical path. Navigation preload and static routing exist to take it off again. Subresource requests are dispatched to an already running worker, so their overhead is small but not zero: an event dispatch, a thread hop and, for cache hits, a Cache Storage read.
The rest of this page treats the network as the thing you optimize, but keep in mind that on a repeat visit many of these requests never reach it. That is why every section below has a note on how the technique interacts with a worker.
Resource priorities and fetchpriority¶
Browsers do not fetch everything at once at the same urgency. Chromium assigns every request a priority (Highest, High, Medium, Low, Lowest) that controls the order in which requests are sent and, over HTTP/2 and HTTP/3, how the server should share bandwidth between them. The defaults, as documented by web.dev's Fetch Priority article, look roughly like this:
| Resource | Default Chromium priority | Notes |
|---|---|---|
| Main document | Highest | |
CSS in <head> | Highest | Non-matching media stylesheets drop to Lowest |
| Fonts | High | Discovered late, from CSS, unless preloaded |
Scripts before the first image in <body> | High | Parser-blocking scripts late in the body get Medium |
async / defer / module scripts | Low | |
| Images | Low initially | Boosted to High once layout finds them in the viewport |
| The first five large images (> 10,000 px²) | Medium | A Chromium heuristic, so hero images start ahead of icons |
fetch() and XHR | High | |
<link rel=prefetch> | Lowest | Intended for the next navigation |
Firefox and Safari have their own schemes; the details differ but the principle is the same: the browser guesses, and it guesses wrong whenever the most important resource looks unimportant (a hero image that is only boosted after layout) or an unimportant resource looks important (a below-the-fold carousel's first images).
The fetchpriority attribute¶
The fetchpriority attribute (on <img>, <link> and <script>) tells the browser that a resource is more or less important than its type suggests. It takes high, low or auto (the default) and adjusts the priority relative to the default rather than setting an absolute level:
<!-- The LCP image: start at High instead of Low, before layout runs. -->
<img src="/img/hero-800.avif" width="800" height="450" alt="Espresso machine"
fetchpriority="high">
<!-- Images in a carousel that start off-screen: deprioritize. -->
<img src="/img/slide-2.avif" width="800" height="450" alt="" fetchpriority="low">
<!-- A script that is async but needed early: raise it above other async scripts. -->
<script src="/js/checkout.js" async fetchpriority="high"></script>
<!-- A preload that should not compete with the LCP image. -->
<link rel="preload" href="/fonts/icons.woff2" as="font" type="font/woff2" crossorigin
fetchpriority="low">
The same knob exists for JavaScript-initiated requests as the priority option of fetch() and new Request():
// Background data the user may need soon: never compete with the page's own
// critical requests.
const response = await fetch("/api/suggestions", { priority: "low" });
web.dev reports that Google Flights improved LCP from 2.6 s to 1.9 s by adding fetchpriority="high" to its hero image (source). Improvements of that size occur when the LCP image is discoverable early but starts at Low priority and waits behind scripts and other images. If the image is discovered late (injected by JavaScript), fetchpriority alone changes little: fix discovery first.
Resource hints: preconnect, dns-prefetch, preload, modulepreload and prefetch¶
Resource hints are <link> elements (or equivalent Link HTTP headers) that tell the browser about work it will need to do before the parser would discover it. They differ in what they do and when the result is used:
| Hint | What the browser does | Scope | Cache the result lives in |
|---|---|---|---|
dns-prefetch | Resolves the hostname | Current page | DNS cache |
preconnect | DNS + TCP/QUIC + TLS handshake | Current page | Connection pool (idle connections are closed after a short time) |
preload | Fetches one resource with the given as destination at its normal priority | Current page | Memory cache for this document, plus the HTTP cache |
modulepreload | Fetches, parses and compiles one module script into the module map | Current page | Module map, plus the HTTP cache |
prefetch | Fetches a resource at Lowest priority for a future navigation | Next page | HTTP cache (Chromium keeps it for about five minutes regardless of cacheability) |
Speculation Rules prefetch / prerender | Fetches (or fully renders) a navigation | Next page | A per-document prefetch cache or prerendered page |
preconnect and dns-prefetch¶
Opening a secure connection to a new origin costs a DNS lookup, a TCP handshake and a TLS handshake over TCP (or a combined QUIC handshake over HTTP/3): typically one to three round trips, which on a mobile network is easily 300–600 ms. preconnect does that work in advance:
<!-- Fonts are fetched in CORS mode; a CORS connection is separate from a
no-CORS one, so the crossorigin attribute must match the eventual request. -->
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
<!-- API calls with credentials use a connection opened without crossorigin=anonymous. -->
<link rel="preconnect" href="https://api.example.com">
<!-- Fallback and cheap hint for origins that may or may not be used. -->
<link rel="dns-prefetch" href="https://images.example-cdn.com">
Rules that follow from the mechanism:
- Match the CORS mode. Browsers pool connections separately for credentialed and anonymous requests.
<link rel=preconnect crossorigin>opens an anonymous-mode connection, which is what@font-facerequests andfetch(..., { mode: "cors", credentials: "omit" })use. A preconnect with the wrong mode opens a connection that is never used. - Only preconnect to origins you will use within a few seconds. Idle connections are closed after a short time (Chromium closes unused preconnected sockets after about 10 seconds), so an early preconnect for a resource requested much later is wasted, and every connection costs CPU on the device and the server. Two to four preconnects is a practical ceiling.
- Use
dns-prefetchfor "maybe" origins. It is much cheaper and can be paired withpreconnectas a fallback for older engines. - Preconnect is pointless for your own origin. The navigation already opened that connection. With HTTP/2 or HTTP/3 connection coalescing, a separate hostname that resolves to the same server and is covered by the same certificate may also reuse the connection.
In a PWA, preconnects in the HTML shell matter on repeat visits too: if the shell comes from Cache Storage and the page then calls an API origin, the connection to that origin is cold. A preconnect in the cached shell, or a Link: <https://api.example.com>; rel=preconnect header on the navigation response when it comes from the network, removes that handshake from the critical path.
preload¶
<link rel="preload"> fetches a specific resource for the current page, with the priority of its declared type, and holds it until the page uses it. It exists for resources that are critical but late-discovered: fonts referenced from CSS, the LCP background image, the data file a client-rendered view needs first, the main module of a lazily imported route.
<!-- A font discovered only after the CSS has been parsed. crossorigin is required
for fonts, even same-origin ones, because fonts are always fetched in CORS mode. -->
<link rel="preload" href="/fonts/inter-var-latin.woff2" as="font" type="font/woff2" crossorigin>
<!-- A responsive LCP image used as a CSS background: preload with the same
candidates and sizes the browser would compute for an <img>. -->
<link rel="preload" as="image" href="/img/hero-800.avif"
imagesrcset="/img/hero-400.avif 400w, /img/hero-800.avif 800w, /img/hero-1600.avif 1600w"
imagesizes="100vw" type="image/avif" fetchpriority="high">
<!-- JSON the first view needs. as="fetch" must be matched by a fetch() with the
same mode and credentials, or the preload is not reused. -->
<link rel="preload" href="/api/home.json" as="fetch" crossorigin>
Everything about preload is about matching: the preloaded response is only reused if the later request matches it. The as value sets the request destination and Accept header; crossorigin sets the mode and credentials; type lets browsers that do not support a format skip the preload; imagesrcset/imagesizes must select the same candidate the <img> or CSS would. Mismatches do not fail loudly: the browser downloads the resource twice. Chromium prints a console warning when a preloaded resource was not used within a few seconds of the load event, which is the easiest way to catch mistakes.
Valid as values include audio, document, embed, fetch, font, image, object, script, style, track, video and worker. A preload without as, or with an unknown value, is ignored or fetched at the wrong priority.
The most common preload mistakes:
- Preloading what the preload scanner already finds. A preload of a stylesheet that is already a
<link rel=stylesheet>in the same<head>achieves nothing. - Preloading too much. Preloads are fetched at High or Highest priority, so ten preloads compete with each other and with the LCP resource. Preload only what is on the critical path.
- Preloading resources for other pages. That is what
prefetchand Speculation Rules are for. - Preloading responses that are not cacheable. A
no-storeresource preloaded by one mechanism and requested by another can end up downloaded twice.
Preloads and your service worker
A preload is an ordinary subresource request from a controlled page, so it is dispatched to your worker's fetch event. If the worker serves the resource from cache, the preload costs almost nothing on repeat visits. If it uses network-first for that URL, the preload's early start is preserved but so is the network cost. The preload must still match the eventual request as seen by the browser, not by the worker.
modulepreload¶
<link rel="modulepreload"> is preload for ES modules. Unlike preload as=script, it fetches the module in the right mode (CORS, with same-origin credentials), and it parses and compiles the module and puts it into the document's module map, so the eventual import resolves without another fetch. Chrome 66, Firefox 115 and Safari 17 support it.
<script type="module" src="/js/main.a1b2c3.js"></script>
<!-- The static imports of main.js, so the whole graph downloads in parallel instead
of in sequential waves. Bundlers can generate these lines for you. -->
<link rel="modulepreload" href="/js/router.d4e5f6.js">
<link rel="modulepreload" href="/js/store.071829.js">
<link rel="modulepreload" href="/js/vendor-lit.3a4b5c.js">
The HTML specification lets browsers also fetch the preloaded module's dependencies, but does not require it, so list every module in the graph that you want fetched in parallel. Vite, for example, emits modulepreload links for the static imports of each entry and injects a small runtime that adds them for dynamically imported chunks too. caniuse notes one limitation: once a module is in the module map, changing attributes such as crossorigin on the link does not trigger a new fetch.
prefetch¶
<link rel="prefetch"> fetches a resource at the lowest priority, when the browser is otherwise idle, for use on a future navigation. Chromium and Firefox support it; in Safari it is still behind a feature flag, so treat it as a Chromium and Firefox optimization. Chromium stores it in the HTTP cache and treats it as usable for about five minutes even if its cache headers would not normally allow reuse, so that the next page can use it. Browsers send Sec-Purpose: prefetch on these requests, which lets your server or CDN distinguish them from real traffic.
In a PWA, rel=prefetch competes with a better tool: the service worker's own cache. Resources you know the user will need (the next route's chunk, the shell of the checkout flow) can be put into Cache Storage during installation or at idle time (see Precaching & Runtime Caching), where they survive far longer than five minutes and work offline. rel=prefetch still has two good uses: prefetching on first visits, before the worker exists, and prefetching cross-origin resources that you do not want in your cache.
Hints as HTTP headers¶
Every hint can be sent in a Link response header instead of markup, which lets the browser act on it before the HTML body has been parsed, and lets a CDN add hints without touching templates:
Link: </css/app.4f9a1c.css>; rel=preload; as=style,
</fonts/inter-var-latin.woff2>; rel=preload; as=font; type="font/woff2"; crossorigin,
<https://api.example.com>; rel=preconnect
A Link header on a response that the service worker serves from Cache Storage is not acted upon as early: the headers were cached with the response, and although browsers process Link headers on the navigation response the worker returns, the head start relative to parsing is gone because the whole cached body is available immediately. For controlled navigations, markup hints in the cached shell are just as good.
103 Early Hints¶
Early Hints (RFC 8297) solve a timing problem that Link headers cannot: headers are sent with the final response, but the final response often waits for "server think-time" (database queries, template rendering). A server that knows which subresources a page will need can send an interim 103 Early Hints response immediately, and the browser starts preconnecting and preloading while the server is still working:
sequenceDiagram
participant B as Browser
participant S as Server
B->>S: GET /products/42
S-->>B: 103 Early Hints (Link: preload app.css, preconnect cdn)
Note over B: Starts fetching app.css and connecting to the CDN
Note over S: Server think-time (database, rendering)
S-->>B: 200 OK (HTML)
Note over B: app.css is already in flight or done The restrictions are strict and the browsers differ:
| Chrome / Edge | Firefox | Safari | |
|---|---|---|---|
preconnect in 103 | ✅ 103 | ✅ 120 | ✅ 17 |
preload in 103 | ✅ 103 | ✅ 123 | ❌ |
modulepreload, prefetch, dns-prefetch in 103 | ❌ | ❌ | ❌ |
Support data as of September 2026, from Chrome's Early Hints documentation. See MDN's 103 page for live compatibility data.
- Navigations only. Chrome ignores 103 responses on subresource requests and on iframe navigations; they apply to top-level document loads.
- HTTP/2 or HTTP/3 only. Chromium ignores 103 responses received over HTTP/1.1, because many HTTP/1.1 intermediaries mishandle interim responses.
- Only the first 103 counts. Chromium processes the first Early Hints response and ignores later ones.
- Cross-origin redirects discard them. If the final response is a redirect to another origin, preloaded resources are discarded.
- Preloads land in the HTTP cache. Only cacheable resources should be preloaded this way; otherwise the page may download them again.
Many CDNs can generate 103 responses from the Link headers they saw on previous responses. If you run your own server, Node.js exposes writeEarlyHints() on both the HTTP/1.1 and the HTTP/2 compatibility API. Because browsers ignore 103 over HTTP/1.1, use the HTTP/2 server (or a proxy that speaks HTTP/2 or HTTP/3 to the browser and forwards interim responses):
import { createSecureServer } from "node:http2";
import { readFileSync } from "node:fs";
const server = createSecureServer({
key: readFileSync("certs/key.pem"),
cert: readFileSync("certs/cert.pem"),
allowHTTP1: true, // HTTP/1.1 clients still work; they simply ignore the 103.
});
// Hints are the same for every product page, so compute them once.
const EARLY_HINTS = [
"</css/app.4f9a1c.css>; rel=preload; as=style",
"</js/main.a1b2c3.js>; rel=preload; as=script; crossorigin",
"<https://images.example-cdn.com>; rel=preconnect",
];
server.on("request", async (req, res) => {
const url = new URL(req.url, "https://example.com");
if (req.method === "GET" && url.pathname.startsWith("/products/")) {
// Send the 103 before any slow work. Only do this for navigations: the
// Sec-Fetch-Dest header tells you whether the request is for a document.
if (req.headers["sec-fetch-dest"] === "document") {
res.writeEarlyHints({ link: EARLY_HINTS });
}
try {
const html = await renderProductPage(url); // database + templating
res.writeHead(200, {
"content-type": "text/html; charset=utf-8",
"cache-control": "no-cache",
// Repeat the hints on the final response for browsers and caches that
// ignored or never saw the 103.
link: EARLY_HINTS.join(", "),
});
res.end(html);
} catch (error) {
console.error(error);
res.writeHead(500, { "content-type": "text/plain" });
res.end("Internal Server Error");
}
return;
}
res.writeHead(404);
res.end();
});
server.listen(8443);
async function renderProductPage(url) {
// Stand-in for your real rendering pipeline.
return `<!doctype html><title>${url.pathname}</title>`;
}
Note the crossorigin on the script preload: module scripts are fetched in CORS mode, and the preload must match (or use rel=modulepreload in the final response's markup, since Early Hints do not support it).
Early Hints and service workers. A 103 is part of the response to a network request. When the service worker answers a navigation from Cache Storage, no request reaches your server, and there is nothing to hint; the cached shell's own <link> elements take over. For network-first navigations the request does reach the server; verify in DevTools (the Network panel shows resources initiated by Early Hints) that your hints are used in your target browsers when the navigation goes through the worker or through navigation preload, rather than assuming it. Early Hints also changes TTFB accounting; see Core Web Vitals for how responseStart and finalResponseHeadersStart treat the interim response.
Speculation Rules: prefetching and prerendering navigations¶
The Speculation Rules API lets a page declare which future navigations the browser may prefetch (download the document response) or prerender (load and render the whole page in a hidden tab, running its JavaScript) ahead of time. When the user then navigates, a prefetched page skips the network round trip for the document, and a prerendered page is activated instantly: LCP can be close to zero.
It replaces <link rel=prerender> (which Chrome turned into a "NoState Prefetch") and is more powerful than rel=prefetch because prefetched navigations are stored in a per-document cache with the correct semantics for a navigation (cookies, redirects, Vary), rather than as a generic subresource.
Syntax¶
Rules are JSON in a <script type="speculationrules"> element, or in an external file referenced by a Speculation-Rules response header:
<script type="speculationrules">
{
"prerender": [
{
"where": {
"and": [
{ "href_matches": "/*" },
{ "not": { "href_matches": ["/logout", "/cart/*", "/api/*"] } },
{ "not": { "selector_matches": "[rel~=nofollow], .no-prerender" } }
]
},
"eagerness": "moderate"
}
],
"prefetch": [
{
"urls": ["/products/", "/deals/"],
"eagerness": "immediate"
}
]
}
</script>
The two rule sources:
- List rules (
"urls": [...], implying"source": "list") speculate specific URLs. Their default eagerness isimmediate. - Document rules (
"where": {...}, implying"source": "document") match links in the page withhref_matches(URL Pattern syntax) andselector_matches(CSS selectors), combined withand,orandnotto any depth. Their default eagerness isconservative. Links added to the page later are matched automatically, which makes document rules a natural fit for client-rendered PWAs.
Other rule fields: referrer_policy (cross-site prefetches require a strict policy such as strict-origin-when-cross-origin), relative_to ("ruleset" by default for rules loaded from a header, or "document"), requires: ["anonymous-client-ip-when-cross-origin"] (cross-origin prefetch through a privacy proxy), expects_no_vary_search (tells the browser what No-Vary-Search header to expect so it can reuse an in-flight prefetch for URLs that differ only in ignored query parameters), tag (sent to your server in the Sec-Speculation-Tags request header; Chrome 136) and target_hint ("_blank" or "_self", for prerendering pages that open in a new tab; Chrome 138).
Inline rules need an explicit CSP allowance: script-src 'inline-speculation-rules' plus a hash or nonce for the rule set. Alternatively, serve the rules as a JSON file referenced by the Speculation-Rules response header.
Eagerness: when a speculation fires¶
Eagerness is a hint about how confident you are that the user will navigate. Chrome maps it to concrete triggers, which have changed over time and differ between desktop and mobile. According to Chrome's prerender documentation:
| Eagerness | Desktop trigger | Mobile trigger |
|---|---|---|
immediate | As soon as the rules are observed | Same |
eager | Pointer held over a link for 10 ms | Link enters the viewport, after 50 ms (since January 2026) |
moderate | Hover for 200 ms, or pointerdown, whichever comes first | Viewport heuristics: 500 ms after scrolling stops, for links close to the user's last tap (since August 2025) |
conservative | pointerdown / touch start | Same |
Chrome also limits how many speculations can be active at once, and replaces the oldest when you exceed the limit:
| Eagerness | Prefetches | Prerenders |
|---|---|---|
immediate | 50 | 10 |
eager, moderate, conservative | 2 (FIFO) | 2 (FIFO) |
Chrome skips speculation entirely when the user has enabled Save-Data, Energy Saver is active on low battery, memory is constrained, the "Preload pages" setting is off, or the page is in a background tab. Treat speculation as an optimization that may not happen.
Choosing eagerness. Prerendering costs the full page load: bandwidth, CPU, memory and server capacity, and it runs your analytics and any side-effecting JavaScript unless you defer them. A good default for most sites is prerender with moderate on same-origin links, plus prefetch with conservative or moderate for everything else, and immediate only for the one or two URLs the majority of users visit next (the step after a login form, the first result of a search).
Writing pages that are safe to prerender¶
A prerendered page loads in a hidden state. document.prerendering is true, and the prerenderingchange event fires when the user activates it. Chrome defers APIs that would be visible to the user or intrusive (permission prompts, notifications, media playback, and similar) until activation, but your own code needs to cooperate:
/**
* Run a callback only once the page is actually shown to the user. For normal
* loads it runs immediately; for prerendered pages it waits for activation.
*/
export function whenActivated(callback) {
if (document.prerendering) {
document.addEventListener("prerenderingchange", () => callback(), { once: true });
} else {
callback();
}
}
// Page views, impressions and "last seen" updates must not be recorded for
// pages the user never sees.
whenActivated(() => {
sendPageView(location.pathname);
});
// Time-based UI (countdown timers, "3 minutes ago" labels) should be refreshed
// at activation, because the page may have been rendered seconds earlier.
whenActivated(() => {
refreshRelativeTimestamps();
});
Server-side, Chrome marks speculative requests with Sec-Purpose: prefetch (and Sec-Purpose: prefetch;prerender for prerenders), which lets you exclude them from request-based analytics and rate-limit them. Measure prerendered loads from activationStart in the navigation entry, not from 0 (see Measuring Performance).
Speculation Rules and service workers¶
This is where PWAs historically lost out. Until recently, Chrome cancelled a speculation-rules prefetch as soon as it detected that the target URL was controlled by a service worker, and the Chrome team reported that service workers were the top actionable reason prefetches failed. Chrome 138 changed that: a prefetch to a URL controlled by a service worker now goes through the worker's fetch handler, and the response the worker produced is stored in the prefetch cache, so the later navigation is served from there (blink-dev announcement).
What that means for your worker:
- Your fetch handler runs for the prefetch, earlier than the navigation would have. A network-first navigation handler fetches from the network during the prefetch; a cache-first or app-shell handler answers from Cache Storage. Either way, the user's navigation then skips both the network and the worker start-up.
- Side effects run at prefetch time. If your worker logs navigations, updates caches or counts visits in its
fetchhandler, those effects happen for pages the user may never open. Keep the navigation handler free of side effects, or accept that its bookkeeping includes speculative requests. - Freshness is bounded by the speculation window. The prefetched response is used as-is when the user navigates. For content that must be fresh to the second, use
conservativeeagerness, which fires onpointerdown, a few hundred milliseconds before the click completes. - Prerendering is a real page load. A prerendered page runs in a hidden renderer and its navigation and subresource requests go through the controlling worker like any other load, so an app-shell worker serves the shell from cache, and the prerender mostly costs rendering time.
Speculation Rules and service-worker caching are complementary, not competing:
| Service worker precache | Speculation Rules prefetch | Speculation Rules prerender | |
|---|---|---|---|
| Lifetime | Until you delete it | The current page (discarded on the next navigation) | The current page |
| Works offline | Yes | No | No |
| Covers | Resources you list at build time | The next document | The next document and everything it loads and runs |
| Engines | All | Chromium | Chromium |
| First visit | Only after installation | Yes | Yes |
| Cost if unused | Storage and one-time bandwidth | Bandwidth, server load | Bandwidth, CPU, memory, server load |
A PWA can use precaching for the shell and static assets (so repeat visits and offline use are fast), and Speculation Rules for documents and data-dependent pages that cannot be precached (product pages, articles), especially on the first visit when the worker has not yet installed. See SPA vs MPA PWAs: speculation rules make the multi-page architecture much more competitive, and View Transitions make those navigations feel app-like.
Browser support for speculation rules¶
Speculation Rules are Chromium-only: Chrome and Edge 109 for list rules, 121 for document rules and eagerness, 136 for tag, 138 for target_hint and for prefetches through service workers. Safari 26.2 added a partial implementation (prefetch only, no prerender) behind a feature flag, and Firefox does not support the API. Browsers that do not understand <script type="speculationrules"> ignore it, so it is a safe progressive enhancement. Feature-detect with HTMLScriptElement.supports?.("speculationrules") if you insert rules from JavaScript.
Debugging speculations
In Chrome DevTools, Application > Background services > Speculative loads lists every rule set, each speculated URL, its status, and the reason a speculation failed or was not used. In the Network panel, speculative requests show Sec-Purpose in their request headers.
Code splitting, lazy loading and the PRPL pattern¶
The PRPL pattern, published by the Chrome team in 2016 for PWAs, is still a useful checklist:
- Push (today: preload) the critical resources for the initial route.
- Render the initial route as soon as possible.
- Pre-cache the remaining routes with the service worker.
- Lazy-load and create remaining routes on demand.
The original "push" referred to HTTP/2 server push, which is gone: Chrome removed support in Chrome 106 and Firefox in version 132, because pushed resources were rarely needed and often already cached. preload, modulepreload and 103 Early Hints do the same job without the downsides. The other three steps map directly onto modern tooling: route-based code splitting, a precache manifest generated at build time, and dynamic import().
Route-based splitting¶
The largest win is almost always to stop shipping every route's code on every page. With dynamic import(), bundlers (Vite, webpack, Rollup, esbuild) emit a separate chunk per route:
// Each route's code is its own chunk. Only the matched route is fetched.
const routes = [
{ pattern: new URLPattern({ pathname: "/" }), load: () => import("./views/home.js") },
{ pattern: new URLPattern({ pathname: "/products/:id" }), load: () => import("./views/product.js") },
{ pattern: new URLPattern({ pathname: "/cart" }), load: () => import("./views/cart.js") },
{ pattern: new URLPattern({ pathname: "/account/*" }), load: () => import("./views/account.js") },
];
const outlet = document.querySelector("#app");
export async function navigate(url) {
const match = routes.find((route) => route.pattern.test(url));
if (!match) {
return renderNotFound(outlet);
}
try {
const view = await loadWithRetry(match.load);
await view.render(outlet, match.pattern.exec(url));
} catch (error) {
// A chunk failed to load: offline, or (more often in a PWA) the chunk was
// deleted by a deployment. See "Chunk load failures after a deployment".
renderChunkError(outlet, error);
}
}
// One retry covers transient network errors in engines that re-fetch a failed
// module (Firefox 155+). Chromium caches the failure in the module map, so there
// the retry rejects immediately and renderChunkError() must offer a reload.
async function loadWithRetry(load, retries = 1) {
try {
return await load();
} catch (error) {
if (retries <= 0 || !navigator.onLine) throw error;
await new Promise((resolve) => setTimeout(resolve, 500));
return loadWithRetry(load, retries - 1);
}
}
URLPattern is available in Chrome 95, Firefox 142 and Safari 26; use a small regular-expression matcher if you support older engines.
The retry deserves a caveat. The HTML specification originally cached a failed module fetch in the document's module map, so a second import() of the same URL rejected without touching the network. The specification has since changed (whatwg/html#10327) so that network and HTTP errors are not cached; Firefox implements the new behavior from version 155 and Safari has it in Technology Preview, but Chromium still caches the failure as of September 2026. In Chromium the only reliable recovery from a failed chunk is a page reload (or a different URL), which is why the error path matters more than the retry.
Component-level and interaction-driven loading¶
Below the route level, three patterns defer code that the first paint does not need:
- Import on visibility. Load the code for a widget when its placeholder scrolls near the viewport.
- Import on interaction. Load a heavy dialog, editor or date picker when the user first hovers, focuses or taps its trigger.
- Import at idle. Warm up chunks the user will probably need once the page is idle.
// Import on visibility: the comments widget is below the fold on article pages.
const commentsSlot = document.querySelector("#comments");
if (commentsSlot) {
const io = new IntersectionObserver(
async (entries, observer) => {
if (!entries.some((entry) => entry.isIntersecting)) return;
observer.disconnect();
const { mountComments } = await import("./widgets/comments.js");
mountComments(commentsSlot);
},
{ rootMargin: "400px 0px" }, // start early enough that it is ready on arrival
);
io.observe(commentsSlot);
}
// Import on interaction: start loading on the first sign of intent (hover or
// focus), use it on click. The promise is cached so the import runs once.
const shareButton = document.querySelector("#share");
let shareModule;
const loadShare = () => (shareModule ??= import("./widgets/share-dialog.js"));
shareButton?.addEventListener("pointerenter", loadShare, { once: true });
shareButton?.addEventListener("focus", loadShare, { once: true });
shareButton?.addEventListener("click", async () => {
const { openShareDialog } = await loadShare();
openShareDialog();
});
// Import at idle: warm the cart chunk once the page has settled. requestIdleCallback
// is not available in Safari, so fall back to a timeout after load.
const idle = globalThis.requestIdleCallback ?? ((cb) => setTimeout(cb, 2000));
addEventListener("load", () => idle(() => import("./views/cart.js")), { once: true });
Interaction-driven loading has an INP cost: if the chunk is not ready when the user clicks, the click's visible response waits for the network. Starting the load on pointerenter or focus hides most of that on desktop; on touch devices there is no hover, so either precache the chunk or show immediate feedback (a spinner in the dialog frame) before awaiting the import. See Runtime Performance for keeping those handlers responsive.
Chunk load failures after a deployment¶
Code splitting creates a problem specific to long-lived pages and PWAs. A user opens the app, a new version is deployed, and later the user navigates to a route whose chunk has a new fingerprinted name. The old page asks for the old chunk name, which no longer exists on the server: the dynamic import rejects (webpack calls this a ChunkLoadError; Vite dispatches a vite:preloadError event on window).
A service worker can make this better or worse:
- Better: if your precache contains every chunk of the version the page is running, old pages keep working until they reload, because the worker serves old chunks from the old cache. This only holds if the new worker does not delete the old caches while old pages are still open. The default lifecycle (the new worker waits until all old clients close) protects you; calling
skipWaiting()and deleting old caches inactivateremoves that protection. See Updating Service Workers. - Worse: if the worker serves an HTML shell with old chunk URLs from its cache after the server has deleted them, and the chunks are not precached, the app is broken until the worker updates.
Two defensive measures work regardless of the setup: keep previous deployments' assets on the server or CDN for at least as long as sessions typically last, and handle import failures by reloading the page once:
// Vite fires this event when a dynamic import's preload fails. For webpack,
// catch errors whose name is "ChunkLoadError" in your route loader instead.
addEventListener("vite:preloadError", (event) => {
// Avoid reload loops: reload at most once per ten minutes per tab.
const key = "chunk-reload-at";
const last = Number(sessionStorage.getItem(key) ?? 0);
if (Date.now() - last > 10 * 60 * 1000) {
sessionStorage.setItem(key, String(Date.now()));
event.preventDefault(); // stop Vite from rethrowing
location.reload();
}
});
How far to split¶
Splitting has diminishing returns. Every chunk is a request with its own headers, its own compression context (a small file compresses worse than the same bytes inside a large one), and its own place in the dependency graph. HTTP/2 and HTTP/3 made many parallel requests cheap, but not free. In practice:
- Split per route and per heavy, rarely used feature (editors, charts, maps, admin tools).
- Keep a shared vendor chunk for libraries used on every route, so that the long-lived, rarely changing part of your code caches well across deployments.
- Avoid hundreds of tiny modules in production. Some teams ship unbundled ES modules in development and bundle for production for this reason.
- Precache the chunks of the most common routes during service worker installation; runtime-cache the rest on first use with a cache-first strategy (fingerprinted URLs never change, so cache-first is always correct for them). See Caching Strategies.
Compression: gzip, Brotli, zstd and dictionaries¶
Text resources (HTML, CSS, JavaScript, JSON, SVG) compress by 60–90%, and compression is usually the single cheapest loading win. The browser advertises what it can decode in Accept-Encoding; the server picks one and labels the response with Content-Encoding, and must send Vary: Accept-Encoding so that caches keep the variants apart.
| Encoding | Token | Chrome / Edge | Firefox | Safari | Best for |
|---|---|---|---|---|---|
| gzip | gzip | ✅ | ✅ | ✅ | Universal fallback |
| Brotli (RFC 7932) | br | ✅ 50 | ✅ 44 | ✅ 11 | Precompressed static assets |
| Zstandard (RFC 8878) | zstd | ✅ 123 | ✅ 126 | ⚠️ 26.3 | Dynamic responses compressed per request |
| Dictionary-compressed Brotli | dcb | ✅ 130 | ❌ | ❌ | Repeat downloads of changed files |
| Dictionary-compressed Zstandard | dcz | ✅ 130 | ❌ | ❌ | Repeat downloads of changed files |
Support data as of September 2026. See caniuse: zstd and caniuse: Brotli for live data.
⚠️ WebKit's Safari 26.3 release notes say zstd is available on iOS 26.3, iPadOS 26.3, visionOS 26.3 and macOS Tahoe 26.3, not in Safari 26.3 on earlier macOS versions, because Safari relies on the operating system's networking stack. Browsers only advertise encodings they can decode, so serving zstd when it is listed in Accept-Encoding is always safe.
Chromium and Firefox advertise br only on HTTPS connections, which for a PWA is a given.
Brotli versus zstd¶
The two modern formats have different sweet spots:
- Brotli at its maximum quality (11) produces the smallest files for web text, thanks to a built-in static dictionary of common web strings, but it is slow to compress. It is ideal for precompressing at build time: compress each fingerprinted asset once, serve it millions of times.
- zstd compresses much faster at comparable ratios at medium levels, and decompresses very fast. It is ideal for dynamic responses (server-rendered HTML, API JSON) compressed on every request, where Brotli at high levels costs too much server CPU and Brotli at low levels loses its size advantage.
For HTTP, zstd has an extra constraint: RFC 9659 limits the decoder window to 8 MB for the zstd content coding, so encoders must not use larger windows (which the highest compression levels would otherwise do). Standard libraries and servers apply this limit, but check if you tune levels yourself.
A build step that precompresses every text asset in both formats lets the server pick the best variant per request without compressing anything at request time:
// Precompress build output with Brotli (quality 11) and, where the runtime supports
// it, zstd. Run after the bundler, before deployment. Node's zlib added zstd in
// recent release lines, so feature-detect it.
import { readdir, readFile, writeFile, stat } from "node:fs/promises";
import { join, extname } from "node:path";
import zlib from "node:zlib";
const ROOT = process.argv[2] ?? "dist";
const TEXT_TYPES = new Set([".html", ".css", ".js", ".mjs", ".json", ".svg", ".xml", ".txt", ".webmanifest", ".wasm"]);
const MIN_SIZE = 1024; // tiny files gain nothing and cost a request header
const hasZstd = typeof zlib.zstdCompressSync === "function";
async function* walk(dir) {
for (const entry of await readdir(dir, { withFileTypes: true })) {
const path = join(dir, entry.name);
if (entry.isDirectory()) yield* walk(path);
else yield path;
}
}
let saved = 0;
for await (const file of walk(ROOT)) {
if (!TEXT_TYPES.has(extname(file))) continue;
const { size } = await stat(file);
if (size < MIN_SIZE) continue;
const input = await readFile(file);
const br = zlib.brotliCompressSync(input, {
params: {
[zlib.constants.BROTLI_PARAM_QUALITY]: 11,
[zlib.constants.BROTLI_PARAM_SIZE_HINT]: input.length,
[zlib.constants.BROTLI_PARAM_MODE]: extname(file) === ".wasm"
? zlib.constants.BROTLI_MODE_GENERIC
: zlib.constants.BROTLI_MODE_TEXT,
},
});
// Only keep a variant if it is actually smaller.
if (br.length < input.length) {
await writeFile(`${file}.br`, br);
saved += input.length - br.length;
}
if (hasZstd) {
const zst = zlib.zstdCompressSync(input, {
params: { [zlib.constants.ZSTD_c_compressionLevel]: 19 },
});
if (zst.length < input.length) await writeFile(`${file}.zst`, zst);
}
const gz = zlib.gzipSync(input, { level: 9 });
if (gz.length < input.length) await writeFile(`${file}.gz`, gz);
}
console.log(`Brotli saved ${(saved / 1024).toFixed(1)} KiB across ${ROOT}`);
Your server or CDN then maps Accept-Encoding to the precompressed file (for example nginx's gzip_static, the brotli_static directive of the ngx_brotli module, or your CDN's "serve precompressed" option) and must still send Vary: Accept-Encoding.
Compression and service worker caches¶
What the service worker sees is the decoded body: the Fetch API transparently decompresses response bodies, while the Content-Encoding header remains in response.headers for information. When your worker writes a response into Cache Storage, it stores the decoded payload. Two consequences:
- Quota accounting uses uncompressed sizes. A 300 KB JavaScript bundle that transfers as 80 KB of Brotli occupies roughly 300 KB (plus overhead) of your origin's storage quota. Budget your precache with that in mind; see Storage Quotas & Persistence.
- Cache hits skip decompression. On a repeat visit, a cached response needs no decompression work, which slightly favors precaching large text assets on low-end devices.
Whether compression helps the first load of those assets is unaffected by the worker: the precache step downloads each asset through the network, compressed like any other request.
Compression dictionary transport¶
Most PWA deployments change a few lines in a large bundle and force every returning user to download the whole file again. Compression Dictionary Transport (RFC 9842, published September 2025; shipped in Chrome 130) lets the server compress the new version using the old version as a dictionary, so only the differences cross the wire:
- The server marks a response as usable as a dictionary with
Use-As-Dictionary: match="/js/app.*.js"(thematchvalue is a URL pattern). - On the next request for a matching URL, Chrome sends
Available-Dictionary: :<sha-256 of the dictionary>:and addsdcbanddcztoAccept-Encoding. - If the server has that dictionary, it responds with
Content-Encoding: dcb(Brotli with a dictionary) ordcz(zstd with a dictionary) and a delta that is often a few percent of the full file.
Chrome's team reported that rolling dictionaries out to Google Search reduced the average HTML payload by 23% compared to standard Brotli (source). For PWAs the fit is natural: fingerprinted bundles are exactly the "old version is a good dictionary for the new version" case. The catch is operational: your CDN or server must keep previous versions, compute deltas, and key its cache on Available-Dictionary. Note also that a service worker precache update fetches the new files from the network through the worker's own fetch() calls, which use the HTTP cache like any other request, so dictionaries can reduce precache update traffic too.
HTTP/2 and HTTP/3¶
Every PWA should be served over HTTP/2 at least, and HTTP/3 where the CDN supports it. The practical differences for loading performance:
| HTTP/1.1 | HTTP/2 | HTTP/3 | |
|---|---|---|---|
| Transport | TCP + TLS | TCP + TLS | QUIC (UDP) with TLS 1.3 built in |
| Requests per connection | One at a time (six connections per host) | Many, multiplexed | Many, multiplexed |
| Head-of-line blocking | Per connection | At the TCP layer (one lost packet stalls every stream) | Per stream only |
| Connection setup | TCP + TLS (2–3 RTT) | TCP + TLS (2–3 RTT) | 1 RTT, 0-RTT on resumption |
| Connection migration (Wi-Fi to cellular) | No | No | Yes (connection IDs) |
| Prioritization | None | RFC 9218 Priority header (the RFC 7540 tree is deprecated) | RFC 9218 Priority header and frames |
| Browser support | All | All | Chrome 87, Firefox 88, Safari 16 |
Support data as of September 2026; see caniuse: HTTP/3.
What this means in practice:
- Domain sharding is an anti-pattern. Splitting assets across
static1.andstatic2.hostnames to get more HTTP/1.1 connections now costs extra handshakes and defeats prioritization. Consolidate onto as few origins as possible. - Mobile networks favor HTTP/3. QUIC's per-stream loss recovery and faster handshakes help most on lossy, high-latency links, which is where PWAs are often used. Browsers discover HTTP/3 through an
Alt-Svcresponse header (so the first connection is usually HTTP/2) or through DNSHTTPSrecords, which let the browser use HTTP/3 from the first request. - Priorities need server support. The browser sends priorities (
Priority: u=0, istyle headers or frames under RFC 9218), but it is the server that decides what to send first. Some CDNs and servers ignore them. If your LCP image is slow despitefetchpriority="high", check the server. - Bundling still matters, but less. HTTP/2 made per-request overhead small, not zero. Header compression (HPACK, QPACK) removes most of the per-request header cost; compression efficiency and the number of dependency waves remain reasons to bundle sensibly.
nextHopProtocol in Resource Timing tells you which protocol each request used (h2, h3, http/1.1). For responses served by your service worker from Cache Storage it is an empty string, which is itself useful for classifying cache hits; see Measuring Performance.
Images¶
Images are the LCP element on most pages and usually the largest share of bytes. Loading them well is a combination of format, size, priority and timing.
Formats¶
| Format | Chrome / Edge | Firefox | Safari | Notes |
|---|---|---|---|---|
| WebP | ✅ | ✅ | ✅ 14 (macOS Big Sur and later) | Safe default for photos and graphics with alpha |
| AVIF | ✅ Chrome 85, Edge 121 | ✅ 93 | ✅ 16.4 (partial 16.1–16.3) | Usually 20–50% smaller than WebP at similar quality; slower to encode |
| JPEG XL | 🧪 Rust decoder behind a flag since Chrome 145; on by default from Chrome 155 | 🧪 intent to ship announced | ⚠️ 17 | Only Safari decodes it by default in stable releases as of September 2026 |
Support data as of September 2026; see caniuse: AVIF and caniuse: JPEG XL. ⚠️ Safari does not decode animated JPEG XL, so caniuse lists its support as partial. Chrome 155, which enables JPEG XL by default, was in beta in September 2026, and Mozilla announced its intent to ship JPEG XL in August 2026. Until those releases reach most of your users, serve JPEG XL only through <picture> or Accept negotiation with an AVIF, WebP or JPEG fallback.
Serve modern formats with <picture> (the browser picks the first <source> whose type it supports) or with server-side content negotiation on the Accept header, which image CDNs do automatically:
<picture>
<source type="image/avif"
srcset="/img/p42-400.avif 400w, /img/p42-800.avif 800w, /img/p42-1200.avif 1200w"
sizes="(min-width: 960px) 600px, 100vw">
<source type="image/webp"
srcset="/img/p42-400.webp 400w, /img/p42-800.webp 800w, /img/p42-1200.webp 1200w"
sizes="(min-width: 960px) 600px, 100vw">
<img src="/img/p42-800.jpg" width="1200" height="900" alt="Espresso machine, front view"
fetchpriority="high" decoding="async">
</picture>
Content negotiation and service workers need care. A response that varies by Accept must carry Vary: Accept, and cache.match() honors Vary: if the page later requests the same URL with a different Accept header, the match fails. Browsers send stable Accept headers for images, so in practice this works, but a worker that precaches images using fetch(url) from the worker (whose Accept is */*) stores a variant that later image requests will not match, and the CDN may have served it a JPEG. Either precache with explicit <picture>-style URLs per format, or construct precache requests with the same Accept header the browser uses for images.
Responsive sizing: srcset and sizes¶
srcset with width descriptors plus sizes lets the browser download the smallest candidate that is sharp on the current screen. The browser computes sizes × device pixel ratio and picks the closest candidate at or above it. Rules of thumb:
- Provide candidates roughly every 1.5× in width, from the smallest layout slot to about 2× the largest.
- Make
sizesmatch your CSS.sizes="100vw"on an image that is 600 CSS pixels wide on desktop makes a 2× laptop download a 2,880-pixel image. - For lazy-loaded images,
sizes="auto"lets the browser use the actual layout width (Chrome 126, Firefox 150, Safari 27; it requiresloading="lazy"). List it first with a fallback:sizes="auto, (min-width: 960px) 600px, 100vw". - Always set
widthandheight(or CSSaspect-ratio) so the browser reserves space before the image loads; otherwise the image causes layout shifts. See Core Web Vitals.
Lazy loading and the LCP image¶
Native loading="lazy" defers off-screen images and iframes until they approach the viewport (images in Chrome 77, Firefox 75, Safari 15.4; iframes in Safari 16.4 and Firefox 121). It is the right default for everything below the fold and the wrong setting for the LCP image: a lazy image is not requested until layout confirms it is near the viewport, which delays it behind CSS and JavaScript. The same goes for JavaScript lazy-loading libraries that swap data-src for src, which also hide the image from the preload scanner.
A reliable pattern for templates:
<!-- First card is likely LCP: eager, high priority, sync decoding is fine. -->
<img src="/img/card-1.avif" width="400" height="300" alt="…" fetchpriority="high">
<!-- The next few may be in the first viewport on large screens: leave them eager
at default priority. -->
<img src="/img/card-2.avif" width="400" height="300" alt="…">
<!-- Everything further down: lazy. -->
<img src="/img/card-9.avif" width="400" height="300" alt="…" loading="lazy" decoding="async">
decoding="async" lets the browser decode off the critical path and present the image slightly later; it helps below-the-fold images and large images inside animations and is unnecessary on the LCP image.
Images in the service worker¶
Images are usually the largest runtime cache in a PWA. Use a cache-first strategy with an expiration policy (maximum entries, maximum age) so that the cache does not grow without bound; Workbox's ExpirationPlugin does this for you (see Workbox Fundamentals). Do not precache images the user may never see; precache only the shell's own icons and illustrations. Be careful with cross-origin images without CORS: they produce opaque responses, which Chromium pads to a large size for quota purposes, so a few hundred opaque images can exhaust a small quota. Serve images with CORS headers and request them with crossorigin if you intend to cache them. See Storage Quotas & Persistence.
Fonts¶
Web fonts affect loading twice: they are late-discovered resources (the browser only knows it needs a font after CSS is parsed and text using it is laid out), and they block or change the rendering of text.
Controlling rendering with font-display¶
The font-display descriptor in @font-face decides what happens while the font loads. Each value is defined by a block period (text is rendered invisibly) and a swap period (text is rendered in a fallback and swapped when the font arrives):
| Value | Block period | Swap period | Result |
|---|---|---|---|
auto | Browser default (usually like block) | ||
block | Short (about 3 s) | Infinite | Invisible text, then the font. Use only for icon fonts |
swap | Extremely small (about 100 ms at most) | Infinite | Fallback immediately, swap whenever the font arrives; can cause layout shift late |
fallback | Extremely small | Short (about 3 s) | Swap only if the font arrives quickly, otherwise keep the fallback for this page |
optional | Extremely small | None | Use the font only if it is available almost immediately (cached); the browser may skip downloading it on slow connections |
For a PWA, optional is attractive: on the first visit text renders immediately in the fallback, the font downloads (and your service worker caches it), and on every later visit the font is available instantly from cache and used from the first frame, with no layout shift at all. swap is the right choice when the brand font must appear on the first visit.
Reducing font layout shift with metric overrides¶
When a fallback font is swapped for the web font, text reflows because the fonts have different metrics. @font-face descriptors size-adjust, ascent-override, descent-override and line-gap-override let you define a fallback face whose metrics match the web font:
@font-face {
font-family: "Inter";
src: url("/fonts/inter-var-latin.woff2") format("woff2");
font-weight: 100 900; /* variable font: one file for every weight */
font-display: swap;
unicode-range: U+0000-00FF, U+0131, U+0152-0153, U+02BB-02BC, U+02C6, U+02DA,
U+02DC, U+2000-206F, U+2074, U+20AC, U+2122, U+2191, U+2193,
U+2212, U+2215, U+FEFF, U+FFFD;
}
/* A local fallback tuned to Inter's metrics, so the swap barely moves text.
Tools such as Fontaine or Capsize compute these numbers from the font files. */
@font-face {
font-family: "Inter Fallback";
src: local("Arial");
size-adjust: 107%;
ascent-override: 90%;
descent-override: 22%;
line-gap-override: 0%;
}
body {
font-family: "Inter", "Inter Fallback", system-ui, sans-serif;
}
The numbers above are illustrative; compute them for your actual font and fallback pair. local() fallbacks depend on the fonts installed on the user's system, so pick fallbacks available on each target platform or list several.
Font loading checklist¶
- WOFF2 only. Every browser that runs a PWA supports it; drop WOFF, TTF and EOT.
- Subset fonts with
unicode-rangeso users download only the scripts they need, and self-host rather than relying on third-party font services: third-party font caches are partitioned per site in all modern browsers, so the "shared cache" argument no longer applies, and a self-hosted font avoids a connection to another origin. - Preload the one or two fonts used above the fold, with
crossorigin(always required for fonts), and nothing else. - Use variable fonts when you need several weights: one file instead of four.
- Precache fonts in the service worker. They are small, fingerprintable and used on every page, which makes them ideal precache candidates, and caching them makes
font-display: optionalwork well.
Performance budgets for bundles and assets¶
A budget turns "keep it fast" into a number that a build can fail. Budgets come in three flavors: quantity (bytes of JavaScript, number of requests), milestone timings (LCP, TBT in the lab), and rule-based (a Lighthouse score). Quantity budgets are the easiest to enforce at build time and the best early warning, because bytes of JavaScript correlate with parse, compile and execution cost on the device.
There is no universal correct budget. A useful way to derive one is to start from the target: for example, "LCP under 2.5 s at the 75th percentile on a mid-range Android phone over a slow 4G connection" leaves a limited number of kilobytes for the critical path once you subtract the round trips for DNS, connection and the HTML. Then assign budgets per resource type and per route, and measure. The table below is an example starting point for a PWA's initial route, not a standard:
| Resource (compressed, initial route) | Example budget | Why |
|---|---|---|
| HTML (including inline critical CSS) | ≤ 30 KB | First round trip carries the head and above-the-fold markup |
| CSS (render-blocking) | ≤ 50 KB | Blocks first paint |
| JavaScript (initial, all chunks) | ≤ 150 KB | Parse, compile and execution on low-end devices |
| Fonts | ≤ 2 files, ≤ 100 KB | Late-discovered; each file is a request on the critical path |
| LCP image | ≤ 150 KB | Directly in the LCP resource load duration |
| Service worker precache (uncompressed) | Set explicitly, for example ≤ 2 MB | Downloaded during the first visit; counts against quota |
The last row is PWA-specific and often forgotten: a precache manifest that grows to 20 MB makes every first visit and every deployment expensive, competes with the first page load for bandwidth, and fills storage on low-end devices.
Enforcing budgets in the build¶
Bundlers have built-in warnings: webpack's performance.maxAssetSize and performance.maxEntrypointSize (250,000 bytes by default, with performance.hints set to "warning" or "error"), and Vite's build.chunkSizeWarningLimit (500 kB by default, a warning only). A dedicated tool such as size-limit fails CI on compressed sizes:
[
{
"name": "Initial JS (home route)",
"path": ["dist/assets/index-*.js", "dist/assets/vendor-*.js"],
"limit": "150 kB",
"brotli": true
},
{
"name": "Critical CSS",
"path": "dist/assets/index-*.css",
"limit": "50 kB",
"brotli": true
}
]
For the precache budget, check the manifest your tooling generates. Workbox's build tools report the total size and count of precached files and accept maximumFileSizeToCacheInBytes (2 MB per file by default) to stop individual large files from being precached; add a separate check on the total:
// Fails the build if the service worker's precache grows past the budget.
// Assumes a Workbox-style manifest injected into dist/sw.js as self.__WB_MANIFEST
// entries; adjust the regex for your tooling.
import { readFile, stat } from "node:fs/promises";
import { join } from "node:path";
const DIST = "dist";
const BUDGET_BYTES = 2 * 1024 * 1024;
const sw = await readFile(join(DIST, "sw.js"), "utf8");
const urls = [...sw.matchAll(/\{"?url"?:\s*"([^"]+)"/g)].map((match) => match[1]);
let total = 0;
for (const url of urls) {
try {
total += (await stat(join(DIST, url))).size;
} catch {
console.warn(`Precache entry not found in ${DIST}: ${url}`);
}
}
console.log(`Precache: ${urls.length} files, ${(total / 1024).toFixed(0)} KiB (uncompressed)`);
if (total > BUDGET_BYTES) {
console.error(`Precache budget exceeded by ${((total - BUDGET_BYTES) / 1024).toFixed(0)} KiB`);
process.exit(1);
}
Timing and score budgets belong in lab runs on every pull request, with Lighthouse CI or WebPageTest; that set-up is covered in Measuring Performance.
Loading a PWA: first visit versus repeat visit¶
Everything above applies to any website. What makes a PWA different is that it has two quite different loading paths, and the service worker's installation happens during the first one.
First visit: keep the worker out of the way¶
On the first visit there is no worker. The page loads over the network, and then your registration code runs, the worker downloads, and its install handler precaches the shell and assets. If registration happens early, precaching competes for bandwidth and CPU with the page's own critical resources, and first-visit LCP suffers. Register after the load event:
if ("serviceWorker" in navigator) {
// Wait for the page's own resources before starting precache downloads.
addEventListener("load", () => {
navigator.serviceWorker
.register("/sw.js", { scope: "/", updateViaCache: "none" })
.catch((error) => console.error("Service worker registration failed:", error));
});
}
Other first-visit levers:
- Keep the precache small (see the budget above), and runtime-cache the rest.
- Let precaching reuse what the page already downloaded. Assets the page loaded are usually in the HTTP cache. Whether the worker's precache requests can be satisfied from it depends on each request's
cachemode and on how your tooling busts caches for unversioned URLs, so check the Network panel during installation: fingerprinted files with a longmax-ageshould not be downloaded twice. - Use Speculation Rules for the next navigation. They work before the worker exists.
Repeat visit: worker start-up and cache reads¶
On a repeat visit the question becomes how quickly the worker can answer. A service worker that is not running must be started before it can handle the navigation (tens to hundreds of milliseconds on mobile devices). Mitigations, all covered in depth on their own pages:
- Serve the shell from Cache Storage for navigations, so that the only cost is worker start-up plus a cache read (App Shell Model).
- Enable navigation preload for network-first navigations, so the network request runs in parallel with worker start-up (Navigation Preload).
- Use static routing to send requests that the worker would only pass through straight to the network or cache, without starting the worker (Static Routing API).
- Do not add a pass-through
fetchhandler. A handler that only callsfetch(event.request)makes every request slower. Chromium skips the worker entirely for fetches when allfetchlisteners are no-ops and warns in the console, but a handler that does anything at all prevents the optimization. See Handling Fetch Events.
Stale code and loading¶
A PWA served from cache can also be too fast in the wrong way: it loads yesterday's code. Loading performance and update strategy interact: an app that checks for an update on every launch and reloads when it finds one pays a second full load. Prefer updating in the background and applying the new version on the next navigation or with a user prompt. See Updating Service Workers.
Browser support¶
Support data as of September 2026. For live data, check MDN and caniuse links in each section, and caniuse.com in general.
| Feature | Chrome / Edge | Firefox | Safari (macOS, iOS) |
|---|---|---|---|
fetchpriority attribute | ✅ 102 | ✅ 132 | ✅ 17.2 |
<link rel=preload> | ✅ | ✅ | ✅ |
<link rel=modulepreload> | ✅ 66 | ✅ 115 | ✅ 17 |
<link rel=preconnect> / dns-prefetch | ✅ | ✅ | ✅ |
103 Early Hints (preconnect) | ✅ 103 | ✅ 120 | ✅ 17 |
103 Early Hints (preload) | ✅ 103 | ✅ 123 | ❌ |
| Speculation Rules (list rules) | ✅ 109 | ❌ | 🧪 |
| Speculation Rules (document rules, eagerness) | ✅ 121 | ❌ | 🧪 |
| Speculation Rules prefetch through service workers | ✅ 138 | ❌ | ❌ |
Brotli (br) | ✅ 50 | ✅ 44 | ✅ 11 |
Zstandard (zstd) | ✅ 123 | ✅ 126 | ⚠️ 26.3 |
Compression dictionaries (dcb, dcz) | ✅ 130 | ❌ | ❌ |
| HTTP/3 | ✅ 87 | ✅ 88 | ✅ 16 |
| AVIF | ✅ 85 (Edge 121) | ✅ 93 | ✅ 16.4 |
loading="lazy" on images | ✅ 77 | ✅ 75 | ✅ 15.4 |
sizes="auto" | ✅ 126 | ✅ 150 | ✅ 27 |
blocking="render" | ✅ 105 | ❌ | ✅ 18.2 |
⚠️ zstd in Safari 26.3 requires macOS Tahoe 26.3, iOS 26.3, iPadOS 26.3 or visionOS 26.3; Safari 26.3 on older macOS versions does not advertise it.
Common pitfalls¶
- Lazy-loading the LCP image.
loading="lazy"or a JavaScript lazy loader on the hero image delays it until after layout. Keep the first image eager and give itfetchpriority="high". - Preloads that do not match. A preload without
crossoriginfor a font, oras=fetchpreloads whose laterfetch()uses different credentials, download twice. Watch for Chromium's "preloaded but not used" console warning. - Too many high-priority requests. Five preloads and three
fetchpriority="high"images mean nothing is high priority. Reserve both for the critical path. - Registering the service worker before
load. Precaching competes with the first page load. - A precache that grows without limit. Every deployment re-downloads changed files for every user; large precaches also count against storage quotas at their uncompressed size.
- Deleting old chunks immediately on deployment. Open tabs and installed PWAs running the previous version fail to load routes. Keep old assets available, and handle chunk load errors.
- Prerendering pages with side effects. Page views, "mark as read", or cart updates triggered on load run for pages the user never opens. Gate them on
document.prerenderingandprerenderingchange. - Assuming Early Hints help controlled navigations. A navigation served from Cache Storage never reaches the server.
- Serving
Content-EncodingwithoutVary: Accept-Encoding, or image content negotiation withoutVary: Accept. Shared caches then serve the wrong variant; service worker caches then miss. - Domain sharding and many third-party origins. Each origin costs a connection; each connection costs handshakes that
preconnectcan only partly hide.
Debugging¶
- Network panel (Chrome DevTools). Enable the Priority column (right-click the header) to see initial and final priorities; the Initiator column shows whether a request came from the parser, a preload, Early Hints or a script. Requests served by a service worker show "(ServiceWorker)" in the Size column and have a gear icon for requests the worker itself made.
- Performance panel. The Network track shows render-blocking requests with a marker, and the Insights sidebar flags render-blocking resources, LCP discovery problems (lazy-loaded or late-discovered LCP images), missing compression and HTTP/1.1 usage. See Browser DevTools.
- Application panel. Background services > Speculative loads for Speculation Rules; Service workers for bypassing the worker (the Bypass for network checkbox) when you want to see pure network behavior.
- Response headers. Check
Content-Encoding,Vary,Cache-ControlandLinkon real responses, from the CDN, not your origin server.curl -sI -H "Accept-Encoding: zstd, br, gzip" https://example.com/shows which encoding is chosen. - Resource Timing in the console.
performance.getEntriesByType("resource")exposesnextHopProtocol,renderBlockingStatus,transferSize,encodedBodySize,decodedBodySizeanddeliveryTypefor each request; see Measuring Performance for how to interpret them for worker-served responses. - WebPageTest. Its waterfall shows connection reuse, HTTP/3 usage, Early Hints and priorities per request, on real devices and network conditions.
Further reading¶
On this site
- App Shell Model: precaching a shell and the LCP trade-offs of client rendering
- Core Web Vitals: LCP subparts, TTFB with service workers and Early Hints
- Runtime Performance: what happens after the bytes arrive
- Measuring Performance: lab and field set-ups that separate first and repeat visits
- Navigation Preload and Static Routing API: removing worker start-up from the critical path
- Precaching & Runtime Caching and HTTP Caching & Service Workers
- SPA vs MPA PWAs: how Speculation Rules change the architecture choice
External references
- web.dev: Optimize resource loading with the Fetch Priority API
- MDN: Speculation Rules API
- Chrome for Developers: Prerender pages in Chrome for instant page navigations
- Chrome for Developers: Faster page loads using server think-time with Early Hints
- WebKit: WebKit features for Safari 26.3 (zstd)
- RFC 9842: Compression Dictionary Transport
- RFC 9218: Extensible Prioritization Scheme for HTTP
- MDN: Responsive images
- MDN: font-display