Streaming Responses¶
A streaming response is one whose body the browser can start using before the last byte has arrived. Browsers parse and render HTML incrementally, and a service worker can build a response from a ReadableStream instead of a finished Blob, so a worker can send a cached page header immediately, pipe the page's content from the network as it arrives and finish with a cached footer. The result is a multi-page app whose navigations paint from the cache in milliseconds while the fresh content streams in. This page covers the Streams API as it applies to service workers, a production-grade stream concatenation helper, a complete streaming service worker, workbox-streams, the headers and status rules that streaming imposes, the server changes it needs, and how to handle errors after the first byte has been sent.
Key takeaways
- A
Responsebody can be anyReadableStreamofUint8Arraychunks.new Response(stream, init)returns immediately; the browser starts parsing as soon as the first chunk is enqueued. - Streaming composition (cached header + network content + cached footer) gives an MPA app-shell-like first paint without becoming an SPA. It needs a server that can return content-only fragments and a precached, versioned shell.
- The status code and headers are fixed at the first byte. Decide redirects, errors and the page
<title>before committing, or accept a fixed200and handle them in-band. - A synthesized response carries only the headers you give it: copy the document's
Content-Security-Policy, COOP/COEP and other security headers from the network response, and never copyContent-LengthorContent-Encoding. - Keep the worker alive with
event.waitUntil()until the stream finishes, propagate cancellation to the network, and turn mid-stream failures into valid closing markup instead of erroring the stream. - Intermediaries break streaming silently: reverse-proxy buffering, compression middleware and anything that calls
response.text()in the worker all turn a stream back into a buffer. - Every engine supports what composition needs:
ReadableStream,TransformStream,pipeTo(),TextEncoderStreamand navigation preload. Newer conveniences such as async iteration andReadableStream.from()are not yet universal.
Why streaming makes navigations faster¶
A browser does not wait for a whole HTML document. The parser consumes bytes as they arrive, the preload scanner discovers stylesheets, scripts and images in the first chunks and starts fetching them, and the renderer can paint as soon as the render-blocking resources in the <head> are available. A document that sends its <head> and page frame in the first few kilobytes gets its CSS requests started and its first paint scheduled long before the server has finished the slow part of the page.
A service worker can either preserve or destroy that behavior:
| What the worker does | Effect on streaming |
|---|---|
No fetch handler, or does not call respondWith() | Browser handles the navigation normally; streams from the network |
respondWith(fetch(request)) or respondWith(event.preloadResponse) | Body is passed through as a stream; streaming preserved |
respondWith(caches.match(request)) | Body streams from disk; very fast, but content may be stale |
Awaits response.text(), edits the string, returns new Response(text) | Whole body buffered in the worker; first byte delayed until the last network byte |
Returns a ReadableStream composed of several sources | Streams, and can mix cached and network bytes in one document |
The last row is what this page is about. It combines the two fastest sources: the part of the page that rarely changes (the <head> with CSS, the header, the navigation) comes from the Cache API with essentially no latency, and the part that changes (the article, the product, the inbox) comes from the network, streamed through as it arrives.
sequenceDiagram
participant B as Browser
participant SW as Service worker
participant C as Cache Storage
participant N as Network
B->>SW: navigation request (preload starts in parallel)
SW->>C: match /partials/head.html
SW->>N: content fragment (navigation preload)
C-->>SW: head partial
N-->>SW: fragment headers (status, title)
SW-->>B: head bytes: parser starts, CSS requested
N-->>SW: fragment body chunks
SW-->>B: content chunks as they arrive
SW->>C: match /partials/foot.html
SW-->>B: footer bytes, close Compared with an SPA app shell, streaming composition keeps real documents: every URL returns complete, crawlable HTML from the server on the first visit, the browser's back/forward cache works normally, and no client-side router or hydration is needed to show content. Compared with network-first navigations, repeat visits start painting from the cache instead of waiting for the server's first byte. App Shell Model and SPA vs MPA PWAs compare the architectures, and the Architecture overview shows which rendering strategies pair with which navigation strategy.
Jake Archibald's 2016 article Stream your way to immediate responses introduced the technique, and the Chrome team's Faster multipage applications with streams describes it for Workbox users, including the caveats about per-page <title> and navigation state covered below.
The Streams API as a service worker uses it¶
The Streams Standard defines three stream types and the machinery to connect them. A service worker needs only part of it, but it needs that part exactly.
ReadableStream: sources, controllers and readers¶
const stream = new ReadableStream(
{
start(controller) {
// Called once, synchronously during construction. May return a promise.
},
async pull(controller) {
// Called whenever the internal queue is below the high-water mark and the
// previous pull() has settled. Enqueue at least one chunk, close, or error.
controller.enqueue(new TextEncoder().encode("<p>chunk</p>"));
},
cancel(reason) {
// The consumer is no longer interested (tab closed, navigation aborted).
// Release resources and cancel upstream work here.
},
},
{ highWaterMark: 1 } // queuing strategy: counts chunks by default
);
The controller (ReadableStreamDefaultController) exposes enqueue(chunk), close(), error(reason) and desiredSize (high-water mark minus queued size; zero or negative means "stop producing"). A stream is consumed through a reader from getReader(), whose read() returns { value, done }; while a reader holds the lock, stream.locked is true and no one else may read it. reader.releaseLock() gives the stream back, and reader.cancel(reason) cancels it.
Rules that matter for responses:
- Response bodies are byte streams. Every chunk you enqueue into a stream used as a
Responsebody must be aUint8Array. Enqueuing a string is accepted by the stream but makes the fetch machinery fail with aTypeErrorwhen it reads that chunk, which breaks the response mid-flight. Encode text withTextEncoderor pipe throughTextEncoderStream. - A body can be read once.
response.bodyis locked aftergetReader(),pipeTo()ortext(). Useresponse.clone()(which tees the stream) when you also need to cache it. cancel()is your cleanup hook. When the user navigates away mid-load, the browser cancels the response body. If your source is reading from a network response, cancel that too, or the download continues in the worker for nothing.
TransformStream: modifying a stream in flight¶
A TransformStream is a pair of streams, writable and readable, with a transformer in between. Whatever is written to one end comes out of the other, possibly changed.
const upper = new TransformStream({
start(controller) {}, // optional
transform(chunk, controller) { // called per written chunk
controller.enqueue(chunk.toUpperCase());
},
flush(controller) {}, // called after the writable side closes
});
controller.enqueue(), controller.error() and controller.terminate() are available in the transformer. With no transformer at all, new TransformStream() is an identity stream: a convenient way to get a readable you can hand to new Response() and a writable you can fill later, from several sources in sequence.
pipeTo, pipeThrough and error propagation¶
readable.pipeTo(writable, options) moves all chunks from a readable to a writable, respecting backpressure, and returns a promise that settles when the pipe finishes. readable.pipeThrough(transform, options) pipes into transform.writable and returns transform.readable, so transforms chain naturally.
The options control how the end and errors cross the pipe:
| Option | Default | When true |
|---|---|---|
preventClose | false | Closing the source does not close the destination. Required to pipe several sources into one writable. |
preventAbort | false | An error in the source does not abort the destination. |
preventCancel | false | An error or close of the destination does not cancel the source. |
signal | none | An AbortSignal; aborting it stops the pipe, and unless prevented, aborts the destination and cancels the source. |
By default, errors travel forward (source error aborts the destination) and cancellation travels backward (destination closed or errored cancels the source). That is exactly what you want for a single pass-through pipe and precisely what you must override when composing.
TextEncoderStream and TextDecoderStream¶
new TextDecoderStream(label = "utf-8", { fatal, ignoreBOM }) turns bytes into strings and correctly handles multi-byte characters split across chunk boundaries, which a naive new TextDecoder().decode(chunk) per chunk does not (it would emit replacement characters where a UTF-8 sequence straddles two chunks). new TextEncoderStream() turns strings back into UTF-8 bytes; it always encodes UTF-8, which is also the only encoding you should serve HTML in. To edit text in a byte stream:
const edited = response.body
.pipeThrough(new TextDecoderStream())
.pipeThrough(someStringTransform)
.pipeThrough(new TextEncoderStream());
If you only concatenate byte streams, do not decode at all: bytes pass through untouched, and decoding and re-encoding costs CPU for nothing.
Backpressure and queuing strategies¶
Each stream has a queue and a high-water mark. When a consumer reads slowly, the queue fills, desiredSize drops to zero, pull() stops being called and pipeTo() stops reading from its source, all the way back to the network socket. In a service worker this means a slow page parse naturally slows the worker's network read instead of accumulating megabytes in worker memory. CountQueuingStrategy({ highWaterMark }) counts chunks; ByteLengthQueuingStrategy({ highWaterMark }) counts bytes. For HTML composition the defaults are fine.
Backpressure has one sharp edge: tee() and response.clone() buffer for the slower branch. If you clone a response to cache it and the cache write stalls, or you never read one branch, the unread branch accumulates every chunk in memory. Always consume both branches.
Streaming HTML composition¶
The pattern splits every page into three parts:
- Head partial (cached):
<!doctype html>,<html>, the whole<head>with CSS and scripts,<body>and the site header and navigation. - Content (network, falling back to cache or an offline block): only what is inside
<main>, rendered by the server when it recognizes a content-only request. - Foot partial (cached): footer, closing tags, late scripts.
The partials are precached at install time under a versioned cache, so the shell and the server's fragments always agree on markup. The first visit, crawlers and browsers without a controlling worker receive normal full documents; only worker-controlled navigations request fragments.
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>{{title}}</title>
<link rel="stylesheet" href="/assets/site.2026-09-25.css">
<script type="module" src="/assets/app.2026-09-25.js" nonce="{{nonce}}"></script>
<meta name="sw-composed" content="2026-09-25">
</head>
<body>
<header class="site-header">…</header>
<main id="content">
Three design rules for partials:
- Put nothing page-specific in the head partial except placeholders the worker can fill (
{{title}},{{nonce}}). Active navigation state, canonical URLs and meta descriptions are either filled the same way or set by a small script in the fragment. - Keep the head partial small and render-ready. The point is that the browser can paint the frame immediately; a head that waits on a slow third-party script loses the benefit.
- Version the partials together with the fragments. The worker tells the server which shell version it holds (see Server requirements), and the server renders compatible fragments or full pages.
A stream concatenation helper¶
The core utility takes a list of sources that are already started (so the network request for the content begins immediately, in parallel with reading the cached head) and emits them in order as one byte stream. A production helper must:
- accept
Response,ReadableStream, strings and binary data, and skipnullsources; - start all sources eagerly but read them strictly in order;
- propagate cancellation to the current source and to every source not yet reached;
- let the caller replace a failed source with fallback markup instead of erroring the whole document;
- expose a
donepromise forevent.waitUntil().
/* Classic-script helpers for a service worker: load with importScripts(). */
const utf8 = new TextEncoder();
/** Normalize a source value to a ReadableStream of Uint8Array, or null. */
function toByteStream(value) {
if (value == null) return null;
if (value instanceof ReadableStream) return value;
if (value instanceof Response) return value.body; // null for 204/304 or HEAD
if (typeof value === "string") return new Response(utf8.encode(value)).body;
if (value instanceof Uint8Array || value instanceof ArrayBuffer || value instanceof Blob) {
return new Response(value).body;
}
throw new TypeError(`Unsupported stream source: ${Object.prototype.toString.call(value)}`);
}
/**
* Concatenate sources into one byte stream.
* @param {Array<Promise<any>|any>} sources started in parallel, emitted in order
* @param {object} [options]
* @param {(error: unknown, index: number, started: boolean) => string|null} [options.onError]
* Return replacement markup to keep going, or null to error the stream.
* @returns {{ stream: ReadableStream<Uint8Array>, done: Promise<void> }}
*/
function concatStreams(sources, { onError } = {}) {
const pending = sources.map((s) => Promise.resolve(s));
pending.forEach((p) => p.catch(() => {})); // handled when reached; avoid unhandled rejections
let index = 0;
let reader = null;
let started = false; // has the current source produced any bytes?
let resolveDone;
let rejectDone;
const done = new Promise((res, rej) => { resolveDone = res; rejectDone = rej; });
function recover(error, controller) {
const replacement = onError ? onError(error, index - 1, started) : null;
if (replacement == null) {
controller.error(error);
cancelRemaining(error);
rejectDone(error);
return false;
}
controller.enqueue(utf8.encode(replacement));
return true;
}
function cancelRemaining(reason) {
for (const p of pending.slice(index)) {
p.then((v) => toByteStream(v)?.cancel(reason)).catch(() => {});
}
}
const stream = new ReadableStream({
async pull(controller) {
// Loop until we enqueue one chunk, close, or error: pull() must make progress.
for (;;) {
if (!reader) {
if (index >= pending.length) {
controller.close();
resolveDone();
return;
}
let body;
try {
body = toByteStream(await pending[index++]);
} catch (error) {
started = false;
if (!recover(error, controller)) return;
continue;
}
if (!body) continue;
reader = body.getReader();
started = false;
}
try {
const { value, done: finished } = await reader.read();
if (finished) {
reader.releaseLock();
reader = null;
continue;
}
started = true;
controller.enqueue(value);
return;
} catch (error) {
reader = null;
if (!recover(error, controller)) return;
return; // replacement enqueued; next pull() moves to the next source
}
}
},
cancel(reason) {
// The browser gave up on the response (user navigated away, tab closed).
reader?.cancel(reason).catch(() => {});
cancelRemaining(reason);
resolveDone(); // nothing left to wait for
},
});
return { stream, done };
}
/**
* Chunk-boundary-safe placeholder replacement for a text stream.
* Keys must not appear inside replacement values.
*/
function replaceStream(replacements) {
const keys = Object.keys(replacements);
const keep = Math.max(0, ...keys.map((k) => k.length - 1));
let buffer = "";
const applyAll = (text) =>
keys.reduce((acc, key) => acc.split(key).join(replacements[key]), text);
return new TransformStream({
transform(chunk, controller) {
buffer = applyAll(buffer + chunk);
// Hold back a tail that could be the start of a placeholder split across chunks.
if (buffer.length > keep) {
controller.enqueue(buffer.slice(0, buffer.length - keep));
buffer = buffer.slice(buffer.length - keep);
}
},
flush(controller) {
if (buffer) controller.enqueue(applyAll(buffer));
},
});
}
/** Escape text for safe insertion into HTML text or attribute values. */
function escapeHtml(text) {
return String(text).replace(/[&<>"']/g, (c) => (
{ "&": "&", "<": "<", ">": ">", '"': """, "'": "'" }[c]
));
}
Why the helper is built on pull() rather than on an identity TransformStream with a pipeTo() loop: both work, but the pull-based version gets backpressure and cancellation for free. When the browser cancels the response, cancel() runs, and the helper cancels the network body it is reading plus every body it has not reached yet. With the pipeTo(writable, { preventClose: true }) loop used on the Navigation Preload page, cancellation reaches only the source currently being piped, which is usually acceptable but leaves later sources to download to completion.
The complete streaming service worker¶
This worker precaches the partials, enables navigation preload, requests content fragments, fills the <title> and CSP nonce into the head partial, falls back to a cached fragment or an offline block, copies security headers, and keeps the worker alive until the last byte.
importScripts("/sw-stream-utils.js");
const SHELL_VERSION = "2026-09-25";
const SHELL_CACHE = `shell-${SHELL_VERSION}`;
const CONTENT_CACHE = "content-v1";
const PARTIALS = {
head: "/partials/head.html",
foot: "/partials/foot.html",
offline: "/partials/offline-content.html",
};
const HEADER_TIMEOUT_MS = 4000; // wait this long for the fragment's status line
const DEFAULT_TITLE = "Example";
const COPIED_HEADERS = [ // headers the document itself needs
"content-security-policy",
"content-security-policy-report-only",
"cross-origin-opener-policy",
"cross-origin-embedder-policy",
"permissions-policy",
"referrer-policy",
"x-content-type-options",
"reporting-endpoints",
];
const BYPASS = [/^\/api\//, /^\/admin\//, /^\/logout/];
self.addEventListener("install", (event) => {
event.waitUntil(
caches.open(SHELL_CACHE).then((cache) => cache.addAll(Object.values(PARTIALS)))
);
});
self.addEventListener("activate", (event) => {
event.waitUntil((async () => {
if (self.registration.navigationPreload) {
await self.registration.navigationPreload.enable();
// Tell the server which shell this worker holds; it renders matching fragments.
await self.registration.navigationPreload.setHeaderValue(`partial;v=${SHELL_VERSION}`);
}
const keys = await caches.keys();
await Promise.all(keys
.filter((k) => k.startsWith("shell-") && k !== SHELL_CACHE)
.map((k) => caches.delete(k)));
})());
});
self.addEventListener("fetch", (event) => {
const { request } = event;
if (request.mode !== "navigate" || request.method !== "GET") return;
const url = new URL(request.url);
if (url.origin !== self.location.origin || BYPASS.some((re) => re.test(url.pathname))) return;
event.respondWith(composeNavigation(event));
});
/** Content fragment: the preload response, or an explicit fragment request. */
async function fetchFragment(event) {
const preloaded = await event.preloadResponse; // undefined if preload is off or unsupported
if (preloaded) return preloaded;
return fetch(event.request.url, {
headers: { "Service-Worker-Navigation-Preload": `partial;v=${SHELL_VERSION}` },
credentials: "include",
redirect: "manual", // let the browser follow redirects for the real URL
cache: "no-cache",
});
}
/** Race a promise against a timer, and clear the timer whichever wins. */
function withTimeout(promise, ms) {
let timer;
const expired = new Promise((_, reject) => {
timer = setTimeout(() => reject(new Error("timeout")), ms);
});
return Promise.race([promise, expired]).finally(() => clearTimeout(timer));
}
async function composeNavigation(event) {
const cache = await caches.open(SHELL_CACHE);
const [head, foot] = await Promise.all([
cache.match(PARTIALS.head),
cache.match(PARTIALS.foot),
]);
// Shell missing (evicted, or install raced): do not compose. The preload response
// is a fragment, so discard it and fetch the full page (a plain request has no fragment header).
if (!head || !foot) {
event.waitUntil(Promise.resolve(event.preloadResponse)
.then((r) => r?.body?.cancel()).catch(() => {}));
return fetch(event.request);
}
// 1. Wait for the fragment's status line and headers, not its body.
let fragment = null;
let source = "network";
const pendingFragment = fetchFragment(event);
try {
fragment = await withTimeout(pendingFragment, HEADER_TIMEOUT_MS);
} catch {
fragment = null; // offline, DNS failure, or too slow
// A fragment that arrives after the timeout is never used: cancel its body so the
// download stops instead of running to completion in the background.
event.waitUntil(pendingFragment.then((r) => r.body?.cancel()).catch(() => {}));
}
// 2. Cases where composition is wrong: hand the response to the browser as is.
if (fragment) {
if (fragment.type === "opaqueredirect") return fragment; // browser follows it
const kind = fragment.headers.get("Content-Mode");
if (kind !== "partial") return fragment; // server sent a full page (old fragment version, error page)
}
// 3. Fallbacks when the network gave us nothing.
if (!fragment) {
fragment = await caches.match(event.request.url, { cacheName: CONTENT_CACHE, ignoreVary: true });
source = fragment ? "cache" : "offline";
fragment ??= await cache.match(PARTIALS.offline);
} else if (fragment.ok) {
// Keep a copy of successful fragments for offline revisits (unless the server says no).
const cc = fragment.headers.get("Cache-Control") ?? "";
if (!/no-store|private/.test(cc)) {
const copy = fragment.clone();
event.waitUntil(
caches.open(CONTENT_CACHE)
.then((c) => c.put(event.request.url, copy))
.catch(() => {}) // quota errors must not break the page
);
}
}
// 4. Values the head partial needs, taken from the fragment's headers.
const title = decodeTitle(fragment?.headers.get("Page-Title")) ?? DEFAULT_TITLE;
const csp = fragment?.headers.get("Content-Security-Policy") ?? "";
const nonce = /'nonce-([A-Za-z0-9+/=_-]+)'/.exec(csp)?.[1] ?? "";
const headBody = head.body
.pipeThrough(new TextDecoderStream())
.pipeThrough(replaceStream({ "{{title}}": escapeHtml(title), "{{nonce}}": escapeHtml(nonce) }))
.pipeThrough(new TextEncoderStream());
const { stream, done } = concatStreams([headBody, fragment, foot], {
onError(error, index, started) {
console.warn("stream source failed", index, error);
if (index !== 1) return null; // a broken cached partial: nothing sensible to show
// Content failed. Close any open markup the fragment might have left and explain.
return `${started ? "</div></section></article>" : ""}
<div class="stream-error" role="alert">
<p>The rest of this page could not be loaded.</p>
<p><a href="${escapeHtml(event.request.url)}">Try again</a></p>
</div>`;
},
});
event.waitUntil(done.catch(() => {})); // keep the worker alive until the last byte
// 5. Headers: our own plus the security headers the document needs.
const headers = new Headers({
"Content-Type": "text/html; charset=utf-8",
"X-Composed-By": `sw;shell=${SHELL_VERSION};content=${source}`,
});
for (const name of COPIED_HEADERS) {
const value = fragment?.headers.get(name);
if (value) headers.set(name, value);
}
const status = source === "network" && fragment.status >= 400 ? fragment.status : 200;
return new Response(stream, { status, headers });
}
function decodeTitle(value) {
if (!value) return null;
try { return decodeURIComponent(value); } catch { return null; }
}
Walkthrough of the decisions in this worker:
- It waits for the fragment's headers before sending the first byte. The status code, redirects, the title and the CSP nonce all live in those headers and cannot be changed after streaming starts. Because navigation preload starts the fragment request in parallel with worker startup, the wait is typically about one server time-to-first-byte, while the fragment itself is smaller and faster to render than a full page. The trade-off, and the alternative, are discussed in Committing early versus waiting for headers.
- Redirects pass straight through. The fallback fragment request uses
redirect: "manual", so a redirect comes back as anopaqueredirectresponse, whichrespondWith()may return for a navigation (whose redirect mode is alsomanual); the browser then follows it with the realLocation. The preload response behaves the same way. Composing a redirect into a200page would leave the user on the wrong URL. - A full page from the server is passed through. The server marks fragments with a
Content-Mode: partialresponse header. Anything else (a full error page from a proxy, a maintenance page, a response for an outdated shell version) is served unmodified, which is always correct. ignoreVary: trueis required when reading cached fragments: the server'sVary: Service-Worker-Navigation-Preloadheader would otherwise make the Cache API refuse to match a plain URL against an entry stored from a request that carried the header.- The fragment is cloned for the cache inside
waitUntil(), so the cache write never delays the response, and both branches of the tee are consumed. - Cancellation is free: if the user navigates away,
concatStreamscancels the fragment body, which aborts the network download.
Register the worker as usual (see Registration & Scope). The helper file must be served from the same origin and is fetched by importScripts() during installation; bump the worker when it changes.
Setting the title, active navigation and other per-page head data¶
The biggest practical limitation of composition is that the head partial is written before the page content is known. The worker above solves the <title> by reading a Page-Title header from the fragment and filling a placeholder. The same technique covers any small per-page value (canonical URL, meta name="description", og: tags, the active navigation item) as long as the server exposes it in fragment headers and the value is escaped before insertion. HTTP header values must be ASCII-safe, so percent-encode on the server and decodeURIComponent() in the worker.
For values that are too large or too dynamic for headers, let the fragment carry them and apply them from a tiny inline script at the top of the fragment:
<script nonce="rAnd0mPerResp0nse">
document.title = "Order #4821 · Example";
document.querySelector('nav a[href="/orders"]')?.setAttribute("aria-current", "page");
</script>
Search engines see the server's full page on their first request, since crawlers do not run your service worker, so composition does not affect what gets indexed. See SEO for PWAs.
Content Security Policy with composed pages¶
A cached head partial is static, but a nonce-based CSP requires a fresh, unguessable nonce per response. Two approaches work:
- Nonce substitution (as in the worker above): the server generates a nonce for the fragment response, sends it in the fragment's CSP header and uses it in the fragment's own scripts; the worker parses the nonce from that header, writes it into the head partial's
{{nonce}}placeholders and copies the CSP header onto the composed response. The nonce stays per-response. Only same-origin script can see it, and such script could already read the DOM. - Hash-based CSP: list
'sha256-…'hashes of the partials' inline scripts, or avoid inline scripts in partials altogether and allow them by URL with'strict-dynamic'from a hashed loader.
What does not work is caching a head partial that contains a literal nonce: every composed page would reuse it, which defeats the purpose of nonces. Note that a fragment served from the content cache while offline replays the nonce it was originally delivered with (its cached CSP header and its scripts agree, so it works); if replaying a nonce is unacceptable in your threat model, strip inline scripts from fragments before caching them or rely on hashes. Content Security Policy covers the service worker side of CSP in detail.
Committing early versus waiting for headers¶
Waiting for fragment headers is simpler and correct, but on a slow server the user looks at a blank screen for one server think time. The alternative is to commit immediately: stream the head partial as soon as it is read from the cache, and only then wait for the fragment.
What you give up by committing early:
| Concern | Waiting for headers | Committing early |
|---|---|---|
| Time to first paint | Server TTFB (with preload) | Cache read, typically a few milliseconds after worker start |
| Status code | Reflects the server (404, 410, 500) | Always 200 |
| Redirects | Followed by the browser | Must be handled in-band, for example a server fragment that contains <script>location.replace(…)</script> plus a link |
<title> and head metadata | Filled from fragment headers | Set by script inside the fragment |
| CSP nonces | Parsed from the fragment | Must use hashes, or a nonce generated in the worker and sent to the server, which the server must accept |
| Error pages | Full server page passed through | Rendered inside the shell |
Committing early fits app-like sections where every URL is valid, redirects are rare and a consistent shell is the priority (an inbox, a dashboard). Waiting for headers fits content sites where status codes, redirects and metadata matter. A hybrid is common: commit early for a known set of routes and wait for headers everywhere else.
To commit early with the worker above, skip step 1, pass fetchFragment(event) (a promise) directly as the second source, and move the redirect and error handling into onError and the fragment markup.
Using workbox-streams¶
Workbox's workbox-streams module (part of Workbox 7.4.x as of September 2026) packages the same idea. Its API:
| Function | Signature | Purpose |
|---|---|---|
strategy() | strategy(sourceFunctions, headersInit) | Returns a route handler. Each source function receives the route handler options (event, request, url, params) and returns a StreamSource or a promise of one |
concatenate() | concatenate(sourcePromises) | Returns { done, stream } for promises of Response, ReadableStream or BodyInit |
concatenateToResponse() | concatenateToResponse(sourcePromises, headersInit) | Returns { done, response } |
isSupported() | isSupported() | Whether the browser can construct a ReadableStream. When it cannot, strategy() waits for all sources and concatenates them into one buffered response |
import { precacheAndRoute, matchPrecache } from "workbox-precaching";
import { registerRoute } from "workbox-routing";
import { NetworkFirst } from "workbox-strategies";
import { strategy as composeStrategies } from "workbox-streams";
import * as navigationPreload from "workbox-navigation-preload";
precacheAndRoute(self.__WB_MANIFEST); // includes /partials/head.html and /partials/foot.html
navigationPreload.enable("partial;v=2026-09-25");
const contentStrategy = new NetworkFirst({
cacheName: "content-v1",
networkTimeoutSeconds: 4,
plugins: [{
// Ask the server for fragments when preload is unavailable.
requestWillFetch: async ({ request }) => {
const headers = new Headers(request.headers);
headers.set("Service-Worker-Navigation-Preload", "partial;v=2026-09-25");
return new Request(request.url, { headers, credentials: "include" });
},
handlerDidError: async () => matchPrecache("/partials/offline-content.html"),
}],
matchOptions: { ignoreVary: true },
});
// matchPrecache() resolves to undefined on a miss, which would break concatenate().
const precached = (url, fallback) => async () => (await matchPrecache(url)) ?? fallback;
const navigationHandler = composeStrategies(
[
precached("/partials/head.html",
'<!doctype html><html lang="en"><head><meta charset="utf-8"><title>Example</title></head><body><main>'),
({ event, request }) => contentStrategy.handle({ event, request }),
precached("/partials/foot.html", "</main></body></html>"),
],
{ "Content-Type": "text/html; charset=utf-8" }
);
registerRoute(({ request, url }) =>
request.mode === "navigate" && !url.pathname.startsWith("/api/"), navigationHandler);
NetworkFirst uses event.preloadResponse automatically when navigation preload is enabled. Under the hood, the Workbox strategy passes the done promise to event.waitUntil(), like the hand-written worker.
What workbox-streams does not do for you (checked against the 7.4.1 source):
- It always commits early. There is no "wait for headers" mode; the status is always
200and the headers are whatever you passed tostrategy(). - It does not copy security headers from the content response, so a composed page runs without the server's CSP unless you add it to
headersInityourself. - Redirects and full pages are not detected. A redirect or full-page response from the network is concatenated into the shell as if it were a fragment.
- A failing source errors the whole stream.
concatenate()rejectsdoneand errors the response when any source read fails; there is no hook to substitute closing markup.handlerDidErrorin the example only covers failures before the content response arrives. - A missing source throws.
matchPrecache()resolves toundefinedwhen the URL is not in the precache (for example after a manifest change), andconcatenate()then fails reading a null body. Guard source functions with a fallback string. - Cancellation is not propagated. When the browser cancels the composed response,
concatenate()resolvesdonebut does not cancel the source readers, so the content download continues until it finishes.
If those matter, wrap contentStrategy.handle() in your own source function that inspects the response and catches errors, or use the hand-written worker. Workbox Fundamentals and Advanced Workbox cover strategies and plugins.
Headers and status codes on synthesized responses¶
When you return new Response(stream, init), the browser uses exactly the status and headers in init. The network response you read from is only a data source; none of its metadata carries over automatically.
- Status
- Must be decided before the first byte.
new Response()accepts 200–599 (and throws aRangeErrorotherwise). Null-body statuses (204,205,304) cannot have a body at all. Content-Type- Always set
text/html; charset=utf-8. The header's charset takes precedence over<meta charset>(only a byte order mark overrides it), so the parser never has to guess the encoding or restart when it meets a late or conflicting meta declaration. Content-LengthandContent-Encoding- Never copy them from a network response.
fetch()bodies are already decoded, and a composed body has a different length. Both headers describe bytes on the wire that no longer exist. - Security headers
- A document's
Content-Security-Policy,Cross-Origin-Opener-Policy,Cross-Origin-Embedder-Policy,Permissions-PolicyandReferrer-Policycome from the response that created it. A composed response without them runs without a CSP and without cross-origin isolation (crossOriginIsolatedbecomesfalse). Copy them from the fragment, as the worker above does, and make the server send the same values on fragment responses as on full pages. Cache-Control- Irrelevant to the Cache API and to the HTTP cache for worker-produced documents, but it does influence other browser features.
Cache-Control: no-storeon the main document has historically made Chromium exclude pages from the back/forward cache; if you want bfcache for composed pages, omit it or useno-cache. See HTTP Caching & Service Workers. Set-Cookie- Ignored on responses the worker synthesizes (it is a forbidden response header name for script). Cookies set by the fragment's network response are already stored by the time the worker sees it.
Server-Timing, custom diagnostics- Add your own marker (
X-Composed-Byabove) to see in DevTools which responses were composed and whether content came from the network, the cache or the offline block.
The response's url is empty for a synthesized Response, which does not matter for navigations: the document URL is the request URL. response.redirected is false. Handling Fetch Events lists the other rules respondWith() enforces.
Server requirements¶
Streaming composition is half a server feature. The server must render fragments on request, send them without buffering, and keep caches from mixing fragments with full pages.
Serving fragments and full pages from the same URL¶
- Detect fragment requests by the
Service-Worker-Navigation-Preloadrequest header (sent on preload requests with the value you configured, and set explicitly by the worker's fallback fetch). - Validate the shell version in the header value. If the server no longer supports that version, return the full page: the worker passes full pages through unchanged, and the next worker update fixes the mismatch.
- Mark fragments with a response header (
Content-Mode: partialabove) so the worker never mistakes a full page for a fragment. - Send
Vary: Service-Worker-Navigation-Preloadon both fragments and full pages, so shared caches and CDNs store them separately. Without it, a CDN can serve a fragment to a first-time visitor (an unstyled page) or a full page into the shell (a page nested in a page). - Send the same security headers on fragments as on full pages.
import express from "express";
import crypto from "node:crypto";
import { renderHeadAndHeader, renderFooter, renderContent } from "./render.js";
const app = express();
const SUPPORTED_SHELLS = new Set(["2026-09-25", "2026-09-10"]);
// A regular expression route works the same in Express 4 and 5.
app.get(/^\/(articles\/[^/]+)?$/, async (req, res, next) => {
const preload = req.get("Service-Worker-Navigation-Preload") ?? "";
const shell = /^partial;v=(.+)$/.exec(preload)?.[1];
const partial = shell !== undefined && SUPPORTED_SHELLS.has(shell);
const nonce = crypto.randomBytes(16).toString("base64");
res.set({
"Content-Type": "text/html; charset=utf-8",
"Vary": "Service-Worker-Navigation-Preload",
"Content-Security-Policy": `script-src 'nonce-${nonce}' 'strict-dynamic'; object-src 'none'; base-uri 'none'`,
"X-Accel-Buffering": "no", // nginx: do not buffer this response
"Cache-Control": "private, no-cache",
});
try {
const page = await loadPageMeta(req.path); // fast: title, status, redirect
if (page.redirect) return res.redirect(301, page.redirect);
res.status(page.status);
if (partial) {
res.set("Content-Mode", "partial");
res.set("Page-Title", encodeURIComponent(page.title));
}
res.flushHeaders(); // status line and headers now
if (!partial) res.write(renderHeadAndHeader(page, nonce));
for await (const chunk of renderContent(page, nonce)) {
res.write(chunk); // stream content as it is produced
}
if (!partial) res.write(renderFooter(page, nonce));
res.end();
} catch (err) {
if (!res.headersSent) return next(err);
// Headers are gone: close the document with an error block instead.
res.end(`<div class="stream-error" role="alert">Something went wrong.</div>`);
}
});
async function loadPageMeta(path) {
// Look up the page cheaply before rendering, so status and redirects are known first.
return { status: 200, title: "Article title", redirect: null, path };
}
app.listen(8080);
The server applies the same rule as the worker: decide status, redirects and headers first, then stream. loadPageMeta() does the cheap lookups that determine them; renderContent() is an async generator that yields HTML as slower data arrives. Streaming SSR frameworks (React's renderToReadableStream(), SvelteKit, Astro and others) already work this way; the Architecture overview discusses them.
Making sure bytes are not buffered on the way¶
Every hop between your renderer and the browser can buffer:
| Hop | Symptom | Fix |
|---|---|---|
Compression middleware (for example Express compression) | Nothing arrives until the compressor's buffer fills or the response ends | Call res.flush() after important chunks (the compression middleware adds it), or compress at the proxy/CDN |
| nginx reverse proxy | Response arrives in one piece | proxy_buffering off; for the location, or send X-Accel-Buffering: no from the app |
| Other proxies, load balancers, serverless adapters | Same | Check that the platform supports streaming responses; some adapters buffer whole responses |
| CDN | Usually streams; some features (HTML rewriting, some optimizations) buffer | Test with a slow origin and curl -N |
| HTTP/1.1 | Needs Transfer-Encoding: chunked when there is no Content-Length | Node sets it automatically when you write() without a length |
Test the whole path, not the origin alone: curl -N -H "Service-Worker-Navigation-Preload: partial;v=2026-09-25" https://example.com/articles/x prints chunks as they arrive, and a response that appears all at once after a delay is being buffered somewhere.
Handling errors mid-stream¶
Errors fall into two classes, and they need different handling.
Before the first byte, you are free: return a different response, a cached page, an offline page, or Response.error() (which the browser shows as a network error). The worker above handles network failures, timeouts, redirects and server error pages at this stage.
After the first byte, the status and headers are committed and part of the document is already rendered. Your options:
- Close the document gracefully (recommended). Catch the failure in the stream source, enqueue markup that closes any open elements and explains the problem, then continue with the footer and close the stream normally. The page is valid, navigable and has a retry link. This is what the
onErrorcallback inconcatStreamsdoes. - Error the stream with
controller.error(). The browser treats the response body as a network error partway through: the part already parsed stays on screen, the rest of the document (including your footer and late scripts) never arrives, and how the load is reported differs between browsers. Avoid this for documents. - Let an exception escape. Equivalent to erroring the stream, plus an unhandled rejection in the worker. Never do this.
Specific failure modes to design for:
- The server fails halfway through rendering. It cannot change the status either. Have the server close its own markup and emit an error block (the Express example's
catchdoes this), so the worker receives a well-formed fragment. - The connection drops mid-body.
reader.read()rejects with aTypeError. The helper'sonErrorreceivesstarted === trueand emits closing tags plus the error block. Tailor the closing tags to your fragment structure, or keep fragments shallow (content directly in<main>) so a truncated fragment needs little closing. - A cached partial is missing or corrupt. Check that both partials exist before composing and fall back to plain network navigation if they do not, as the worker does. Storage can be evicted independently of the service worker registration; see Storage Quotas & Persistence.
- The worker is terminated mid-stream. Browsers stop idle workers, and a worker streaming without an extended lifetime can be stopped before the stream finishes, truncating the page.
event.waitUntil(done)prevents that for the duration of the stream (within the browser's maximum event lifetime). - Slow content after a fast shell (commit-early mode). A header followed by a long blank wait looks broken. Put a lightweight loading indicator at the end of the head partial and hide it with CSS once content arrives (for example
main:has(> *:not(.loading)) .loading { display: none; }), or emit a skeleton block from the worker when the fragment has not produced a byte within a second.
Streaming beyond composed HTML¶
The same primitives are useful elsewhere in a PWA.
Streaming JSON into the page incrementally¶
For long lists, newline-delimited JSON (NDJSON) lets the page render items as they arrive rather than after the whole array has been parsed:
/** Yield one parsed object per line of an NDJSON response, as bytes arrive. */
export async function* readNdjson(response) {
if (!response.ok || !response.body) throw new Error(`HTTP ${response.status}`);
const reader = response.body.pipeThrough(new TextDecoderStream()).getReader();
let buffer = "";
try {
for (;;) {
const { value, done } = await reader.read();
if (done) break;
buffer += value;
let newline;
while ((newline = buffer.indexOf("\n")) >= 0) {
const line = buffer.slice(0, newline).trim();
buffer = buffer.slice(newline + 1);
if (line) yield JSON.parse(line);
}
}
if (buffer.trim()) yield JSON.parse(buffer);
} finally {
reader.releaseLock(); // lets the caller cancel response.body if it stops early
}
}
// Usage: render the first results while the rest are still downloading.
const controller = new AbortController();
const res = await fetch("/api/search?q=offline&format=ndjson", { signal: controller.signal });
for await (const item of readNdjson(res)) {
appendResult(item);
}
The example uses getReader() rather than for await (const chunk of stream), because async iteration of ReadableStream arrived in Safari only with Safari 27 (September 2026).
Download progress from a stream¶
response.body also lets you measure progress, which fetch() does not report directly. Read chunks, add up value.byteLength, and compare with Content-Length when present (it is the encoded size, so progress for compressed responses can pass 100%; treat it as an estimate). For large downloads that must survive the tab closing, use Background Fetch instead.
Streaming request bodies¶
Chromium 105 and later can send a ReadableStream as a request body with fetch(url, { method: "POST", body: stream, duplex: "half" }). Streaming uploads are restricted to HTTP/2 and HTTP/3, always trigger a CORS preflight, cannot use no-cors, and reject on any redirect except 303. MDN's compatibility data lists the duplex option as unsupported in Chrome for Android, Firefox and Safari as of September 2026, so treat request streaming as a progressive enhancement with a buffered fallback. The Chrome team's streaming requests article includes a feature-detection snippet.
Compressing and decompressing in a worker¶
CompressionStream and DecompressionStream ("gzip", "deflate", "deflate-raw") are available in Chrome 80, Firefox 113 and Safari 16.4 and later; Firefox 147 and Safari 18.4 also support "brotli", Chromium does not yet. They are useful for compressing large JSON exports stored in IndexedDB or OPFS, not for HTTP responses, which the browser decompresses for you.
Measuring the benefit¶
Measure composition against your current navigation strategy with real users, not only in the lab.
- Put a marker in the head partial (
<meta name="sw-composed">above) and report it with your RUM beacon, together with whethernavigator.serviceWorker.controllerwas set, so you can segment FCP and LCP by composed, uncontrolled and non-composed navigations. - From
performance.getEntriesByType("navigation")[0], compareworkerStart(worker startup),responseStart(first byte from the worker) andresponseEnd(last byte). For composed pages,responseStartshould be close to the preload's server TTFB in wait-for-headers mode and close toworkerStartin commit-early mode. - Track LCP, not only FCP: if the LCP element lives in the content fragment, composition helps FCP more than LCP, and server rendering time still dominates LCP.
- Run the change as an experiment, for example by varying the worker behavior per client with a flag stored in IndexedDB, and compare distributions, not averages.
Measuring Performance and Core Web Vitals cover RUM collection, and Navigation Preload shows how to A/B test worker-level changes.
Browser support¶
| Feature | Chrome / Edge | Firefox | Safari (macOS and iOS) | Notes |
|---|---|---|---|---|
ReadableStream constructor, Response with a stream body | ✅ 52+ | ✅ 65+ | ✅ 10.1+ | Everything composition needs |
pipeTo() / pipeThrough() | ✅ 59+ | ✅ 100+ / 102+ | ✅ 10.1+ | |
TransformStream | ✅ 67+ | ✅ 102+ | ✅ 14.1+ | |
TextEncoderStream / TextDecoderStream | ✅ 71+ | ✅ 105+ | ✅ 14.1+ | |
WritableStreamDefaultController.signal | ✅ 98+ | ✅ 100+ | ✅ 16.4+ | |
Navigation preload (event.preloadResponse) | ✅ 59+ | ✅ 99+ | ✅ 15.4+ | See Navigation Preload |
Async iteration of ReadableStream | ✅ 124+ | ✅ 110+ | ✅ 27+ | Use getReader() loops for older Safari |
ReadableStream.from() | ❌ | ✅ 117+ | ✅ 27+ | Not in Chromium as of September 2026 |
Byte streams (type: "bytes", BYOB readers) | ✅ 89+ | ✅ 102+ | ✅ 26.4+ | |
Transferable ReadableStream (postMessage) | ✅ 87+ | ✅ 103+ | ✅ 27+ | Safari does not transfer WritableStream or TransformStream |
Streaming request bodies (duplex: "half") | ⚠️ 105+ desktop | ❌ | ❌ | Not in Chrome for Android per MDN; HTTP/2+ only |
CompressionStream gzip / deflate | ✅ 80+ | ✅ 113+ | ✅ 16.4+ | Brotli: Firefox 147+, Safari 18.4+, not Chromium |
Support data as of September 2026. See MDN: Streams API and caniuse: Streams for live data.
Common pitfalls¶
- Buffering in the worker.
await response.text(),await response.blob()or string concatenation in afetchhandler destroys streaming. Transform with streams or pass bodies through. - Enqueuing strings. A response stream must contain
Uint8Arraychunks; strings fail when the browser reads them. UseTextEncoderorTextEncoderStream. - Forgetting
event.waitUntil(). The worker can be stopped before the stream finishes and the page is truncated. - Dropping security headers. A composed document without the server's
Content-Security-Policyor COOP/COEP silently runs without them. Copy them from the fragment response. - Copying
Content-LengthorContent-Encoding. They describe the network bytes, not your composed body. - Composing redirects and error pages. A redirect streamed into a
200shell leaves the user on the wrong URL; a full server error page inside the shell nests two documents. Inspect the fragment's headers first. - No
Varyon fragment responses. A CDN serves fragments to first-time visitors or full pages into the shell. - Unversioned partials. A new server layout with an old cached shell produces broken pages until the worker updates. Send the shell version to the server and let it fall back to full pages.
- Caching the composed document. Cache fragments and partials separately; a cached composed page freezes an old shell around old content.
- Cloning without reading both branches. An unread
clone()ortee()branch buffers the whole body in memory. - Relying on async iteration or
ReadableStream.from()without checking support. Chromium lacksfrom(), and Safari only gained both in Safari 27. - Reverse-proxy or middleware buffering. The worker streams perfectly, but nothing arrives until the server finishes. Test the whole path with
curl -N.
Debugging¶
- DevTools Network panel: composed navigations show "(ServiceWorker)" in the Size column in Chromium. The Timing tab shows service worker startup and "respondWith" durations, and the Headers tab shows the headers your worker set (look for your
X-Composed-Bymarker and the copied CSP). - Slow-network throttling makes streaming visible: with a throttled profile, the head partial should paint immediately and content should appear progressively. If everything appears at once, something buffers.
- Console in the worker context: in Chromium, choose the service worker in the console's context selector (or open it from Application > Service workers) to see the
stream source failedwarnings fromonError. - Simulate failures: stop the server mid-response, or add a test route that writes half a fragment and destroys the socket, to verify that the page closes gracefully.
- Check the server alone:
curl -Nwith and without theService-Worker-Navigation-Preloadheader should return a fragment and a full page respectively, both streaming.
Browser DevTools covers service worker debugging in each browser, and Pitfalls & Anti-Patterns lists other ways a fetch handler goes wrong.
Further reading¶
On this site
- Navigation Preload: starting the content request in parallel with worker startup
- Handling Fetch Events:
respondWith(), response types and rules - Architecture: how streaming SSR and other rendering strategies pair with worker strategies
- App Shell Model: the SPA alternative to composition
- Precaching & Runtime Caching: keeping partials and fragments cached and versioned
- Workbox Fundamentals:
workbox-streamsin context - Content Security Policy: nonces and hashes with service workers
- Loading Performance: FCP, LCP and the critical rendering path
External references
- Streams Standard (WHATWG) and Fetch Standard
- MDN: Streams API, ReadableStream, TransformStream
- web.dev: Streams, the definitive guide
- Chrome for Developers: workbox-streams
- Chrome for Developers: Faster multipage applications with streams
- Jake Archibald: Stream your way to immediate responses
- Chrome for Developers: Streaming requests with the fetch API
- Service Workers specification
- nginx
proxy_bufferingand Expresscompressionmiddleware