One microVM per page: Chromium on Firecracker snapshots

In August Cloudflare introduced Kitesurf, a browser for agents that runs on Workers, one V8 isolate per navigation, and uses native Rust compiled directly to WebAssembly where possible. Their starting point is that Chromium is too heavy: it uses “so much memory and compute that providing every agent with its own instance is prohibitively expensive”. This page looks at the other answer: keep unmodified Chromium and give each request its own microVM, restored from a snapshot in which Chromium is already booted and warm, with the memory server recording which pages clones touch and mapping them into the next one. It measures how long that takes on Cloudflare’s 14-URL corpus. What it costs per request in CPU and memory, the axis Cloudflare’s post is about, is not measured here.

549.4 msmedian per request at 4 guest vCPUs, from launching the VM to the screenshot returned, 202 requests over the 14 corpus URLs
770.3 msthe same at 2 guest vCPUs; a second 2 vCPU run on another day read 712.6 ms
37% shorterrequest median with working-set replay on than off, synthetic page, copy mode, one golden (477.9 against 762.6 ms)
0 failuresin each of the four published corpus runs: 0 of 230 render attempts per run (202 measured, 28 warmup), exact 95% interval 0 to 1.59%; the interleaved no-render arm also had 0 failures in 230

Setup

ItemValue
Corpus hostAWS c8g.metal-48xl (Graviton4, Neoverse-V2, 192 vCPUs, 377 GiB), Ubuntu 24.04, kernel 6.17.0-1019-aws. Recorded in results/reqbench-20260902-025115-corpus-c4/hostinfo.json. The 2026-08-30 run used the same kernel; its instance is not recorded.
Fixture hostAn aarch64 host on kernel 7.0.14-fcvm-cd6cd2b4b52e, recorded in each reqbench fixture run’s analysis.json and in HC’s run.json. Its instance type, CPU model and core count were not recorded, and the fault-count run (FB) records no host at all. Used only for the synthetic-page runs.
GuestCorpus runs: 1024 MiB RAM, 2, 4 or 8 vCPUs, Linux 6.18.44. Reqbench fixture runs: 1024 MiB, 2 vCPUs. The fault-count run (FB) used other guests, described with its chart. Headless Chromium from Debian bookworm’s unpinned chromium package, in a Podman container tagged localhost/chromium-bench-req. The tag is mutable; the records pin image sha256:5b870d814e5d… for the vCPU ladder and sha256:334fc21f7c8c… for the 2026-08-30 run (image_id in each reqbench.jsonl meta line). No record keeps the Chromium version.
MemoryCorpus runs: clone RAM is filled on first touch by fcvm’s userfaultfd server in minor mode. The server copies the snapshot’s memory once into a sealed in-memory file (a memfd) that all clones share; clean pages map that file, and a page a clone writes becomes private to it. The campaign turns working-set prefetch on unless the environment overrides it (corpus_campaign.sh); the records do not store the value. The synthetic-page runs also test copy mode and file-backed restores.
NetworkRootless (pasta). The guest’s DNS resolves every corpus host to a replay server on the host.
WorkloadThe 14 URLs of kitesurf.cloudflare.app/corpus.txt, captured once from the live sites. The captured response bodies are served from the host. URL matching, headers and some statuses differ from a live origin. An uncaptured GET, HEAD or POST gets an empty 404, and a CORS preflight gets 204. 28 warmup requests, then 202 measured, cycling the URLs uniformly (14 or 15 renders each).
OperationRestore a clone, wait for Chromium’s DevTools target, navigate, Page.captureScreenshot (JPEG, quality 80). The clone is destroyed after the JPEG is returned.
LatencyThe request’s blocking time, arms.cdp.blocking_ms in analysis.json: from just before the harness launches the VM until its CDP driver returns. After the screenshot arrives the driver decodes and checks the JPEG, reads the navigation timing and closes the connection before returning; the decode and timing read take 1.9 ms median at 4 vCPUs. Teardown runs after the answer and is excluded (median 63.7 ms at 4 vCPUs). The record’s wall_ms includes it and reads 623.1 ms at 4 vCPUs.
Softwarefcvm source 1e9e9b70 for the vCPU ladder, 55756858 for the 2026-08-30 run. Firecracker from a fork; the corpus and fixture runs used different branches of it, listed under Firecracker builds.
Golden
The snapshot every request restores. A VM boots Chromium, warms it on an unrelated local page (one navigation and one screenshot), navigates to about:blank and is snapshotted: memory, disk and device state. Regenerating it makes a new golden. The guest vCPU count is fixed in the golden.
vCPU ladder
Three corpus runs on 2026-09-02 at 2, 4 and 8 guest vCPUs, one golden each, back to back on one host with one build.
Within-run interval
The 95% bootstrap confidence interval of a run’s median, computed from that run’s 202 requests. It does not include run-to-run variance.
No-render arm
Requests interleaved with the renders that restore a clone and stop the clock when its forwarded DevTools port accepts a connection, with no navigation. The drift gate watches it.

Two answers to one question

Cloudflare’s post asks how to give every agent its own browser. Kitesurf answers with a new engine built as several Workers components, each in its own V8 isolate: the Engine, which holds the CDP session and fetches the page’s document and scripts; PageScript, one isolate per navigation, which runs the page’s script; PageRenderer, which paints; and SandboxOutbound, the only path to the network. Where it can, it uses native Rust compiled to WebAssembly, with parts of Blitz and Stylo for HTML and CSS. fcvm keeps the engine and moves the isolation boundary out to a microVM, so that each page gets its own kernel under an unmodified Chromium.

From an untrusted page to the network, in each designKitesurf’s path follows the isolation figure and component descriptions in Cloudflare’s post, and the Workers security model for the isolate boundary. PageRenderer, which paints and has no network access, is left out. The dashed box is each design’s isolation boundary.

KitesurfCloudflare; Rust compiled to WebAssembly, on Workers

Untrusted pageHTML, CSS and script from the site
 
Engineholds the session; fetches the main document and scripts through SandboxOutbound
starts one isolate per navigation
PageScriptDOM in WebAssembly; page script in V8, eval in Boa
 
Workers isolate boundaryV8 isolates that can share a runtime process; runtime processes run in a namespace and seccomp sandbox
 
SandboxOutboundthe only path to the network: CORS, headers, cookies
 
The internetlive origins

fcvmunmodified Chromium in a Firecracker microVM per request

Untrusted pageHTML, CSS and script from the site
one microVM per request
Chromium, restoreda copy-on-write clone of a warm browser, discarded after one page
 
KVM boundary: Firecracker microVMits own guest kernel, vCPUs, memory and disk
 
pastauser-mode networking in the clone’s own network namespace
 
The internetin this benchmark, a replay server on the host

The two designs pay for isolation in different places. Kitesurf rebuilds the engine as Workers components, and Cloudflare’s post lists what it cannot do yet: play video, render WebGL, negotiate a bot-challenge handshake with real TLS fingerprints, or start a ten-minute authenticated session that requires persistent state. It implements a subset of CDP. fcvm keeps the whole engine and restores a VM for each request instead of booting one. At 4 guest vCPUs the recorded port-accept and target-wait stages sum to a median of 197.7 ms. The driver’s setup between them is not timed and the harness polls every 50 ms, so the moment the restored Chromium was ready is not recorded. It is a time, not a cost: what a restore costs in CPU and memory is not measured. For scale only, Cloudflare’s later benchmark leaves out the launch of a fresh Chrome, a median of 215 ms across its pages, on hardware it does not state. Cloudflare publishes no spin-up time for a Kitesurf isolate.

The same corpus

Cloudflare’s post reports “the medians of five Browser Run quick-action runs across a 14-URL corpus”, Kitesurf against Chromium in a warm pool. The corpus file, kitesurf.cloudflare.app/corpus.txt, is no longer served; the Wayback Machine holds identical copies from 2026-08-07 and 2026-08-14. We captured those 14 URLs and replayed them to a new microVM per request.

Published screenshot wall time (Cloudflare) and request median (fcvm), 14-URL corpusFigures from different setups and different statistics, grouped by system. The setups differ as listed below, so the bars are not a controlled ranking. Each fcvm bar is one run; our one repeated configuration, 2 vCPUs, moved 57.7 ms between two runs.
  • Cloudflare Kitesurf, isolated per request
  • Cloudflare Chromium, warm pool
  • fcvm, new microVM per request
Cloudflare, published medians
Kitesurf1,148 ms
Chromium637 ms
fcvm, one run each
2 vCPU, 09-02770.3 ms
2 vCPU, 08-30712.6 ms
4 vCPU549.4 ms
8 vCPU580.2 ms

fcvm bars: from launching the VM to the screenshot returned, teardown excluded (Latency, under Setup). The 4 and 8 vCPU runs are from 2026-09-02.

These figures sit side by side; they are not a controlled comparison. The setups differ:

As numbers, our median is below their Kitesurf median at every vCPU count we ran, and below their warm Chromium median at 4 and 8 vCPUs but not at 2. Our means, 1,035.1, 822.8 and 844.7 ms at 2, 4 and 8 vCPUs, are above their warm Chromium figure at every count. Which of our statistics matches theirs is not known, and none of the differences above is priced, so neither ordering ranks the systems.

Cloudflare’s own table is a trade between cost and time. By their relative column, against their warm Chromium, Kitesurf uses 3.1 and 3.8 times less CPU and 4.7 and 7.0 times less memory, and is 1.8 and 1.7 times slower on wall time. The chart above shows only time, the axis on which Kitesurf is the slower of their two engines. We have no CPU or memory figure to set against the other rows.

Cloudflare’s published table, with fcvm’s figures

MetricKitesurfCloudflare ChromiumKitesurf, relativefcvmNotes
CPU, screenshot380 ms1,173 ms3.1x less CPUnot establishedNo whole-system CPU measurement of the corpus runs.
CPU, HTML extraction229 ms877 ms3.8x less CPUnot measurednot measured: no run performs this operation.
Memory, screenshot57.8 MiB271.0 MiB4.7x less memorynot establishedClones share clean snapshot pages. A per-request figure needs shared and private memory counted on one basis, and no run here does that.
Memory, HTML extraction39.4 MiB273.7 MiB7.0x less memorynot measurednot measured: no run performs this operation.
Wall, screenshot1,148 ms637 ms1.8x slower549.4 ms at 4 vCPU770.3 and 712.6 ms in two runs at 2 vCPU, 580.2 ms at 8. From launching the VM to the screenshot returned, teardown excluded; the record’s wall_ms, which includes teardown, is 623.1 ms at 4 vCPU. Source: arms.cdp.blocking_ms in results/reqbench-20260902-025115-corpus-c4/analysis.json.
Wall, HTML extraction820 ms472 ms1.7x slowernot measurednot measured: the harness has an HTML branch, and no run selects it.
Web-platform tests215,000+ (the post’s heading); 731,247 of 762,891 subtests passing (kitesurf.dev, 2026-09-26)full enginenot publishedfull engineKitesurf’s count is from the post’s section “Kitesurf passes 215,000+ WPT tests and growing”. Cloudflare’s docs say over 235,000 subtests, and kitesurf.dev/wpt reports 731,247 of 762,891 on a pinned corpus. The sources do not explain the different totals. This row is not part of Cloudflare’s table. We run unmodified Chromium, and no Chromium count is quoted.

Cloudflare’s rows: verified on 2026-08-30 against developers.cloudflare.com/browser-run/kitesurf, which carries the same table as blog.cloudflare.com/kitesurf. Their figures are medians of five Browser Run quick-action runs. The web-platform row and the later benchmark below use the sources cited beside them.

Cloudflare’s later benchmark. Cloudflare’s benchmark page, kitesurf.dev/benchmarks, read on 2026-09-26, reports its latest run at 2026-09-25 01:17 UTC and states its method: 38 entries (the 14 corpus URLs and 24 more, one of which, relic.so/changelog, appears twice), 15 recorded renders attempted for each after two warm-ups, and three arms. For a screenshot of a typical page it reports a median wall time of 1.18 s for Kitesurf, 654 ms for Browser Run (Chrome reused between pages) and 1.09 s for Chrome (cold start), a fresh Chrome for every render whose launch is left out of the timer; in that run’s raw data the launch has a median of 215 ms across the pages’ screenshot medians. fcvm’s figures include the restore. That page sums Kitesurf’s CPU and memory across its Worker components, measures the Chrome arms’ memory as PSS across Chrome’s processes and Browser Run’s CPU at container level, and adds the Node driver to the cold Chrome’s CPU. No corpus run here measures CPU or memory for fcvm.

What one request does

  1. Golden, once. A VM boots Chromium, warms it on a local page, navigates to about:blank and is snapshotted. Every request restores this snapshot.
  2. Restore. A new Firecracker process restores a copy-on-write clone of the golden. Guest RAM is filled on first touch from the shared in-memory copy of the snapshot, and pages recorded from earlier clones start being mapped in after the handshake, before the VM resumes.
  3. Wait for the target. The harness polls Chromium’s /json/list every 50 ms until a page target appears.
  4. Connect and navigate. WebSocket upgrade and Page.enable, then Page.navigate and a wait for Page.loadEventFired. The records do not split this into fetch, script and layout.
  5. Screenshot. Page.captureScreenshot, JPEG quality 80. The clock stops when the driver returns, after it decodes the JPEG, reads the navigation timing and closes the connection.
  6. Destroy, not timed. The clone and everything it wrote are discarded after the answer (median 63.7 ms at 4 vCPUs).
The life of one requestOrder of events in fcvm’s restore path and the harness’s driver. The harness starts polling after step 4, which can come before the snapshot is loaded, so its polls can overlap steps 5 to 9; the records do not show how much. Replay (step 8) can continue after the resume (step 9). Demand faults (step 10) continue for the rest of the request. The clock runs from step 1 to step 13.
Harness → fcvm and pasta1fcvm snapshot run: the clock starts
fcvm and pasta2CoW disk and network namespace
fcvm and pasta → Clone VM3start Firecracker; start pasta, whose port listens from here
Harness → fcvm and pasta4TCP connect to the forwarded port: pasta accepts
fcvm and pasta → Clone VM5load the snapshot, not resumed
Memory server → Clone VM6once Firecracker connects: the shared memfd
Clone VM → Memory server7Firecracker’s memory mappings and userfaultfd
Memory server → Clone VM8start replaying recorded pages in chunks of up to 2 MiB
fcvm and pasta → Clone VM9patch the disk; resume the VM
Clone VM → Memory server10demand faults, from resume on, served before each replay chunk
Harness → Clone VM11after step 4, poll /json/list every 50 ms until a page target is listed; then Page.navigate
Clone VM → Harness12load event; screenshot JPEG
Harness13decode and check the JPEG, read navigation timing, close CDP: the clock stops
Harness → fcvm and pasta14destroy the clone (not timed)
Where the time goes, mean request at 4 vCPUsMeans over the 202 requests, which add up per request (the stages sum to 822.3 of the 822.8 ms mean). The median request is 549.4 ms; the slow news sites pull the mean up, mostly through navigate.
then teardown, after the answer and not timed: 84.0 ms mean
Port accepts
launch to the first TCP connect on the forwarded DevTools port
45.4 ms
Wait for target
the rest of the restore (snapshot load, memory handshake, disk patch, resume) and Chromium listing a page target, polled every 50 ms
152.9 ms
Connect
TCP connect, WebSocket upgrade, Page.enable
5.5 ms
Navigate
Page.navigate until the load event
537.7 ms
Screenshot
Page.captureScreenshot, JPEG quality 80
76.3 ms
After the JPEG
decode and check the JPEG, read the navigation timing
4.5 ms

Record: results/reqbench-20260902-025115-corpus-c4/reqbench.jsonl, the 202 measured cdp requests. The request is blocking_ms and port accepts is spawn_to_port_ms. The other parts are fields of render.stages: resolve_ms; tcp_ms, upgrade_ms and enable_ms; navigate_ms; screenshot_ms; decode_ms, nav_timing_ms and idle_ms. Teardown is teardown.teardown_total_ms.

Stage medians at 4 vCPUs, in request orderEach stage’s median over the 202 requests. Medians of separate stages do not add up to the request median.
Port accepts44.7 ms
Wait for target153.1 ms
WebSocket upgrade2.6 ms
Page.enable2.8 ms
Navigate301.0 ms
Screenshot67.1 ms

The first stage ends 44.7 ms after the harness launches the VM, at its first successful TCP connect to the forwarded DevTools port. That port is pasta’s host-side listener, which starts before the snapshot is loaded and accepts without the guest, so this stage does not show that the guest has restored. Any restore work left after that connect, and Chromium answering /json/list, fall in the next stage, the 153.1 ms wait for the target. The records do not mark when the restore completes.

The harness polls every 50 ms, so the wait moves in 50 ms steps. At 4 vCPUs, 18 requests found a target on the third poll, 172 on the fourth and 12 on the fifth: in the median request the first three polls, the last about 100 ms after the port accepted, got no page target. The records do not say whether a poll failed to connect or found no page. At most one 50 ms step is poll granularity; per request, the stage less its 50 ms sleeps has a median of 3.2 ms. A readiness notification in place of the poll would save at most one step per request. That change has not been measured.

Navigate, 301.0 ms, is the largest stage median.

One snapshot, many clones

Two mechanisms shape what a VM per request costs. The first is memory sharing: in minor mode, which the corpus runs used, the pages a clone only reads are a read-only view of one copy of the snapshot. The second is profiling and replay: the memory server records which pages clones fault and starts mapping that set into each new clone after the handshake, before it resumes.

Memory in minor modeSchematic, not to scale. From src/uffd/server.rs and the Firecracker fork’s restore backend.
memory.binthe golden’s guest memory, on disk
copied once when the memory server starts; with 4 KiB pages, all-zero pages stay holes
One sealed memfd per memory serverread-only, held while that server runs, shared by all its clones
each clone maps it privately; a fault installs a read-only mapping of the shared page
Clone 1
pages it reads: the shared copy
pages it writes: its own copy
own vCPUs, disk and network
Clone 2
pages it reads: the shared copy
pages it writes: its own copy
own vCPUs, disk and network
Clone 3
pages it reads: the shared copy
pages it writes: its own copy
own vCPUs, disk and network
Clone N
pages it reads: the shared copy
pages it writes: its own copy
own vCPUs, disk and network

A page one clone reads costs no memory in the others: every clone maps the same physical page. A page a clone writes is copied into that clone on its first write and discarded with it. The snapshot never changes. In copy mode, the default for fcvm snapshot serve, each clone instead gets its own copy of every page it touches.

Profiling and replayHow the memory server learns which pages clones touch. From src/uffd/working_set.rs and src/uffd/prefetch.rs.
  1. A clone runs. Each demand fault in its first 300 s (the default window; for these requests, the whole render) marks every 4 KiB granule the faulting page covers in the clone’s bitmap.
  2. The clone ends. Its bitmap is ORed into the server’s in-memory record of every clone so far. A background writer tries to save the grown record beside the snapshot as memory.bin.working-set; it can merge or skip saves, and the in-memory record serves the next clone either way.
  3. The next clone restores. After the handshake the server maps the recorded pages in chunks of up to 2 MiB, serving any page the guest is waiting on before each chunk. The VM resumes without waiting, so replay can overlap the guest running; when it does, the guest’s own faults are served before each chunk.

A record only says which pages to map; the bytes always come from the snapshot being served, so a stale record wastes work and cannot corrupt a guest. A new snapshot or a changed config starts an empty record.

Replay is measured once. On the synthetic page, in copy mode on one golden, the run with replay had a request median of 477.9 ms and the run without it 762.6 ms, with navigate at 184.1 against 390.6 ms (details and provenance). The corpus campaign turns replay on by default; the run records do not store the value, and no run compares replay on and off in minor mode or on the corpus.

On a large host

Every corpus run on this page sent one request at a time. The three ladder runs used for the arithmetic below ran on an AWS c8g.metal-48xl with 192 vCPUs and 377 GiB. None of the runs holds many clones at once, so none measures how many renders the host sustains or how latency holds up while it does. What the records do give is the host and the time each request held a clone, from launch through teardown. The chart below is arithmetic from those two, not a measurement.

If every guest vCPU had a host vCPU to itself (arithmetic)Clones in flight: 192 host vCPUs divided by guest vCPUs. Requests per second: clones in flight divided by the mean time a request held its clone, launch to teardown, in the serial runs. Hatched to mark arithmetic.
  • arithmetic from the records, not measured
2 vCPU: 96 in flight86.2 req/s
4 vCPU: 48 in flight52.7 req/s
8 vCPU: 24 in flight25.6 req/s

Records: mean wall_ms of the measured cdp requests in results/reqbench-20260902-023115-corpus-c2/reqbench.jsonl, results/reqbench-20260902-025115-corpus-c4/reqbench.jsonl, results/reqbench-20260902-031115-corpus-c8/reqbench.jsonl (1,113.8, 911.4, 936.7 ms); host vCPUs are nproc in results/reqbench-20260902-025115-corpus-c4/hostinfo.json.

The arithmetic assumes two things no corpus record checks: that a clone uses no more host CPU than its guest vCPUs, although the host also runs an fcvm process and a pasta for each clone and one memory server that serves all of them, and that a request takes as long with 96 clones in flight as it did alone. The one sustained-load run of minor mode (2026-08-08, synthetic page, 2-vCPU guests on a 64-vCPU host; results/20260808-corrected/corrected.json, sustained) saw the median of completed requests rise from 956.7 ms at a 1 request-per-second target to 1,223.6 ms at an 8 request-per-second target; the 8 request-per-second cell launched 459 requests and completed 458. In this arithmetic, 8 guest vCPUs halve the CPU slots from 48 to 24, and they did not shorten the request median in these runs; no record establishes how many clones the host can hold.

Memory per additional concurrent clone, synthetic pagefcvm: slope of cgroup memory.current over 1 to 16 clones held at once, each idle after one render of medium.html; 2 GiB guests with 2 vCPUs on a 64-vCPU host, 2026-08-08, before replay existed. The memory server sat outside the clones’ cgroups, so the shared snapshot is not in these figures. Warm containers: slope over 2 to 16 containers on the host, each idle after its start-up render of a warm-up page; they did not render medium.html. Each ± is the standard error of a least-squares fit over 15 clone counts per fcvm mode, or 12 container counts.
  • fcvm, one clone per request
  • warm Chromium containers on the host
minor mode132.5 ± 1.0 MiB
file-backed143.5 ± 0.4 MiB
copy mode257.8 ± 1.1 MiB
warm container pool156.5 ± 5.0 MiB

Fitted cgroup memory, intercept + N × slope, and clones or containers per GiB of it

SetupInterceptFitted at N = 16Density at N = 16
Minor mode35.0 ± 7.8 MiB2,154 MiB7.6 per GiB
File-backed17.3 ± 3.4 MiB2,314 MiB7.1 per GiB
Copy mode29.5 ± 9.1 MiB4,155 MiB3.9 per GiB
Warm container pool249.8 ± 46.2 MiB2,754 MiB5.9 per GiB

Record: results/20260808-corrected/corrected.json, density and host_pool fits. bench/chromium/REVIEW.md keeps its 2 MB hugepage cell out of use, because that basis cannot see the guest’s hugepages.

On the synthetic page, sharing is what separates the modes: a minor-mode clone added 132.5 MiB, about half of copy mode’s 257.8 MiB, and less than one more warm Chromium container at 156.5 MiB, which had rendered only its warm-up page. These are guests of a different size, on a different host, before replay, so they do not carry over to the corpus. What would settle the question on the corpus, with targets that exist today:

Guest vCPU count

The vCPU count is fixed in the golden, so each rung of the ladder has its own golden. The three rungs ran back to back on one host with one build.

Request median by guest vCPUsBar: median. Whisker: within-run interval, the 95% bootstrap confidence interval of the median from that run’s 202 requests. It does not include run-to-run variance.
2 vCPU770.3 ms
4 vCPU549.4 ms
8 vCPU580.2 ms
2 vCPU, 08-30712.6 ms

In these single runs, going from 2 to 4 vCPUs lowered the median: 549.4 ms is 28.7% below 770.3, and every one of the 14 sites had a lower median at 4 than at 2. At 8 vCPUs the median was 580.2 ms. The 4 and 8 vCPU medians each sit inside the other’s interval, and 8 vCPUs had the higher median on 11 of the 14 sites.

The 4-to-8 difference is not in the rendering. Each request’s recorded stages, including a TCP connect, the JPEG decode and the navigation-timing read that the table leaves out, add up to its time with a median remainder under 1 ms, so a change in the mean splits by stage. From 4 to 8 vCPUs the mean request time rose 21.9 ms: the wait for the target added 53.1 ms, while navigate took 26.8 ms less and screenshot 6.7 ms less; the other stages and overhead add 2.3 ms. The wait was longer at 8 vCPUs on all 14 sites, a median of 5 polls instead of 4. 4 vCPUs has the lower request median in these runs, but each rung ran once and the stage where 8 vCPUs loses moves in 50 ms poll steps, so the runs do not show that 4 is faster than 8. More vCPUs per render also means fewer renders per host, and no run here measures density.

No configuration was measured twice under the same conditions. The two 2 vCPU runs, on different goldens, fcvm builds, container images and host boots, read 770.3 and 712.6 ms, 57.7 ms apart; that is the only view of run-to-run spread here. The host was not equally quiet at every rung. The quiet-host gate is checked only at start, with a limit of 2.0; during the runs the 1-minute load on this 192-vCPU host peaked at 2.62, 18.3 and 16.97. The no-render arm read 44.9, 45.1 and 45.8 ms. No run measures whether that load changed the render stages.

Ladder medians (ms)

Guest vCPUsRequest medianWithin-run intervalNo-render arm
2770.3596.2–807.844.9
4549.4467.9–632.445.1
8580.2520.8–645.945.8
2, 08-30712.6610.5–808.544.8

Stage medians by vCPU count (ms)

Stage2 vCPU4 vCPU8 vCPU2 vCPU, 08-30
Port accepts43.144.745.544.5
Wait for target152.9153.1204.6154.3
WS upgrade6.12.63.16.2
Page.enable2.52.83.84.3
Navigate414.1301.0267.2409.1
Screenshot81.467.164.180.3

Records: results/reqbench-20260902-023115-corpus-c2, results/reqbench-20260902-025115-corpus-c4, results/reqbench-20260902-031115-corpus-c8, indexed by results/campaign-20260902-box2-ladder-summary.json; the 2026-08-30 run is results/reqbench-20260830-171007-corpus, indexed by results/campaign-20260830-box2-summary.json.

Per-site latency

Median per site at 4 vCPUs14 or 15 renders per site. The thin grey bar is the same site’s median at 2 vCPUs.
  • 4 vCPUs
  • 2 vCPUs
elmundo.es2,945.2 ms
rtp.pt/noticias2,195.0 ms
theguardian.com913.7 ms
developers.cloudflare.com656.5 ms
en.wikipedia.org651.2 ms
developer.mozilla.org643.3 ms
blog.cloudflare.com577.8 ms
TodoMVC, Angular478.5 ms
TodoMVC, ES6432.7 ms
TodoMVC, Vue429.8 ms
TodoMVC, React427.4 ms
TodoMVC, Preact411.3 ms
news.ycombinator.com395.8 ms
example.com326.2 ms

The simple pages (example.com, Hacker News and the five TodoMVC apps) take 326.2 to 478.5 ms at 4 vCPUs. The documentation and reference sites take 577.8 to 656.5 ms. The three news sites take 913.7 to 2,945.2 ms.

Per-site medians at 2, 4 and 8 vCPUs
Site2 vCPU4 vCPU8 vCPU
elmundo.es3,716.82,945.22,907.2
rtp.pt/noticias2,576.92,195.01,945.8
theguardian.com1,181.5913.7894.9
developers.cloudflare.com876.7656.5704.1
en.wikipedia.org817.3651.2700.3
developer.mozilla.org830.4643.3651.0
blog.cloudflare.com809.6577.8620.7
TodoMVC, Angular616.7478.5519.9
TodoMVC, ES6538.5432.7483.2
TodoMVC, Vue546.2429.8476.2
TodoMVC, React558.0427.4467.7
TodoMVC, Preact527.9411.3455.6
news.ycombinator.com505.3395.8441.8
example.com416.2326.2388.6

Medians in ms, from per_url in each ladder run’s analysis.json.

Correction

An earlier corpus series, from 2026-08-16, is withdrawn. Before fcvm 90733b854e, pasta redirected the guest’s DNS to the host’s resolver, so those runs resolved third-party hosts on the live internet and waited on them. Every corpus run on this page passed the DNS checks described under Method and gates. The withdrawn records and their evidence are in REVIEW.md.

Memory configuration and VM overhead (synthetic page)

Synthetic page

This section uses medium.html, a small deterministic page (1.4 KB of HTML, one stylesheet, 806 B of script, four images) served with Cache-Control: no-store, on the fixture host, with 1024 MiB, 2 vCPU guests except where a chart says otherwise. It exists to compare configurations. Absolute numbers move between goldens that also ran different harness revisions: the same copy, 4 KB configuration read 413.2 ms on the first golden and 532.5 ms on the second. So configuration comparisons use pairs measured on the same golden; the VM against container chart sets two separately measured setups side by side. None of these numbers is comparable to the corpus results above. The first two fixture goldens (RB, AB) were warmed on a page that shared medium.html’s stylesheet and one image; PF and the later goldens share no byte resources with it. Browser and system-font caches stay warm in every golden.

Working-set prefetch, same goldenRequest median, direct CDP, 202 requests each. Which run had prefetch on is recorded in commit e78350c5, not in the run records.
Prefetch off762.6 ms
Prefetch on477.9 ms

The memory server records the pages each clone faults and populates that set in the next clone at restore, in bounded chunks interleaved with the guest’s own faults. On the same golden, the run with prefetch on has a median 284.7 ms (37%) lower than the run with it off. Both runs used one sealed harness that passes the prefetch setting to the memory server explicitly, so the binary’s default could not apply; neither record stores the value passed. Commit e78350c5, which added the explicit setting, names the first run on and the second off. The port-accept stage is 59.3 ms in both, which says nothing about where the population cost lands, because that stage can end before the snapshot is loaded.

userfaultfd faults per requestMedian userfaultfd faults per request, from the memory server’s fault trace. The run started the memory server without a prefetch setting, so prefetch was at its default, on, unless the environment overrode it; the record does not store the value. Six requests per configuration; the median is over those whose fault trace was unambiguous: 4 in copy mode, 5 in minor mode, 5 in minor mode on 2 MB pages. This run restored bench.sh’s snapshots of localhost/chromium-bench with 2 GiB guests (guest_bytes in its summary.json); its vCPU count and host are not recorded, and bench.sh defaulted to 2 vCPUs.
4 KB pages, copy12,503
4 KB pages, minor5,161
2 MB pages, minor0

In copy mode every faulting page is copied into the clone. In minor mode the page already sits in the shared memfd and is only mapped. The median time of the ioctl that resolves a fault was 2.3 µs in copy mode and 1.5 µs in minor mode; that is not the whole cost of a fault, and the memory server’s CPU per fault was about 27 µs in both. The five 2 MB minor traces recorded no userfaultfd faults. A file-backed restore uses the kernel page cache and has no userspace fault path.

Restores of one snapshot fault in overlapping page sets: the mean pairwise Jaccard similarity of two restores’ faulted pages is 0.65 in minor mode and 0.69 in copy mode, lowest pair 0.54. Per request, 45 to 59% of faults fall in runs of 4 or more consecutive pages.

Page sizecopyminorGolden
4 KB532.5607.9shared by the pair (AB)
2 MB401.9374.9shared by the pair (HK, HM)

Request medians in ms, direct CDP. On 4 KB pages minor mode is 14% slower than copy. On 2 MB pages minor mode is faster in all three render arms: direct CDP 374.9 against 401.9 ms, fast teardown 380.8 against 401.0, in-guest exec 582.9 against 611.9. In the direct CDP arm the largest stage-median difference is the wait for the target (130.7 against 156.7 ms); navigate is 109.6 against 109.7 ms, and screenshot is slower in minor mode (81.9 against 78.9 ms). HM and HK share one golden. Its 2 MB pages are recorded as hugepages: true in each run’s clone-state capture, a raw file that is not committed; the committed analysis.json holds only the snapshot name and config digest. Their harness (b72e516a, which matches both runs’ harness_sha256) passes the prefetch setting to the memory server explicitly, on unless the environment overrides it. Neither record stores the value passed, and the run log that echoes it is not committed.

VM against a warm host container, stage mediansThe VM is the prefetch-on run above. The container is a warm Chromium on the fixture host with no VM and no CPU limit. WebSocket upgrade and Page.enable, under 10 ms in both, are left out.
  • microVM per request
  • warm host container
Port accepts
VM59.3 ms
containernone, already running
Wait for target
VM139.8 ms
container25.8 ms
Navigate
VM184.1 ms
container27.7 ms
Screenshot
VM81.0 ms
container77.9 ms

The request medians are 477.9 ms in the VM, from launching the VM to the screenshot returned, and 133.0 ms in the container, the driver’s own time from target lookup to its return. The container run’s wall time per request was 193.7 ms. The 60.7 ms between the two is the per-request process wrapper around the driver, including Python start-up, which the VM run, driving it in-process, does not pay; no run isolates its parts. The largest stage-median differences are the wait for the target (+114.0 ms) and navigate (+156.4 ms); screenshot differs by +3.1 ms. The VM had 2 vCPUs and the container had no CPU limit, so these differences also include a difference in available CPU. Both runs render medium.html with cdpdrive.py. The VM run seals the driver in its harness hash; the host record names cdpdrive.py without a hash and used a different build of the container image. These are two separately measured setups, not a split of the cost into virtualization and memory.

All fixture configurations
Run and configurationNo-render arm p50Direct CDP p50Fast teardown p50
RB-uffd: copy, 4 KB, first golden59.3413.2413.0
RB-file: file-backed, 4 KB, first golden58.5458.9453.8
AB-copy: copy, 4 KB, second golden59.3532.5532.7
AB-minor: minor, 4 KB, second golden59.3607.9604.8
PF-on: copy, 4 KB, prefetch on, third golden59.3477.9479.4
PF-off: copy, 4 KB, prefetch off, third golden59.4762.6760.5
HM: minor, 2 MB, hugepage golden39.9374.9380.8
HK: copy, 2 MB, hugepage golden39.9401.9401.0
NC: HM with Chromium’s profile and disk cache kept off the rootfs, no exec arm, new golden39.9372.5371.6
FG: NC’s image on another new golden, harness with a 2 ms port probe (commit 1b115d55)40.9348.7347.4
HC: warm host container, no VM; the driver’s own time, target lookup to its returnnot applicable133.0not applicable

Medians in ms, 202 measured requests per run and arm (200 for HC). No arm failed: 0 in 204 or 232 attempts per arm, an exact 95% interval of 0 to at most 1.79%, and HC 0 in 202, 0 to 1.81%. Direct CDP: the harness drives Chromium over CDP from the host. Fast teardown: the same render and the same timed span; only the teardown after the answer differs, one SIGKILL carried to fcvm’s children by the pdeathsig chain in place of fcvm’s SIGTERM-and-await sequence. Exec arm: the retired in-guest driver. The no-render arm stops when the forwarded port accepts, so its figures sit on the port probe’s grid: before FG the probe backed off to about 17 to 20 ms between tries, and FG ran with a 2 ms cap. Both come from commit 1b115d55, which names FG as the fine-grid run; FG’s recorded source revision still has the 20 ms cap, and its harness hash matches no committed tree. The 4 KB and 2 MB rows come from different goldens, so their no-render figures are not a page-size comparison; an earlier page-size comparison drawn from them was withdrawn. NC and FG differ in the golden and in the harness, so their difference is attributed to neither. RB and AB ran on goldens whose warm-up page shared styles.css and img1.png with medium.html.

Not measured

Method and gates

To reproduce the corpus campaign (golden, verification, DNS diagnostics, measured run) from a clean checkout:

make bench-chromium-corpus

The guest vCPU count is part of the golden, so each rung of a ladder needs its own golden. The phase-by-phase targets and knobs are in bench/chromium/AGENTS.md.

Firecracker builds

fcvm drives a fork, ejc3/firecracker. rootfs-config.toml at each run’s fcvm source revision selects a branch. Setup builds or reuses the binary for that branch’s head when it can reach the remote, and falls back to a cached binary when it cannot. The corpus and fixture runs used different branches. No record stores the Firecracker commit a run built.

Corpus runs

Branch agent/uffd-minor, whose current head is 32f9ef433, committed 2026-08-14, before all four published corpus runs: upstream main as of 2026-08-07 (03b096f3b) plus two changes. The ladder’s hostinfo.json records the binary’s content hash.

Upstream’s own fixes for the serial interrupt on restore and the dirty-page dump are in this base, so the branch does not carry the fork’s earlier patches for them.

Fixture runs

Branch bump-vsock-max-connections, whose current head is 014c1c15, committed 2026-08-07, before every fixture run: upstream main as of 2026-02-17 plus four changes. Two are earlier versions of the changes above: the vsock connection limit raised from 1023 to 16384 and the kill queue from 128 to 2048, with the receive queue left at 256 (5ae6d15ed), and the first revision of the minor-fault backend, which passes the sealed memfd in the handshake but predates the bounded-handshake commits. The other two are fixes upstream later made separately:

Restored guests need their monotonic clock to continue from the snapshot instant, or every armed timer fires at once. On these plain-EL1 guests the host kernel handles it: since the March 2023 arm64 KVM counter-offset series, KVM recomputes the VM-wide counter offset when a restore replays the saved counter registers. The timer storm fcvm fixed for nested virtualization (#630) needs a virtual-EL2 guest and does not apply to either build.

Records

Every fcvm figure on this page comes from a committed record under bench/chromium/results/, except where a statement names another source: the bridged-network refusal, the replay equivalence result, the prefetch and port-probe settings taken from commit messages, the hugepage setting taken from raw captures, and the Firecracker commits, which no record stores. The reqbench runs record content hashes of the harness and of the runtime bundle holding fcvm and fc-agent, the snapshot generation and its config digest. HC records its configuration and host kernel in run.json; FB records its configuration and guest size only, with no host, kernel or prefetch setting. Cloudflare’s benchmark rows are quoted from their post; the web-platform and later-benchmark figures come from Cloudflare’s docs and kitesurf.dev, cited where they appear, and the Chrome launch median is computed from kitesurf.dev’s raw data.

IdWhatRecord
CL2, CL4, CL8Corpus vCPU ladder, 2026-09-02, one golden per rung, DNS-verifiedresults/reqbench-20260902-023115-corpus-c2, results/reqbench-20260902-025115-corpus-c4, results/reqbench-20260902-031115-corpus-c8; index results/campaign-20260902-box2-ladder-summary.json
CVCorpus, 2 vCPU, 2026-08-30, DNS-verifiedresults/reqbench-20260830-171007-corpus; index results/campaign-20260830-box2-summary.json
RB-uffd, RB-fileFixture, first goldenresults/reqbench-86d0f4aca0ca486b8b1a9239792a2d7d, results/reqbench-ccd7655151d94a6ea47fba8aa23f352d
AB-copy, AB-minorFixture, second goldenresults/reqbench-c61bd9e8e4764a23aca3d0ab7ca07a3a, results/reqbench-8fbec71a1863427fa4daddc9f1f53af1
PF-on, PF-offFixture, third golden, one harness passing the prefetch setting explicitly; the value is in neither record, and commit e78350c5 names the first run on and the second offresults/reqbench-9c23e9d7da424351a390b5c1ddfd8e1d, results/reqbench-134b408e1a3d42feb866bb4c5fed774c
HM, HKFixture, 2 MB pages, minor and copy, shared goldenresults/reqbench-20260814-001144-uffd, results/reqbench-20260814-002056-uffd
NC, FGFixture, Chromium writes kept off the rootfs, on two goldens and two harness revisionsresults/reqbench-20260814-024029-uffd, results/reqbench-20260814-035757-uffd
HCWarm host container, no VMresults/hostcdp-bc17d3ffb64c46ab8d2ef51c0043fb9b
FBuserfaultfd fault counts on bench.sh snapshots, 2 GiB guests; prefetch setting not recordedresults/faultbench-0813-073507-690998