One microVM per page: Chromium on Firecracker snapshots
In August Cloudflare introduced Kitesurf, a browser for agents that runs on Workers, one V8 isolate per navigation, and uses native Rust compiled directly to WebAssembly where possible. Their starting point is that Chromium is too heavy: it uses “so much memory and compute that providing every agent with its own instance is prohibitively expensive”. This page looks at the other answer: keep unmodified Chromium and give each request its own microVM, restored from a snapshot in which Chromium is already booted and warm, with the memory server recording which pages clones touch and mapping them into the next one. It measures how long that takes on Cloudflare’s 14-URL corpus. What it costs per request in CPU and memory, the axis Cloudflare’s post is about, is not measured here.
Setup
| Item | Value |
|---|---|
| Corpus host | AWS c8g.metal-48xl (Graviton4, Neoverse-V2, 192 vCPUs, 377 GiB), Ubuntu 24.04, kernel 6.17.0-1019-aws. Recorded in results/reqbench-20260902-025115-corpus-c4/hostinfo.json. The 2026-08-30 run used the same kernel; its instance is not recorded. |
| Fixture host | An aarch64 host on kernel 7.0.14-fcvm-cd6cd2b4b52e, recorded in each reqbench fixture run’s analysis.json and in HC’s run.json. Its instance type, CPU model and core count were not recorded, and the fault-count run (FB) records no host at all. Used only for the synthetic-page runs. |
| Guest | Corpus runs: 1024 MiB RAM, 2, 4 or 8 vCPUs, Linux 6.18.44. Reqbench fixture runs: 1024 MiB, 2 vCPUs. The fault-count run (FB) used other guests, described with its chart. Headless Chromium from Debian bookworm’s unpinned chromium package, in a Podman container tagged localhost/chromium-bench-req. The tag is mutable; the records pin image sha256:5b870d814e5d… for the vCPU ladder and sha256:334fc21f7c8c… for the 2026-08-30 run (image_id in each reqbench.jsonl meta line). No record keeps the Chromium version. |
| Memory | Corpus runs: clone RAM is filled on first touch by fcvm’s userfaultfd server in minor mode. The server copies the snapshot’s memory once into a sealed in-memory file (a memfd) that all clones share; clean pages map that file, and a page a clone writes becomes private to it. The campaign turns working-set prefetch on unless the environment overrides it (corpus_campaign.sh); the records do not store the value. The synthetic-page runs also test copy mode and file-backed restores. |
| Network | Rootless (pasta). The guest’s DNS resolves every corpus host to a replay server on the host. |
| Workload | The 14 URLs of kitesurf.cloudflare.app/corpus.txt, captured once from the live sites. The captured response bodies are served from the host. URL matching, headers and some statuses differ from a live origin. An uncaptured GET, HEAD or POST gets an empty 404, and a CORS preflight gets 204. 28 warmup requests, then 202 measured, cycling the URLs uniformly (14 or 15 renders each). |
| Operation | Restore a clone, wait for Chromium’s DevTools target, navigate, Page.captureScreenshot (JPEG, quality 80). The clone is destroyed after the JPEG is returned. |
| Latency | The request’s blocking time, arms. in analysis.json: from just before the harness launches the VM until its CDP driver returns. After the screenshot arrives the driver decodes and checks the JPEG, reads the navigation timing and closes the connection before returning; the decode and timing read take 1.9 ms median at 4 vCPUs. Teardown runs after the answer and is excluded (median 63.7 ms at 4 vCPUs). The record’s wall_ms includes it and reads 623.1 ms at 4 vCPUs. |
| Software | fcvm source 1e9e9b70 for the vCPU ladder, 55756858 for the 2026-08-30 run. Firecracker from a fork; the corpus and fixture runs used different branches of it, listed under Firecracker builds. |
- Golden
- The snapshot every request restores. A VM boots Chromium, warms it on an unrelated local page (one navigation and one screenshot), navigates to
about:blankand is snapshotted: memory, disk and device state. Regenerating it makes a new golden. The guest vCPU count is fixed in the golden. - vCPU ladder
- Three corpus runs on 2026-09-02 at 2, 4 and 8 guest vCPUs, one golden each, back to back on one host with one build.
- Within-run interval
- The 95% bootstrap confidence interval of a run’s median, computed from that run’s 202 requests. It does not include run-to-run variance.
- No-render arm
- Requests interleaved with the renders that restore a clone and stop the clock when its forwarded DevTools port accepts a connection, with no navigation. The drift gate watches it.
Two answers to one question
Cloudflare’s post asks how to give every agent its own browser. Kitesurf answers with a new engine built as several Workers components, each in its own V8 isolate: the Engine, which holds the CDP session and fetches the page’s document and scripts; PageScript, one isolate per navigation, which runs the page’s script; PageRenderer, which paints; and SandboxOutbound, the only path to the network. Where it can, it uses native Rust compiled to WebAssembly, with parts of Blitz and Stylo for HTML and CSS. fcvm keeps the engine and moves the isolation boundary out to a microVM, so that each page gets its own kernel under an unmodified Chromium.
KitesurfCloudflare; Rust compiled to WebAssembly, on Workers
fcvmunmodified Chromium in a Firecracker microVM per request
The two designs pay for isolation in different places. Kitesurf rebuilds the engine as Workers components, and Cloudflare’s post lists what it cannot do yet: play video, render WebGL, negotiate a bot-challenge handshake with real TLS fingerprints, or start a ten-minute authenticated session that requires persistent state. It implements a subset of CDP. fcvm keeps the whole engine and restores a VM for each request instead of booting one. At 4 guest vCPUs the recorded port-accept and target-wait stages sum to a median of 197.7 ms. The driver’s setup between them is not timed and the harness polls every 50 ms, so the moment the restored Chromium was ready is not recorded. It is a time, not a cost: what a restore costs in CPU and memory is not measured. For scale only, Cloudflare’s later benchmark leaves out the launch of a fresh Chrome, a median of 215 ms across its pages, on hardware it does not state. Cloudflare publishes no spin-up time for a Kitesurf isolate.
The same corpus
Cloudflare’s post reports “the medians of five Browser Run quick-action runs across a 14-URL corpus”, Kitesurf against Chromium in a warm pool. The corpus file, kitesurf.cloudflare.app/, is no longer served; the Wayback Machine holds identical copies from 2026-08-07 and 2026-08-14. We captured those 14 URLs and replayed them to a new microVM per request.
- Cloudflare Kitesurf, isolated per request
- Cloudflare Chromium, warm pool
- fcvm, new microVM per request
fcvm bars: from launching the VM to the screenshot returned, teardown excluded (Latency, under Setup). The 4 and 8 vCPU runs are from 2026-09-02.
These figures sit side by side; they are not a controlled comparison. The setups differ:
- Statistic and timer. Cloudflare’s figures are “the medians of five Browser Run quick-action runs”; the post does not say how each run combines its 14 URLs or what its timer covers. Ours is the median of 202 requests, from launch to the screenshot returned. Our mean at 4 vCPUs is 822.8 ms.
- Network. Cloudflare’s corpus names live public URLs, and the post does not state how benchmark traffic was controlled. Our corpus is replayed from the same host, so our numbers contain no internet round trips.
- Uncaptured requests. The replay answers an uncaptured GET, HEAD or POST with an empty 404: 1,261 of them in the 4 vCPU run, median 0.2 ms at the server, longest 341 ms. A live load would fetch those resources and whatever they trigger. Which way this moves our times is not measured.
- Browser state. Their Chromium column is a warm pool; Cloudflare’s later benchmark page describes that arm as “Chrome reused between pages”. Each of our requests gets its own clone and discards it afterwards, so no browser state carries from one request to the next, as in their Kitesurf column. Unlike Kitesurf, each clone starts from a Chromium that was booted and warmed on an unrelated local page before the snapshot: neither a cold engine nor a browser that has seen the measured pages.
- Feature scope. Cloudflare’s post says Kitesurf cannot yet play video, render WebGL, negotiate a bot-challenge handshake with real TLS fingerprints, or start a ten-minute authenticated session that requires persistent state. We run unmodified Chromium.
- Hardware. Cloudflare does not state theirs. Ours is in Setup.
As numbers, our median is below their Kitesurf median at every vCPU count we ran, and below their warm Chromium median at 4 and 8 vCPUs but not at 2. Our means, 1,035.1, 822.8 and 844.7 ms at 2, 4 and 8 vCPUs, are above their warm Chromium figure at every count. Which of our statistics matches theirs is not known, and none of the differences above is priced, so neither ordering ranks the systems.
Cloudflare’s own table is a trade between cost and time. By their relative column, against their warm Chromium, Kitesurf uses 3.1 and 3.8 times less CPU and 4.7 and 7.0 times less memory, and is 1.8 and 1.7 times slower on wall time. The chart above shows only time, the axis on which Kitesurf is the slower of their two engines. We have no CPU or memory figure to set against the other rows.
Cloudflare’s published table, with fcvm’s figures
| Metric | Kitesurf | Cloudflare Chromium | Kitesurf, relative | fcvm | Notes |
|---|---|---|---|---|---|
| CPU, screenshot | 380 ms | 1,173 ms | 3.1x less CPU | not established | No whole-system CPU measurement of the corpus runs. |
| CPU, HTML extraction | 229 ms | 877 ms | 3.8x less CPU | not measured | not measured: no run performs this operation. |
| Memory, screenshot | 57.8 MiB | 271.0 MiB | 4.7x less memory | not established | Clones share clean snapshot pages. A per-request figure needs shared and private memory counted on one basis, and no run here does that. |
| Memory, HTML extraction | 39.4 MiB | 273.7 MiB | 7.0x less memory | not measured | not measured: no run performs this operation. |
| Wall, screenshot | 1,148 ms | 637 ms | 1.8x slower | 549.4 ms at 4 vCPU | 770.3 and 712.6 ms in two runs at 2 vCPU, 580.2 ms at 8. From launching the VM to the screenshot returned, teardown excluded; the record’s wall_ms, which includes teardown, is 623.1 ms at 4 vCPU. Source: arms. in results/reqbench-20260902-025115-corpus-c4/analysis.json. |
| Wall, HTML extraction | 820 ms | 472 ms | 1.7x slower | not measured | not measured: the harness has an HTML branch, and no run selects it. |
| Web-platform tests | 215,000+ (the post’s heading); 731,247 of 762,891 subtests passing (kitesurf.dev, 2026-09-26) | full engine | not published | full engine | Kitesurf’s count is from the post’s section “Kitesurf passes 215,000+ WPT tests and growing”. Cloudflare’s docs say over 235,000 subtests, and kitesurf.dev/wpt reports 731,247 of 762,891 on a pinned corpus. The sources do not explain the different totals. This row is not part of Cloudflare’s table. We run unmodified Chromium, and no Chromium count is quoted. |
Cloudflare’s rows: verified on 2026-08-30 against developers.cloudflare.com/browser-run/kitesurf, which carries the same table as blog.cloudflare.com/kitesurf. Their figures are medians of five Browser Run quick-action runs. The web-platform row and the later benchmark below use the sources cited beside them.
Cloudflare’s later benchmark. Cloudflare’s benchmark page, kitesurf.dev/benchmarks, read on 2026-09-26, reports its latest run at 2026-09-25 01:17 UTC and states its method: 38 entries (the 14 corpus URLs and 24 more, one of which, relic.so/changelog, appears twice), 15 recorded renders attempted for each after two warm-ups, and three arms. For a screenshot of a typical page it reports a median wall time of 1.18 s for Kitesurf, 654 ms for Browser Run (Chrome reused between pages) and 1.09 s for Chrome (cold start), a fresh Chrome for every render whose launch is left out of the timer; in that run’s raw data the launch has a median of 215 ms across the pages’ screenshot medians. fcvm’s figures include the restore. That page sums Kitesurf’s CPU and memory across its Worker components, measures the Chrome arms’ memory as PSS across Chrome’s processes and Browser Run’s CPU at container level, and adds the Node driver to the cold Chrome’s CPU. No corpus run here measures CPU or memory for fcvm.
What one request does
- Golden, once. A VM boots Chromium, warms it on a local page, navigates to
about:blankand is snapshotted. Every request restores this snapshot. - Restore. A new Firecracker process restores a copy-on-write clone of the golden. Guest RAM is filled on first touch from the shared in-memory copy of the snapshot, and pages recorded from earlier clones start being mapped in after the handshake, before the VM resumes.
- Wait for the target. The harness polls Chromium’s
/json/listevery 50 ms until a page target appears. - Connect and navigate. WebSocket upgrade and
Page.enable, thenPage.navigateand a wait forPage.. The records do not split this into fetch, script and layout.loadEventFired - Screenshot.
Page., JPEG quality 80. The clock stops when the driver returns, after it decodes the JPEG, reads the navigation timing and closes the connection.captureScreenshot - Destroy, not timed. The clone and everything it wrote are discarded after the answer (median 63.7 ms at 4 vCPUs).
fcvm snapshot run: the clock starts/json/list every 50 ms until a page target is listed; then Page.navigate- Port accepts
- launch to the first TCP connect on the forwarded DevTools port
- 45.4 ms
- Wait for target
- the rest of the restore (snapshot load, memory handshake, disk patch, resume) and Chromium listing a page target, polled every 50 ms
- 152.9 ms
- Connect
- TCP connect, WebSocket upgrade, Page.enable
- 5.5 ms
- Navigate
- Page.navigate until the load event
- 537.7 ms
- Screenshot
- Page.captureScreenshot, JPEG quality 80
- 76.3 ms
- After the JPEG
- decode and check the JPEG, read the navigation timing
- 4.5 ms
Record: results/reqbench-20260902-025115-corpus-c4/reqbench.jsonl, the 202 measured cdp requests. The request is blocking_ms and port accepts is spawn_to_port_ms. The other parts are fields of render.stages: resolve_ms; tcp_ms, upgrade_ms and enable_ms; navigate_ms; screenshot_ms; decode_ms, nav_timing_ms and idle_ms. Teardown is teardown.teardown_total_ms.
The first stage ends 44.7 ms after the harness launches the VM, at its first successful TCP connect to the forwarded DevTools port. That port is pasta’s host-side listener, which starts before the snapshot is loaded and accepts without the guest, so this stage does not show that the guest has restored. Any restore work left after that connect, and Chromium answering /json/list, fall in the next stage, the 153.1 ms wait for the target. The records do not mark when the restore completes.
The harness polls every 50 ms, so the wait moves in 50 ms steps. At 4 vCPUs, 18 requests found a target on the third poll, 172 on the fourth and 12 on the fifth: in the median request the first three polls, the last about 100 ms after the port accepted, got no page target. The records do not say whether a poll failed to connect or found no page. At most one 50 ms step is poll granularity; per request, the stage less its 50 ms sleeps has a median of 3.2 ms. A readiness notification in place of the poll would save at most one step per request. That change has not been measured.
Navigate, 301.0 ms, is the largest stage median.
One snapshot, many clones
Two mechanisms shape what a VM per request costs. The first is memory sharing: in minor mode, which the corpus runs used, the pages a clone only reads are a read-only view of one copy of the snapshot. The second is profiling and replay: the memory server records which pages clones fault and starts mapping that set into each new clone after the handshake, before it resumes.
src/uffd/server.rs and the Firecracker fork’s restore backend.A page one clone reads costs no memory in the others: every clone maps the same physical page. A page a clone writes is copied into that clone on its first write and discarded with it. The snapshot never changes. In copy mode, the default for fcvm snapshot serve, each clone instead gets its own copy of every page it touches.
src/uffd/working_set.rs and src/uffd/prefetch.rs.- A clone runs. Each demand fault in its first 300 s (the default window; for these requests, the whole render) marks every 4 KiB granule the faulting page covers in the clone’s bitmap.
- The clone ends. Its bitmap is ORed into the server’s in-memory record of every clone so far. A background writer tries to save the grown record beside the snapshot as
memory.; it can merge or skip saves, and the in-memory record serves the next clone either way.bin. working-set - The next clone restores. After the handshake the server maps the recorded pages in chunks of up to 2 MiB, serving any page the guest is waiting on before each chunk. The VM resumes without waiting, so replay can overlap the guest running; when it does, the guest’s own faults are served before each chunk.
A record only says which pages to map; the bytes always come from the snapshot being served, so a stale record wastes work and cannot corrupt a guest. A new snapshot or a changed config starts an empty record.
Replay is measured once. On the synthetic page, in copy mode on one golden, the run with replay had a request median of 477.9 ms and the run without it 762.6 ms, with navigate at 184.1 against 390.6 ms (details and provenance). The corpus campaign turns replay on by default; the run records do not store the value, and no run compares replay on and off in minor mode or on the corpus.
On a large host
Every corpus run on this page sent one request at a time. The three ladder runs used for the arithmetic below ran on an AWS c8g.metal-48xl with 192 vCPUs and 377 GiB. None of the runs holds many clones at once, so none measures how many renders the host sustains or how latency holds up while it does. What the records do give is the host and the time each request held a clone, from launch through teardown. The chart below is arithmetic from those two, not a measurement.
- arithmetic from the records, not measured
Records: mean wall_ms of the measured cdp requests in results/reqbench-20260902-023115-corpus-c2/reqbench.jsonl, results/reqbench-20260902-025115-corpus-c4/reqbench.jsonl, results/reqbench-20260902-031115-corpus-c8/reqbench.jsonl (1,113.8, 911.4, 936.7 ms); host vCPUs are nproc in results/reqbench-20260902-025115-corpus-c4/hostinfo.json.
The arithmetic assumes two things no corpus record checks: that a clone uses no more host CPU than its guest vCPUs, although the host also runs an fcvm process and a pasta for each clone and one memory server that serves all of them, and that a request takes as long with 96 clones in flight as it did alone. The one sustained-load run of minor mode (2026-08-08, synthetic page, 2-vCPU guests on a 64-vCPU host; results/20260808-corrected/corrected.json, sustained) saw the median of completed requests rise from 956.7 ms at a 1 request-per-second target to 1,223.6 ms at an 8 request-per-second target; the 8 request-per-second cell launched 459 requests and completed 458. In this arithmetic, 8 guest vCPUs halve the CPU slots from 48 to 24, and they did not shorten the request median in these runs; no record establishes how many clones the host can hold.
memory.current over 1 to 16 clones held at once, each idle after one render of medium.html; 2 GiB guests with 2 vCPUs on a 64-vCPU host, 2026-08-08, before replay existed. The memory server sat outside the clones’ cgroups, so the shared snapshot is not in these figures. Warm containers: slope over 2 to 16 containers on the host, each idle after its start-up render of a warm-up page; they did not render medium.html. Each ± is the standard error of a least-squares fit over 15 clone counts per fcvm mode, or 12 container counts.- fcvm, one clone per request
- warm Chromium containers on the host
Fitted cgroup memory, intercept + N × slope, and clones or containers per GiB of it
| Setup | Intercept | Fitted at N = 16 | Density at N = 16 |
|---|---|---|---|
| Minor mode | 35.0 ± 7.8 MiB | 2,154 MiB | 7.6 per GiB |
| File-backed | 17.3 ± 3.4 MiB | 2,314 MiB | 7.1 per GiB |
| Copy mode | 29.5 ± 9.1 MiB | 4,155 MiB | 3.9 per GiB |
| Warm container pool | 249.8 ± 46.2 MiB | 2,754 MiB | 5.9 per GiB |
Record: results/20260808-corrected/corrected.json, density and host_pool fits. bench/chromium/REVIEW.md keeps its 2 MB hugepage cell out of use, because that basis cannot see the guest’s hugepages.
On the synthetic page, sharing is what separates the modes: a minor-mode clone added 132.5 MiB, about half of copy mode’s 257.8 MiB, and less than one more warm Chromium container at 156.5 MiB, which had rendered only its warm-up page. These are guests of a different size, on a different host, before replay, so they do not carry over to the corpus. What would settle the question on the corpus, with targets that exist today:
- CPU per request.
make bench-chromium-corpuswith its default arms. Its fast-teardown arm records each child process’s CPU time; the published runs left that arm out. About 25 minutes per vCPU count. - Memory with many clones at once.
make bench-chromium-corpus-extra PHASES=memory MEM_NS=2,16,48,96, which interleaves N clones against N warm containers and reports shared and private memory on two bases. About one to two hours. - Throughput under load. No target drives the corpus at a fixed arrival rate yet;
make bench-chromium-scaledoes it for the synthetic page only.
Guest vCPU count
The vCPU count is fixed in the golden, so each rung of the ladder has its own golden. The three rungs ran back to back on one host with one build.
In these single runs, going from 2 to 4 vCPUs lowered the median: 549.4 ms is 28.7% below 770.3, and every one of the 14 sites had a lower median at 4 than at 2. At 8 vCPUs the median was 580.2 ms. The 4 and 8 vCPU medians each sit inside the other’s interval, and 8 vCPUs had the higher median on 11 of the 14 sites.
The 4-to-8 difference is not in the rendering. Each request’s recorded stages, including a TCP connect, the JPEG decode and the navigation-timing read that the table leaves out, add up to its time with a median remainder under 1 ms, so a change in the mean splits by stage. From 4 to 8 vCPUs the mean request time rose 21.9 ms: the wait for the target added 53.1 ms, while navigate took 26.8 ms less and screenshot 6.7 ms less; the other stages and overhead add 2.3 ms. The wait was longer at 8 vCPUs on all 14 sites, a median of 5 polls instead of 4. 4 vCPUs has the lower request median in these runs, but each rung ran once and the stage where 8 vCPUs loses moves in 50 ms poll steps, so the runs do not show that 4 is faster than 8. More vCPUs per render also means fewer renders per host, and no run here measures density.
No configuration was measured twice under the same conditions. The two 2 vCPU runs, on different goldens, fcvm builds, container images and host boots, read 770.3 and 712.6 ms, 57.7 ms apart; that is the only view of run-to-run spread here. The host was not equally quiet at every rung. The quiet-host gate is checked only at start, with a limit of 2.0; during the runs the 1-minute load on this 192-vCPU host peaked at 2.62, 18.3 and 16.97. The no-render arm read 44.9, 45.1 and 45.8 ms. No run measures whether that load changed the render stages.
Ladder medians (ms)
| Guest vCPUs | Request median | Within-run interval | No-render arm |
|---|---|---|---|
| 2 | 770.3 | 596.2–807.8 | 44.9 |
| 4 | 549.4 | 467.9–632.4 | 45.1 |
| 8 | 580.2 | 520.8–645.9 | 45.8 |
| 2, 08-30 | 712.6 | 610.5–808.5 | 44.8 |
Stage medians by vCPU count (ms)
| Stage | 2 vCPU | 4 vCPU | 8 vCPU | 2 vCPU, 08-30 |
|---|---|---|---|---|
| Port accepts | 43.1 | 44.7 | 45.5 | 44.5 |
| Wait for target | 152.9 | 153.1 | 204.6 | 154.3 |
| WS upgrade | 6.1 | 2.6 | 3.1 | 6.2 |
| Page.enable | 2.5 | 2.8 | 3.8 | 4.3 |
| Navigate | 414.1 | 301.0 | 267.2 | 409.1 |
| Screenshot | 81.4 | 67.1 | 64.1 | 80.3 |
Records: results/reqbench-20260902-023115-corpus-c2, results/reqbench-20260902-025115-corpus-c4, results/reqbench-20260902-031115-corpus-c8, indexed by results/campaign-20260902-box2-ladder-summary.json; the 2026-08-30 run is results/reqbench-20260830-171007-corpus, indexed by results/campaign-20260830-box2-summary.json.
Per-site latency
- 4 vCPUs
- 2 vCPUs
The simple pages (example.com, Hacker News and the five TodoMVC apps) take 326.2 to 478.5 ms at 4 vCPUs. The documentation and reference sites take 577.8 to 656.5 ms. The three news sites take 913.7 to 2,945.2 ms.
Per-site medians at 2, 4 and 8 vCPUs
| Site | 2 vCPU | 4 vCPU | 8 vCPU |
|---|---|---|---|
| elmundo. | 3,716.8 | 2,945.2 | 2,907.2 |
| rtp. | 2,576.9 | 2,195.0 | 1,945.8 |
| theguardian. | 1,181.5 | 913.7 | 894.9 |
| developers. | 876.7 | 656.5 | 704.1 |
| en. | 817.3 | 651.2 | 700.3 |
| developer. | 830.4 | 643.3 | 651.0 |
| blog. | 809.6 | 577.8 | 620.7 |
| TodoMVC, Angular | 616.7 | 478.5 | 519.9 |
| TodoMVC, ES6 | 538.5 | 432.7 | 483.2 |
| TodoMVC, Vue | 546.2 | 429.8 | 476.2 |
| TodoMVC, React | 558.0 | 427.4 | 467.7 |
| TodoMVC, Preact | 527.9 | 411.3 | 455.6 |
| news. | 505.3 | 395.8 | 441.8 |
| example. | 416.2 | 326.2 | 388.6 |
Medians in ms, from per_url in each ladder run’s analysis.json.
Correction
An earlier corpus series, from 2026-08-16, is withdrawn. Before fcvm 90733b854e, pasta redirected the guest’s DNS to the host’s resolver, so those runs resolved third-party hosts on the live internet and waited on them. Every corpus run on this page passed the DNS checks described under Method and gates. The withdrawn records and their evidence are in REVIEW.md.
Memory configuration and VM overhead (synthetic page)
Synthetic page
This section uses medium.html, a small deterministic page (1.4 KB of HTML, one stylesheet, 806 B of script, four images) served with Cache-Control: no-store, on the fixture host, with 1024 MiB, 2 vCPU guests except where a chart says otherwise. It exists to compare configurations. Absolute numbers move between goldens that also ran different harness revisions: the same copy, 4 KB configuration read 413.2 ms on the first golden and 532.5 ms on the second. So configuration comparisons use pairs measured on the same golden; the VM against container chart sets two separately measured setups side by side. None of these numbers is comparable to the corpus results above. The first two fixture goldens (RB, AB) were warmed on a page that shared medium.html’s stylesheet and one image; PF and the later goldens share no byte resources with it. Browser and system-font caches stay warm in every golden.
The memory server records the pages each clone faults and populates that set in the next clone at restore, in bounded chunks interleaved with the guest’s own faults. On the same golden, the run with prefetch on has a median 284.7 ms (37%) lower than the run with it off. Both runs used one sealed harness that passes the prefetch setting to the memory server explicitly, so the binary’s default could not apply; neither record stores the value passed. Commit e78350c5, which added the explicit setting, names the first run on and the second off. The port-accept stage is 59.3 ms in both, which says nothing about where the population cost lands, because that stage can end before the snapshot is loaded.
localhost/chromium-bench with 2 GiB guests (guest_bytes in its summary.json); its vCPU count and host are not recorded, and bench.sh defaulted to 2 vCPUs.In copy mode every faulting page is copied into the clone. In minor mode the page already sits in the shared memfd and is only mapped. The median time of the ioctl that resolves a fault was 2.3 µs in copy mode and 1.5 µs in minor mode; that is not the whole cost of a fault, and the memory server’s CPU per fault was about 27 µs in both. The five 2 MB minor traces recorded no userfaultfd faults. A file-backed restore uses the kernel page cache and has no userspace fault path.
Restores of one snapshot fault in overlapping page sets: the mean pairwise Jaccard similarity of two restores’ faulted pages is 0.65 in minor mode and 0.69 in copy mode, lowest pair 0.54. Per request, 45 to 59% of faults fall in runs of 4 or more consecutive pages.
| Page size | copy | minor | Golden |
|---|---|---|---|
| 4 KB | 532.5 | 607.9 | shared by the pair (AB) |
| 2 MB | 401.9 | 374.9 | shared by the pair (HK, HM) |
Request medians in ms, direct CDP. On 4 KB pages minor mode is 14% slower than copy. On 2 MB pages minor mode is faster in all three render arms: direct CDP 374.9 against 401.9 ms, fast teardown 380.8 against 401.0, in-guest exec 582.9 against 611.9. In the direct CDP arm the largest stage-median difference is the wait for the target (130.7 against 156.7 ms); navigate is 109.6 against 109.7 ms, and screenshot is slower in minor mode (81.9 against 78.9 ms). HM and HK share one golden. Its 2 MB pages are recorded as hugepages: true in each run’s clone-state capture, a raw file that is not committed; the committed analysis.json holds only the snapshot name and config digest. Their harness (b72e516a, which matches both runs’ harness_sha256) passes the prefetch setting to the memory server explicitly, on unless the environment overrides it. Neither record stores the value passed, and the run log that echoes it is not committed.
- microVM per request
- warm host container
The request medians are 477.9 ms in the VM, from launching the VM to the screenshot returned, and 133.0 ms in the container, the driver’s own time from target lookup to its return. The container run’s wall time per request was 193.7 ms. The 60.7 ms between the two is the per-request process wrapper around the driver, including Python start-up, which the VM run, driving it in-process, does not pay; no run isolates its parts. The largest stage-median differences are the wait for the target (+114.0 ms) and navigate (+156.4 ms); screenshot differs by +3.1 ms. The VM had 2 vCPUs and the container had no CPU limit, so these differences also include a difference in available CPU. Both runs render medium.html with cdpdrive.py. The VM run seals the driver in its harness hash; the host record names cdpdrive.py without a hash and used a different build of the container image. These are two separately measured setups, not a split of the cost into virtualization and memory.
All fixture configurations
| Run and configuration | No-render arm p50 | Direct CDP p50 | Fast teardown p50 |
|---|---|---|---|
| RB-uffd: copy, 4 KB, first golden | 59.3 | 413.2 | 413.0 |
| RB-file: file-backed, 4 KB, first golden | 58.5 | 458.9 | 453.8 |
| AB-copy: copy, 4 KB, second golden | 59.3 | 532.5 | 532.7 |
| AB-minor: minor, 4 KB, second golden | 59.3 | 607.9 | 604.8 |
| PF-on: copy, 4 KB, prefetch on, third golden | 59.3 | 477.9 | 479.4 |
| PF-off: copy, 4 KB, prefetch off, third golden | 59.4 | 762.6 | 760.5 |
| HM: minor, 2 MB, hugepage golden | 39.9 | 374.9 | 380.8 |
| HK: copy, 2 MB, hugepage golden | 39.9 | 401.9 | 401.0 |
| NC: HM with Chromium’s profile and disk cache kept off the rootfs, no exec arm, new golden | 39.9 | 372.5 | 371.6 |
| FG: NC’s image on another new golden, harness with a 2 ms port probe (commit 1b115d55) | 40.9 | 348.7 | 347.4 |
| HC: warm host container, no VM; the driver’s own time, target lookup to its return | not applicable | 133.0 | not applicable |
Medians in ms, 202 measured requests per run and arm (200 for HC). No arm failed: 0 in 204 or 232 attempts per arm, an exact 95% interval of 0 to at most 1.79%, and HC 0 in 202, 0 to 1.81%. Direct CDP: the harness drives Chromium over CDP from the host. Fast teardown: the same render and the same timed span; only the teardown after the answer differs, one SIGKILL carried to fcvm’s children by the pdeathsig chain in place of fcvm’s SIGTERM-and-await sequence. Exec arm: the retired in-guest driver. The no-render arm stops when the forwarded port accepts, so its figures sit on the port probe’s grid: before FG the probe backed off to about 17 to 20 ms between tries, and FG ran with a 2 ms cap. Both come from commit 1b115d55, which names FG as the fine-grid run; FG’s recorded source revision still has the 20 ms cap, and its harness hash matches no committed tree. The 4 KB and 2 MB rows come from different goldens, so their no-render figures are not a page-size comparison; an earlier page-size comparison drawn from them was withdrawn. NC and FG differ in the golden and in the harness, so their difference is attributed to neither. RB and AB ran on goldens whose warm-up page shared styles.css and img1.png with medium.html.
Not measured
- CPU and memory per request on the corpus. Cloudflare publishes both; we have no whole-system CPU figure and no memory figure on a reconciled basis.
- HTML extraction, Cloudflare’s second operation. No run performs it.
- Density. How many concurrent renders a host sustains, and at what latency.
- Run-to-run variance. No configuration was measured twice under the same conditions. Five independent runs per configuration would put error bars on the medians.
- Other network modes. Only rootless (pasta) is published. A bridged-network run was refused by the zero-failure gate (the sample gate under Method) at 1 failure in 205 requests. No record of that run is retained; its investigation is in #820, closed 2026-08-15.
- Other engines, including WebKit.
Method and gates
- Sample gate. A run publishes only with at least 200 measured non-warmup requests and zero failures over every attempt, warmups included; otherwise the analyzer exits 5. Zero failures in every published cell: each of the four corpus runs has 0 failures in 230 render attempts (202 measured plus 28 warmups), an exact 95% (Clopper-Pearson) interval of 0 to 1.59% per run, and 0 in the no-render arm’s 230.
- Quiet host. A run refuses to start above a 1-minute load of 2.0. The check runs only at start.
- Runtime seal. The harness hashes itself, fcvm and fc-agent into a bundle and refuses to measure a golden made under a different bundle.
- Drift gate. The no-render arm runs interleaved through every run. The analyzer compares its first and second halves and marks the run unpublishable unless the 95% confidence interval of the shift lies within 10 ms either way. Seven committed fixture runs failed this gate (
publishableis false in theiranalysis.json), none with a significant shift: five moved under 0.2 ms, and every one of their intervals reached about 17 ms on at least one side. All seven ran before the 2 ms port probe. None appears on this page. - DNS verification. Each corpus run resolves every corpus host from a clone three times (before the settle wait, just before measuring and just after), and each check passes only if the replay server’s own log shows an A answer of 10.0.2.2 for all 10 corpus hosts in the rows it added during that check. A sampler records every 10 s of the measured run that the replay server owns port 53 and dnsmasq is inactive. The server logs each query before it sends the reply and does not record a failed send, so its log (5,306 queries in the 2026-08-30 run) shows replies prepared, not datagrams delivered; delivery during the checks is shown by the answers the clone itself resolved. The verdict is written to
dns-evidence.json. - Replay fidelity. 12 of the 14 sites pass the replay equivalence check: DOM similarity at least 0.90 and under 10% of pixels changed, where a pixel counts as changed when the grayscale value of its per-channel difference is 16/255 or more. Most pass with DOM similarity 1.00 and no changed pixel. The two that fail are news sites whose pixel differences are confined to ad slots and consent overlays, whose third-party requests were not captured. The check’s report is not committed; its result is in the message of commit
f7d54690. - Lints.
bench/checks the ladder, stage-median and per-site tables and the headline median against the committed records (each figure must equal a median or interval bound of a DNS-verified run), rejects withdrawn figures, and keeps Cloudflare’s medians out of sentences that quote our within-run intervals.chromium/ test_reqbench.py bench/checks Cloudflare’s rows against the published table and the page’s structure. Fixture figures and chart labels are not checked against records.chromium/ test_report_page.py
To reproduce the corpus campaign (golden, verification, DNS diagnostics, measured run) from a clean checkout:
make bench-chromium-corpus
The guest vCPU count is part of the golden, so each rung of a ladder needs its own golden. The phase-by-phase targets and knobs are in bench/chromium/AGENTS.md.
Firecracker builds
fcvm drives a fork, ejc3/firecracker. rootfs-config.toml at each run’s fcvm source revision selects a branch. Setup builds or reuses the binary for that branch’s head when it can reach the remote, and falls back to a cached binary when it cannot. The corpus and fixture runs used different branches. No record stores the Firecracker commit a run built.
Corpus runs
Branch agent/uffd-minor, whose current head is 32f9ef433, committed 2026-08-14, before all four published corpus runs: upstream main as of 2026-08-07 (03b096f3b) plus two changes. The ladder’s hostinfo.json records the binary’s content hash.
- Minor-fault memory backend. During the userfaultfd handshake the fcvm memory server passes a sealed memfd holding the snapshot’s memory. Firecracker maps guest RAM
MAP_PRIVATEover it and registers it for minor faults. Every minor-mode run here depends on it. - vsock connection limit raised from 1023 to 16384, with the kill queue and receive queue sized to match.
Upstream’s own fixes for the serial interrupt on restore and the dirty-page dump are in this base, so the branch does not carry the fork’s earlier patches for them.
Fixture runs
Branch bump-vsock-max-connections, whose current head is 014c1c15, committed 2026-08-07, before every fixture run: upstream main as of 2026-02-17 plus four changes. Two are earlier versions of the changes above: the vsock connection limit raised from 1023 to 16384 and the kill queue from 128 to 2048, with the receive queue left at 256 (5ae6d15ed), and the first revision of the minor-fault backend, which passes the sealed memfd in the handshake but predates the bounded-handshake commits. The other two are fixes upstream later made separately:
- Serial interrupt on restore, a fix plus its regression test. Reported upstream as #5730; upstream fixed it in #5764 (2026-03-31).
- Dirty-page dump alignment, a snapshot-writing fix. Reported upstream as #5696; upstream shipped its own fix in v1.14.2 (2026-02-27).
Restored guests need their monotonic clock to continue from the snapshot instant, or every armed timer fires at once. On these plain-EL1 guests the host kernel handles it: since the March 2023 arm64 KVM counter-offset series, KVM recomputes the VM-wide counter offset when a restore replays the saved counter registers. The timer storm fcvm fixed for nested virtualization (#630) needs a virtual-EL2 guest and does not apply to either build.
Records
Every fcvm figure on this page comes from a committed record under bench/chromium/results/, except where a statement names another source: the bridged-network refusal, the replay equivalence result, the prefetch and port-probe settings taken from commit messages, the hugepage setting taken from raw captures, and the Firecracker commits, which no record stores. The reqbench runs record content hashes of the harness and of the runtime bundle holding fcvm and fc-agent, the snapshot generation and its config digest. HC records its configuration and host kernel in run.json; FB records its configuration and guest size only, with no host, kernel or prefetch setting. Cloudflare’s benchmark rows are quoted from their post; the web-platform and later-benchmark figures come from Cloudflare’s docs and kitesurf.dev, cited where they appear, and the Chrome launch median is computed from kitesurf.dev’s raw data.
| Id | What | Record |
|---|---|---|
| CL2, CL4, CL8 | Corpus vCPU ladder, 2026-09-02, one golden per rung, DNS-verified | results/reqbench-20260902-023115-corpus-c2, results/reqbench-20260902-025115-corpus-c4, results/reqbench-20260902-031115-corpus-c8; index results/campaign-20260902-box2-ladder-summary.json |
| CV | Corpus, 2 vCPU, 2026-08-30, DNS-verified | results/reqbench-20260830-171007-corpus; index results/campaign-20260830-box2-summary.json |
| RB-uffd, RB-file | Fixture, first golden | results/reqbench-86d0f4aca0ca486b8b1a9239792a2d7d, results/reqbench-ccd7655151d94a6ea47fba8aa23f352d |
| AB-copy, AB-minor | Fixture, second golden | results/reqbench-c61bd9e8e4764a23aca3d0ab7ca07a3a, results/reqbench-8fbec71a1863427fa4daddc9f1f53af1 |
| PF-on, PF-off | Fixture, third golden, one harness passing the prefetch setting explicitly; the value is in neither record, and commit e78350c5 names the first run on and the second off | results/reqbench-9c23e9d7da424351a390b5c1ddfd8e1d, results/reqbench-134b408e1a3d42feb866bb4c5fed774c |
| HM, HK | Fixture, 2 MB pages, minor and copy, shared golden | results/reqbench-20260814-001144-uffd, results/reqbench-20260814-002056-uffd |
| NC, FG | Fixture, Chromium writes kept off the rootfs, on two goldens and two harness revisions | results/reqbench-20260814-024029-uffd, results/reqbench-20260814-035757-uffd |
| HC | Warm host container, no VM | results/hostcdp-bc17d3ffb64c46ab8d2ef51c0043fb9b |
| FB | userfaultfd fault counts on bench.sh snapshots, 2 GiB guests; prefetch setting not recorded | results/faultbench-0813-073507-690998 |