Disk-only clone

Cold-boot fresh, fully independent VMs from a captured disk — with a free choice of memory size, CPU, and hugepages.

fcvm design issue #625 status: implemented (as-built design below) code-grounded · verified file:line
Implementation outcome (2026-06). The feature shipped with a simpler design than this RFC's policy enums. StoragePolicy and ContainerStatePolicy were dropped: the guest self-detects a cold boot of an already-provisioned disk via a provisioned marker (/var/lib/fcvm/provisioned, written by fc-agent after first-boot provisioning). When present, fc-agent skips the storage wipe and image import, re-mounts the existing storage loopback (never mkfs), podman starts the captured fcvm-container, and regenerates identity. No MMDS plumbing — the disk itself carries the signal. The host keys solely off rootfs_override (boot the captured disk; skip image resolution and the snapshot cache). The same machinery powers in-place reboot: every VM builds its launch plan up front and all lifecycle paths boot through one shared configure_and_boot_firecracker primitive, so a guest reboot relaunches the VM from its own disk — semantically identical to a disk-only clone cold boot. The capture path (freeze → reflink → unfreeze, no vCPU pause) shipped as designed. See DESIGN.md → Storage & Cloning for the as-built reference.

TL;DR

fcvm's existing snapshot commands clone a running VM's memory + disk and resume it mid-execution via a UFFD page server. Disk-only clone captures only the disk at a consistent point and cold-boots brand-new VMs from it — no memory image, no UFFD, no shared state. Each clone is a fully independent machine that boots its kernel from scratch and re-launches the container against the prepared filesystem.

# capture only the disk of a running VM (paused briefly for a consistent reflink)
fcvm snapshot create --pid <vm_pid> --tag base --disk-only

# cold-boot independent copies — memory/cpu/hugepages chosen freely at run time
fcvm snapshot run --tag base --name c1 --hugepages --memory 8192 --cpu 4
fcvm snapshot run --tag base --name c2          # another fresh, independent copy

Proposed CLI. This work introduces --disk-only on snapshot create and promotes --tag + --memory/--cpu/--hugepages/--env onto snapshot run (today's run uses --snapshot with cpu/mem internal-only — see P4).

The headline architectural finding: this feature already exists as the empty cell of a 2×2 matrix. The cold-boot path and the snapshot-restore path already share the same disk-reflink machinery (DiskManager::create_cow_disk). Disk-only clone is (captured disk) × (cold boot) — a combination the current code never wires up, but every piece of which is already built and tested. The design's job is to make that matrix explicit instead of adding a fourth ad-hoc path.

Review provenance: this design was hardened by two adversarial passes (local codex + /code-review) before any code was written. They caught four would-be data-loss / broken-build bugs — missing guest quiescing, an unmodeled container-state collision, ungated image resolution, and an over-broad overlay rejection — now folded into the abstraction, risks, and PR plan below. Findings and fixes are flagged inline with (codex-flagged).

The problem & the two asks

From #625, verbatim:

“i want functionality that can clone a vm but not its memory state … get a vm booted and then we clone the disk and then can start fresh copies from that cloned disk state.”

“i want the ability to start that fresh clone with huge tb's even if the original one wasnt started that way.” (huge tb's = hugepages / huge-TLB memory)

Why this is useful

Boot a VM, do expensive one-time work inside it — install packages, warm an image cache, seed a database, run a build — then capture the disk once and fan out N cold copies that each start from that prepared state. Unlike a memory clone, the copies are not frozen mid-execution: they boot fresh, so they can run a different command, get different resources, and have no shared memory pages or in-flight state to corrupt.

The second ask falls out for free

“Start the clone with hugepages even if the original wasn't” is hard for a memory clone: Firecracker rejects the file-backed memory backend for hugepage snapshots, so hugepages force the UFFD restore path and the backing must match the snapshot (src/commands/snapshot.rs:1001, 1016). A disk-only clone has no memory image to restore, so that entire constraint evaporates — the clone allocates fresh guest RAM and can request --hugepages, any --memory, any --cpu with zero coupling to the source. See Hugepages.

How it differs from the existing snapshot/clone

Existing memory+disk clone (snapshot serve/run --pid)Disk-only clone (NEW)
Capturesmemory.bin + vmstate.bin + disk.rawdisk.raw only (+ :ro extra disks)
RestoreResume mid-execution; UFFD serves pages on demandCold boot: fresh kernel, fc-agent re-launches container
Shared stateClones share physical pages (CoW) via a serve processNone — fully independent processes
Serve stepRequired (snapshot serve → UFFD socket)None
Memory/cpu/hugepagesConstrained by the snapshot backingFree at run time
NetworkIP reconstructed from vmstate; restore-mode quirksFresh allocation; guest reconfigures eth0 from boot args
Boot timeMilliseconds (resume)Seconds (full boot) — the tradeoff for independence
Use caseInstant warm forks of a live processFan-out of a prepared filesystem template

The proper abstraction

Today there are two boot paths that look unrelated but in fact share their disk plumbing and differ along exactly two orthogonal axes. The design makes those axes first-class instead of bolting on a third special case.

Axis A — RootfsSource: where does rootfs.raw come from?

DiskManager::new(vm_id, base_rootfs, vm_dir) reflinks whatever base_rootfs points at into the per-VM rootfs.raw (src/storage/disk.rs:26,41). It is already source-agnostic:

Axis B — BootMode: how does the guest come up?

The matrix — disk-only clone is the empty cell

BootMode = ColdBootBootMode = MemoryRestore
RootfsSource
= BaseLayer
✓ podman run
Today's fresh VM boot
— n/a
no memory image for a fresh base
RootfsSource
= CapturedDisk
★ disk-only clone
NEW — the combination this design fills in
✓ snapshot run --pid
Today's UFFD memory clone

Both occupied cells already share create_cow_disk. Disk-only clone = the existing CapturedDisk reflink (from the UFFD cell) composed with the existing ColdBoot machinery (from the podman run cell). Almost no new mechanism — just a new wiring of two proven ones.

Axis C — three policies the matrix exposes

Placing the empty cell next to its neighbours reveals that a disk-only clone is not just “cold boot with a different disk” — three smaller policies become explicit. The first review of this design collapsed them into one coarse StoragePolicy; an adversarial pass showed that hides two data-loss bugs, so they are split out.

C1 · QuiescePolicy (capture side)

What state must be flushed to the disk before the reflink. A memory snapshot is consistent only because it also captures the dirty page cache; a disk-only capture has no such safety net, so it must sync/fsfreeze (or stop the workload) first. Detailed in the capture path.

C2 · StoragePolicy (guest side): provision vs. preserve

On a normal cold boot fc-agent provisions storage — it wipes podman state and (re)installs the image. On a disk-only cold boot the storage is already prepared inside the captured disk, so fc-agent must preserve it:

StoragePolicyUsed byfc-agent behavior
Provision (default)BaseLayer cold bootwipe storage, set up storage.conf, load/pull the image
PreserveCapturedDisk cold bootkeep captured storage; skip every reset + the destructive btrfs setup; verify image present; re-attach a read-only store only if one existed

C3 · ContainerStatePolicy (guest side)

The subtlety an adversarial review surfaced: fc-agent always issues podman run --name fcvm-container with no --replace (fc-agent/src/container.rs:948). A captured VM already has an fcvm-container record and writable layer, so a naïve re-run either collides on the name or, if pre-removed, deletes the prepared layer. The clone must choose explicitly:

ContainerStatePolicySemanticsPreserves writable layer?
StartCaptured (default, v1)podman start the captured container — re-runs its entrypoint against its exact writable layer; no name collision (it's a start, not a run)yes
FreshContainerremove the captured record, podman run a new one — keeps the image + volumes, discards the old container's writable layerno (image + volumes survive)

v1 defaults to StartCaptured: it is the most faithful reading of “start fresh copies from that cloned disk state” — everything written inside the source (including the container's own filesystem) reappears in each copy — and it sidesteps the name collision because podman start fcvm-container resumes an existing record rather than creating a duplicate. This requires fc-agent to branch from the current unconditional podman run --name fcvm-container (container.rs:948) to podman start for the captured container. FreshContainer is the opt-in for “same image/volumes, clean container.” The P4 “file written inside the container layer survives” test exercises the StartCaptured default.

These policies surface on the wire as a few #[serde(default)] MMDS fields whose defaults equal today's behavior, so every existing plan deserializes unchanged. This is where the feature can destroy data if done wrong — see fc-agent and Risks.

Why this is the “proper” abstraction, not a fourth path: it names the two axes the code already varies along (RootfsSource × BootMode), places all three real boot scenarios in one matrix, and factors the genuinely new behavior into three small, defaulted policies instead of one ad-hoc branch. No existing path is special-cased; the disk-only clone is a composition plus three explicit choices. It also lines up with the planned hypervisor abstraction — RootfsSource/BootMode map onto its neutral RestoreSource/SnapshotKind types, and cold-boot-from-disk is VMM-agnostic (both Firecracker and Cloud Hypervisor do direct kernel boot).

Reuse seams (verified file:line)

Threading rootfs_override: Option<PathBuf> + StoragePolicy through the existing path touches a remarkably small surface, because everything downstream is already parameterized on base_rootfs and args.

#SeamLocationChange
1RunArgs entry fieldssrc/cli/args.rs (end of struct, ~:300)add rootfs_override: Option<PathBuf>, storage_policy
2base_rootfs resolutionsrc/commands/podman/mod.rs:297match rootfs_override { Some(p)=>p, None=>ensure_rootfs(..) }
3snapshot-cache diversionsrc/commands/podman/mod.rs:326–333,351,381gate on rootfs_override.is_none() so a clone does not divert into the UFFD path
3bimage resolution + localhost exportsrc/commands/podman/mod.rs:320 (cache ref), :451 (export)also gate on rootfs_override — a baked-in clone must NOT require the original host image tag to still exist or re-export it (codex-flagged gap)
4disk reflinksrc/commands/podman/vm_config.rs:805–811unchanged — create_cow_disk reflinks whatever base it's given
5MachineConfig (cpu/mem/hugepages)src/commands/podman/vm_config.rs:927–935unchanged — already huge_pages: if args.hugepages {Some("2M")}
6MMDS boot-planbuild_and_send_mmds, vm_config.rs:1098; to_mmds_json, config.rs:357add one storage_policy field to the plan
7network setupprepare_vm, mod.rs:668–725unchanged — fresh TAP + allocate_loopback_ip + network.setup()

The dispatch wrapper (cmd_snapshot_run on a DiskOnly tag) synthesizes a RunArgs from the snapshot's metadata plus CLI overrides, sets rootfs_override = Some(disk.raw) + storage_policy = Preserve + no_snapshot = true, then calls the ordinary prepare_vm → run_vm_loop → cleanup_vm_context trio (mod.rs:155 / :1060 / :1205). All network / vsock / volume / listener / health wiring is reused verbatim.

fc-agent: the Preserve storage policy

On a normal boot the guest agent destroys podman storage and rebuilds it. A disk-only clone must not — the captured disk already contains the writable container layer and (for some image modes) the image itself. This is the highest-risk part of the feature.

The wipe risk (citations corrected after review). There are four reachable destroyers: podman system reset --force at container.rs:451 (root), :424 (btrfs setup), :64 (overlay mount), and a reset inside create_vm_user at container.rs:1103 (called from agent.rs:225) when a non-root --user is set. And the btrfs setup is destructive even before its reset: it truncates /var/lib/containers/btrfs.img (container.rs:357) and runs mkfs.btrfs -f (container.rs:377), reached via agent.rs:189 for non-overlay modes. A Preserve boot must skip all of these, not just the resets.

Branch points a Preserve plan must guard

StepLocationProvisionPreserve
write_early_storage_conf()agent.rs:112doskip — keep captured storage.conf
btrfs storage setup branchagent.rs:189–198donon-destructive variant — mount existing btrfs.img, skip reset/truncate/mkfs (see below)
reset_podman_state()agent.rs:246–248doskip — the wipe
image-prep match (mount / load / pull)agent.rs:251–274doskip — new Preserve arm
create_vm_user reset (non-root)container.rs:1103 (via agent.rs:225)doskip
btrfs setup: truncate + mkfs.btrfs -fcontainer.rs:357,377 (via agent.rs:189)doskip the destroy — but still mount the existing btrfs.img (it's a loopback, not persistent across boot — see below)
container start (--name fcvm-container, no --replace)container.rs:948run newper ContainerStatePolicy: podman start the captured container (v1) or remove-then-run
overlay unmount at cleanupagent.rs:436–438cond.n/a — no /mnt/image-store unless re-attached
verify image presentget_image_digest, container.rs:697—do — bail clearly if absent (no load/pull fallback)
exec server, FUSE/disk/NFS mounts, chronyd, exit-notifyagent.rs:115–443dodo — identical

On the wire this is #[serde(default)] storage_policy/container_state_policy fields on the Plan struct (fc-agent/src/types.rs), defaulting to today's behavior — so existing plans deserialize unchanged (proven by the minimal-plan test). Changing fc-agent requires rebuilding the initrd (content-addressed by the fc-agent binary SHA).

Subtlety codex's second pass caught — Preserve ≠ "skip the btrfs branch entirely". On a non-native-btrfs rootfs the image store is a loopback at /var/lib/containers/btrfs.img that is created+formatted+mounted at boot (container.rs:328/357/377/396). The file is reflinked (it's inside rootfs.raw), but the loopback mount does not survive a cold boot. So Preserve must skip only the destructive steps (truncate + mkfs) while still mounting the existing btrfs.img — otherwise a btrfs/archive clone boots with the image present on disk but not mounted, i.e. invisible to podman.

The image-delivery wrinkle (refined after review)

The first draft said “reject overlay mode.” An adversarial pass showed that conflates two different things — what matters is where the image physically lives, which is determined by delivery, not the image_mode label:

How the source got its imageWhere the image livesDisk-only capture
Remote pull (image_mode=None, no device) — fc-agent pulls into rootfs at boot (agent.rs:264)baked into rootfs.raw✓ works — reflink carries it
btrfs / archive — fc-agent podman loads into rootfs storagebaked into rootfs.raw✓ works
localhost overlay — RO .storage-v2.img attached as additionalImageStore (mod.rs:541)separate content-addressed device, NOT in rootfsv1 reject / v2 re-attach

So the v1 guard must key on the effective delivery — “was a separate read-only overlay store device attached?” — which is exactly what MMDS records: it emits image_mode only alongside an image_device (config.rs:377). It must not key on the raw resolve_image_mode(), which defaults to Overlay even for remote pulls (mod.rs:77) that have no device and bake into the rootfs; and not on “an image_device exists,” because btrfs/archive also attach a device (vm_config.rs:481) yet podman load it into the rootfs. The reject condition is precisely overlay-mode AND a store device was attached. In every mode the container's writable layer lives in /var/lib/containers on rootfs.raw; that is what Preserve + ContainerStatePolicy protect.

v1 — rootfs-baked images

Remote-pull, btrfs, and archive: the image is in the reflinked rootfs; Preserve skips the wipe and runs from local storage. Only the separate-store localhost-overlay capture is rejected (clear error). Smallest, safest first cut.

v2 — overlay re-attach

Record image_mode + digest; on cold boot re-attach the same content-addressed .storage device (shared, always present) exactly like a normal overlay boot, then Preserve keeps the writable layer. No baking needed.

The capture path

Disk-only capture is create_snapshot_core (src/commands/common.rs:1553) with the memory half removed and a guest quiesce added. The factored create_disk_only_snapshot_core keeps the pause/resume bracket, adds the quiesce, and drops the Firecracker memory dump.

Step in create_snapshot_corelinesDisk-only
acquire snapshot semaphore; derive dirs1564–1574keep
determine diff base / parent memory1585–1604skip (no memory ⇒ no diff chain)
memory disk-space check (memory_mib×1MiB)1619–1646skip (reflink is O(1))
pause VM (patch_vm_state "Paused")1683–1692keep — consistency
Firecracker create_snapshot (memory dump)1696–1702skip — the call to drop
reflink disk (+ extra disks)1774–1803keep — the core; do it unconditionally after pause
resume VM (patch_vm_state "Resumed")1807–1817keep — “resume no matter what”
diff merge1843–1882skip
write config.json; atomic rename into place1903–1927keep — finalize, kind = DiskOnly, parent = None

Consistency — pause is NOT enough (corrected after review). The memory-snapshot flow is consistent because it captures the dirty guest page cache in memory.bin at common.rs:1696. Disk-only drops that, so vCPU pause alone (common.rs:1683) leaves un-written-back dirty pages and in-flight journal state that never reach rootfs.raw — the reflink would capture a torn filesystem. A disk-only capture must quiesce the guest first.

QuiescePolicy — flush the guest before the reflink

Before reflinking, fcvm runs a quiesce command inside the guest over the existing exec vsock channel — the same path fcvm exec uses (host src/commands/exec.rs:146 → fc-agent exec server agent.rs:120 → guest command fc-agent/src/exec.rs:377). No new control channel is needed, and because it's a runtime command (not the boot-time Plan), capture can use it without any fc-agent build change.

LevelAction in guestGuarantee
Syncsync(2)best-effort only — a write window remains between sync and the vCPU pause, so not race-free on its own
Freeze (the correctness floor)fsfreeze --freeze each writable mount → reflink → fsfreeze --unfreezefilesystem is crash-consistent and no writes can land across the reflink
StopContainerstop the workload, then freezestrongest: also drains userspace buffers (needed if the app buffers writes itself)

Ordering: freeze (exec vsock) → vCPU pause → reflink → vCPU resume → unfreeze (exec vsock). Freeze is the floor because Sync alone leaves a window before the pause. This is the gap codex's first review flagged ("needs disk quiescing").

Unfreeze no matter what. A frozen filesystem left frozen wedges the source VM. Today's snapshot code returns immediately on a pause failure (common.rs:1683) — acceptable when nothing is frozen, but disk-only capture must fsfreeze --unfreeze on every exit path (reflink error, pause failure, resume failure), with the same “resume no matter what” discipline extended to the freeze.

Metadata the DiskOnly config must carry

SnapshotMetadata (src/storage/snapshot.rs:66–112) already records image, vcpu, memory, network, volumes, hugepages, extra_disks, user, ports, forward-localhost, network-mode, ipv6-prefix, tty, interactive. To reconstruct a cold-boot plan it additionally needs fields that are currently not persisted anywhere:

FieldWhy neededStatus today
kind: SnapshotKinddistinguish Full vs DiskOnly at run dispatchdoes not exist (#[serde(default)] ⇒ Full)
effective image deliverydid a separate read-only store get attached? (overlay) vs rootfs-baked (btrfs/archive/remote). v1 rejects only the separate-store casenot stored; raw resolve_image_mode (mod.rs:58) even defaults to Overlay for remote — so persist the effective delivery, i.e. whether MMDS attached a store device (config.rs:377)
container_cmdre-run the original command on cold bootno field in VmConfig at all
kernel_profilea nested/btrfs-profile source must cold-boot under the SAME kernel, not the defaultchosen each run (mod.rs:220,278), never persisted — rootfs_type alone is insufficient
privileged, rootfs_type, non_blocking_outputreproduce the boot plannot persisted

Env vars are intentionally NOT persisted (they may hold secrets; state files are world-readable — mod.rs:613–614, types.rs:109). Disk-only run therefore re-supplies env at run time via a new --env on snapshot run, never baking secrets into the snapshot directory.

Hugepages & resources come free

The second ask — “start the clone with hugepages even if the original wasn't” — needs no new mechanism. It is purely a consequence of cold-booting.

Memory clone (constrained)

Hugepages force the UFFD restore path; Firecracker rejects the file backend for hugepage snapshots, and the backing must match what was captured.

// src/commands/snapshot.rs:1016
if hugepages || env("FCVM_FORCE_UFFD") {
    // implicit in-process UFFD server …
}

Disk-only clone (free)

No memory image ⇒ no backing to match. The cold-boot path already feeds args straight into the machine config:

// src/commands/podman/vm_config.rs:927
MachineConfig {
  vcpu_count: args.cpu,
  mem_size_mib: args.mem,
  huge_pages: if args.hugepages { Some("2M") } else { None },
}

--hugepages, --memory, and --cpu are already implemented for the fresh-boot path; the disk-only run just exposes them on snapshot run and lets them flow through unchanged. The only existing guard that still applies is the 2 MiB memory-alignment check for hugepages (mod.rs:162).

Top risks & mitigations

Ranked; the first three are data-loss / silent-corruption class and gate the design.

  1. Torn filesystem — pause alone doesn't flush the page cache. (codex-flagged) Without a memory snapshot, dirty guest pages never reach the disk. Mitigation: QuiescePolicy — sync/fsfreeze (or stop the workload) before the reflink; freeze brackets the copy.
  2. fc-agent destroys the captured storage. Four resets plus a truncate+mkfs btrfs setup. Mitigation: StoragePolicy::Preserve guards all of them (container.rs:64/424/451/1103 + 357/377); Provision stays the default. Covered by the “file inside the container layer survives the clone” e2e test.
  3. Container name collision / writable-layer loss. (codex-flagged) podman run --name fcvm-container has no --replace. Mitigation: ContainerStatePolicy — v1 StartCaptured (podman start, preserves the writable layer, no collision); FreshContainer opt-in.
  4. Clone needs the original host image tag. (codex-flagged) prepare_vm always resolves/exports the image. Mitigation: gate image-resolution + localhost-export (mod.rs:320,451) on rootfs_override — a rootfs-baked clone needs neither.
  5. Separate-store (localhost-overlay) image not in the rootfs. Mitigation: v1 rejects only when a separate read-only overlay store device was actually attached (the effective delivery, recorded in metadata) — never the raw resolve_image_mode(), which defaults to Overlay even for remote pulls that bake into the rootfs; v2 re-attaches the content-addressed store.
  6. Wrong kernel for a nested/btrfs-profile source. Mitigation: persist kernel_profile in metadata and cold-boot under it; rootfs_type alone is insufficient.
  7. Secrets leaking into a world-readable snapshot. Mitigation: env vars are never persisted; re-supplied at run time via --env.
  8. Footgun: --pid (UFFD) on a DiskOnly tag. Mitigation: dispatch rejects it; a DiskOnly tag has no memory image to serve.

Stacked PR sequence — green at every step

The work is decomposed so each PR compiles, passes tests locally, and is independently mergeable. Early PRs are purely additive (new fields default to existing behavior), so nothing changes until the final wiring PR turns the feature on. Each branch is stacked on the previous one (main → P1 → P2 → …).

main P1 metadata P2 capture P3 fc-agent P4 run + e2e P5 docs/pages
no behavior changehost codefc-agent + initrd rebuild
  1. P1 — metadata foundation (no behavior change). Add SnapshotKind { Full, DiskOnly } with #[serde(default)] on SnapshotConfig; persist all fields the later PRs read — effective image delivery (overlay-store-attached vs rootfs-baked), kernel_profile, container_cmd, privileged, rootfs_type, non_blocking_output — into VmConfig (src/state/types.rs:102) and mirror into SnapshotMetadata (built at common.rs:1395). Update all SnapshotConfig{…} literals (prod builder + ~6 test fixtures) to set kind: Full. unit tests for serde round-trip + back-compat (old config deserializes as Full); existing snapshot suite unchanged.
  2. P2 — disk-only capture (+ guard the old run/serve paths). --disk-only on snapshot create; factor create_disk_only_snapshot_core (freeze via exec vsock → pause → reflink → resume → unfreeze, with unfreeze on every exit path; skip memory dump + memory disk-check); reject only separate-store capture (overlay mode with an attached store device — the effective delivery, not raw resolve_image_mode). Quiesce uses the existing exec vsock (exec.rs:146), so it needs no fc-agent change and can land here. Reject a DiskOnly tag in the existing cmd_snapshot_run (snapshot.rs:707,1076) AND snapshot serve (snapshot.rs:291) with a clear error, so P2 leaves no crashing path on main. capture from a live VM → disk.raw present, no memory.bin/vmstate.bin, kind=DiskOnly; freeze invoked + unfreeze-on-failure; Overlay rejection; `snapshot run/serve` on a DiskOnly tag fail with a clear message, not a missing-file panic.
  3. P3 — fc-agent StoragePolicy + ContainerStatePolicy. Add #[serde(default)] storage_policy + container_state_policy to the boot Plan (defaults = today's behavior). Guard all destroyers behind Preserve — the three resets (container.rs:64/424/451), the create_vm_user reset (container.rs:1103), the destructive btrfs truncate/mkfs (container.rs:357/377) while still mounting the existing btrfs.img, early-storage-conf, image-prep; add the Preserve arm (verify image present, run from local storage). Implement StartCaptured (podman start the captured container — v1 default, no name collision, preserves the writable layer) and FreshContainer (remove-then-run). (Quiesce is NOT here — it's a runtime exec command from P2, not a boot-Plan field.) Rebuild initrd. additive (defaults = today); fc-agent unit tests for Plan deserialization + each guard; full existing host suite still green.
  4. P4 — cold-boot run wiring + the end-to-end test. RunArgs.rootfs_override + policies through prepare_vm; gate the snapshot-cache diversion and the image-resolution/localhost-export (mod.rs:320,451) on rootfs_override; promote --memory/--cpu/--hugepages/--env to real args on snapshot run; dispatch cmd_snapshot_run on kind == DiskOnly → synthesize RunArgs from metadata (incl. kernel_profile), reject --pid; set process_type=Clone + snapshot_name. e2e: a file written inside the container layer of the source survives the clone & the container runs; hugepage upgrade (source without → clone with, healthy); two clones independent; clone works after the original host image tag is removed; --pid-on-DiskOnly rejected.
  5. P5 — docs & published design. README + .claude/CLAUDE.md + DESIGN.md coherence; this page on GitHub Pages. (The design page itself lands first, ahead of code, as this PR.) Pages workflow deploys; doc links resolve; lint/format clean.

Why this ordering keeps CI green: P1–P3 are strictly additive with defaults equal to current behavior, so each merges without changing any existing path. The feature only activates in P4, which lands together with its end-to-end test — there is never a half-wired intermediate state on main. fc-agent (P3) precedes the run wiring (P4) so the initrd that P4's test needs already exists.

Testing plan

Following the repo's rule that a regression test must fail without the fix and pass with it, and that new tests are run locally, not merely compiled:

TestAssertsLands in
serde back-compata pre-kind config deserializes as Full; round-tripsP1
disk-only capture artifactdisk.raw present; no memory.bin/vmstate.bin; kind=DiskOnlyP2
Separate-store capture rejectedclear error only when an overlay store device was attached; remote-pull/btrfs/archive capture succeedP2
DiskOnly rejected by run/servesnapshot run and serve on a DiskOnly tag error clearly (no missing-memory panic)P2
Plan deserializationmissing storage_policy ⇒ Provision; Preserve parsesP3
marker survives the clonewrite a file inside source → capture → cold-boot clone → file present & container runsP4
hugepage upgradesource booted without hugepages → clone with --hugepages boots healthyP4
clone independencetwo clones; a write in one is not visible in the otherP4
--pid on DiskOnly rejectedUFFD restore refused for a disk-only tagP4

P4's tests require a rebuilt initrd (from P3). All run via make test-root FILTER=… locally before push; CI runs the full host/container × arm64/x64 × snapshot matrix.

Review gates (issue success criteria)

Fit with the hypervisor abstraction

fcvm has a separate RFC to become hypervisor-agnostic (Firecracker + Cloud Hypervisor — docs/hypervisor-abstraction.md, epic #632). Disk-only clone is designed to land now on the Firecracker path while mapping cleanly onto those future neutral types:

Disk-only conceptHypervisor-abstraction neutral type
RootfsSource::CapturedDiskRestoreSource (VMM-neutral; reflink/bind-mount redirect is already VMM-agnostic)
SnapshotKind::DiskOnlySnapshotKind (the doc already plans Full vs diff; DiskOnly is a third, memory-less kind)
BootMode::ColdBootthe trait's configure+start (direct kernel boot — supported by both FC and CH)
StoragePolicy (guest)part of the VMM-neutral boot-plan over vsock (guest channel is byte-identical across backends)

Because a disk-only clone needs no snapshot format, UFFD, or drive-retarget — the three things that differ most between Firecracker and Cloud Hypervisor — it is the most portable clone variant and a natural early win for the CH backend (which the RFC schedules as “P1: cold boot + run a container, no snapshots”).