Optimizing Vercel Sandbox snapshots for sub-second restores
Vercel Sandbox snapshot restore performance has been dramatically improved from 40 seconds to under 1 second through parallelized downloads and decompression, plus a local disk cache with 95% hit rates. These optimizations power the new Automatic Persistence feature for instant stop-and-resume cycles.
When we recently shipped in Vercel Sandbox to let teams capture and restore a sandbox's entire filesystem state, our initial engineering focus was entirely on reliability, making sure the system would never fail to snapshot or lose data.filesystem snapshots
Once that foundation was stable, our attention turned to performance. p75 snapshot restores were taking over 40 seconds, and through parallelization and local caching, we brought that under one second.
Vercel Sandbox runs on the same infrastructure as our internal builds product, . Each sandbox is an isolated container inside a Firecracker microVM.Hive
A snapshot is a compressed copy of the sandbox's disk. We're working with two different files:
When you call , we compress the into a and upload it to S3. When you call with a snapshot, we download the and decompress it back. Without compression, every snapshot operation transfers hundreds of MBs to low GBs over the network, adding seconds to tens of seconds to every restore.sandbox.snapshot().img.vhsSandbox.create().vhs
With reliability in place, we turned to the restore path, which was painfully sequential. We'd download the entire file from S3 in a single request, wait for it to finish, then decompress it in a single thread..vhs
Snapshots range from 200MB to a few GBs, so that single S3 download alone could take several seconds to tens of seconds. We used the HTTP header to instead, with the AWS Go SDK's API handling the orchestration. After benchmarking different concurrency levels and chunk sizes, we ended up with 2-5x faster downloads.Rangetransfermanagerdownload chunks in parallel
We applied the same thinking to decompression. Our format stores a header and a frame for each allocated region of the disk image, so instead of decoding and decompressing frames one by one, we switched to one decoder feeding N decompression goroutines. That made the to restore 2-4x faster, depending on snapshot size..vhs.vhs.img
Even with both downloading and decompressing parallelized, the pipeline still wrote downloaded data to disk before decompression could begin. Piping S3 range request streams directly into decompression eliminated that intermediary step, cutting end-to-end restore time by another 2x.
As you might have noticed, we so far only talked about improving the slow path, when we need to retrieve a snapshot from S3 on a cache miss. Well, we actually didn't have a fast path, so it was all cache misses. Yeah, we really didn't focus on performance at first.
Our sandboxes run on metal instances with NVMe disks, giving us several terabytes of fast local storage that was mostly sitting unused. We put it to work with a local disk cache using LRU (least recently used) eviction, sized by total disk space rather than number of entries. We cache the decompressed directly rather than the compressed , so a cache hit skips both the download and the decompression. Once the cache fills up, the least recently used snapshots get evicted to make room..img.vhs
Most customers reuse a "base" snapshot across many sandboxes, which gives us a 95% cache hit rate. On those hits, boot time is bounded only by starting the microVM and container.
p75 dropped from 40s to sub-second, and p95 went from 50s to 5s. With our cache hit rate, most sandbox boots skip the download and decompression pipeline entirely.
There's more we can do. Cache affinity, for example, would route sandboxes to metal instances that already have the requested snapshot cached, potentially eliminating the cold path for popular snapshots. But that risks creating thundering herds and hotspotting certain machines, so we're being deliberate about it.
Long term, we want the cold path fast enough that caching is a bonus, not a requirement.
These optimizations already power , now in beta, which automatically snapshots a named sandbox's filesystem when you stop it and restores everything on resume. With sub-second restores, that stop-and-resume cycle feels instant.Automatic Persistence
Filesystem snapshots are available today for all Vercel Sandboxes. Check the to get started.Sandbox documentation
What a snapshot looks like on disk
Parallelize you shall
We… didn't cache?
From 40 seconds to sub-second
The raw disk image (), which can be several GBs
.imgA compressed version in our custom format (Vercel Hive Snapshot), which is what gets uploaded to and downloaded from S3
VHS
Source: original entry ↗
More from Vercel
Follow Vercel to get its new changes in your feed and email digest.
OpenAI Decisions API now available on AI Gateway
OpenAI's Decisions API is now accessible through Vercel's AI Gateway with an OpenAI-compatible endpoint, enabling decision models to answer typed questions and return probabilities, choices, and scores for routing, triage, and guardrails use cases. Support is available across the OpenAI SDK, AI SDK, HTTP API, and CLI with the latest versions.
Timestamp attributes now supported in Vercel Flags
Vercel Flags now supports timestamp attributes for entities, allowing you to create time-based targeting rules. Use this feature to run limited-time campaigns, show content between specific dates, or target users based on registration date.
Glyph Cluster now available in stealth on AI Gateway
Glyph Cluster, a reasoning model for coding and long-context analysis, is now available as a stealth model on Vercel's AI Gateway for Pro and Enterprise plan teams with purchased AI Gateway credits at no cost during the stealth period. The model supports function calling, streams responses, and can be accessed via AI SDK, OpenAI-compatible APIs, and coding agents.