How we changed v8 isolates with Firecracker MicroVMs

About a billion Edge Functions run on Netlify each day — Sunweb personalizing pages, Loto-Québec routing visitors on a cookie examine, and a whole bunch of 1000’s of different websites doing all the things from personalization to routing to auth. All of it runs on a full JavaScript runtime that scales with our clients’ visitors.
This poses a major technical problem, as we attempt to make the latency as little as doable. To run tens, typically a whole bunch of 1000’s, of edge capabilities per second, we have to course of every request, route it accurately, allocate compute capability, and boot each our platform code and the shopper’s code. All of that has to occur inside milliseconds.
Over the previous a number of months, our staff has rebuilt the transport systems behind Edge Functions, working carefully with the staff at Unikraft, who wrote about the experience from their side. In the previous, requests went out to a hosted execution service. Today, they run on MicroVMs inside our personal edge community — roughly 5x sooner on the median. That shift additionally improves safety and reliability, and opens up extra potentialities for operating complicated compute on the edge.
This doesn’t change how Edge Functions are written or used — URL imports, npm packages, Node built-ins, netlify.toml declarations, native growth — all of it really works precisely because it did earlier than. It’s now sooner and extra resilient. In this text we’d wish to share extra in regards to the new structure and our learnings on constructing a brand new compute platform that’s capable of serve excessive quantity at a low efficiency overhead.
The numbers, first
An edge operate runs in entrance of a website, on each request that matches it. The time it takes is time a buyer spends ready, so milliseconds right here depend for greater than they do virtually anyplace else.
A heat invocation — routing to a compute node, coming into a MicroVM, operating the operate, producing response headers — now prices:
- ~5–6ms at median (p50), down from 25–40ms on our earlier transport systems
- 47.4% sooner p99 invocations
- 99.998% availability
- 5x sooner edge operate log supply
A chilly invocation is value stating too. When a request arrives in a area that no compute node has seen earlier than, it must fetch the related pictures earlier than it’s capable of run something. This occurs on about 1.2% of invocations and takes about 9ms on common.
What occurs throughout request time
What follows is the trail a single request takes, so as: it arrives on the edge node, it’s was a specification, is routed to a compute node, after which handed off to a MicroVM that will or could not exist already primarily based on whether or not it’s a chilly or heat invocation.
Request arrives on the edge node
Every request lands on the Netlify edge node closest to the consumer. The node terminates the TLS connection and checks the request path in opposition to the Edge Functions’ routes for that deploy.
If nothing matches, the request carries on to the cache and onto the origin as ordinary. If a route does match, that is the purpose the place the request used to depart our community. With our previous transport systems it went out over the web, ran the sting operate, and got here again to us to move on. With the brand new compute platform, the request is forwarded to a compute node inside our community.

Creating an Edge Function service
When the compute node receives the request with the machine specification and the service ID, it first checks to see if a service with that ID already exists. If it does, it forwards the request into the service for it to ship into the MicroVM. A service lets us have a number of MicroVMs related to the identical website’s Edge Functions, and permits us to configure parameters for when to scale MicroVMs out and in. For instance, we configure every service to solely permit a set variety of requests to be dealt with by a MicroVM earlier than we shut it down, to keep away from MicroVMs operating indefinitely. We use the identical parameters to know when to eagerly boot up one other MicroVM in anticipation of 1 shutting down.
If a service for the location’s edge capabilities doesn’t exist already on the compute node, one is created, and we examine to see if we’ve got all the pictures within the machine specification on disk. If any are lacking, they’re fetched from the sting node and written to disk. This method means we solely fetch the sting operate pictures which are receiving visitors in that area.
The edge node writes a spec
Before the request goes anyplace, the sting node writes a specification for the machine that may run a operate. The spec names three pictures: the runtime, our platform picture, and the sting operate picture. It additionally units the CPU, reminiscence, and connection limits.
The spec travels with the request, on each request. Its hash and site-specific info is computed to change into a service ID. This permits for isolation, since two deploys with totally different code or totally different surroundings variables are totally different companies, and so they by no means share a MicroVM.
This isolation issues most for failures we don’t wish to be doable. A doubtlessly compromised deploy runs in a separate MicroVM, and even when it escapes the runtime, it can not poison different clients or the compute layer itself. V8 isolates, regardless of their title, don’t present this degree of isolation.
Choosing a compute node to run an edge operate
Each area has a bunch of compute nodes. The edge node picks one for the service utilizing rendezvous hashing: the identical service lands on the identical node each time, which is what retains a MicroVM heat and the code already on disk and in cache as soon as it’s been learn. This stickiness offers us a caching technique. If we unfold requests evenly throughout the swarm, we’d find yourself with a better degree of chilly begins.
It’s necessary to keep in mind that whereas sending each request for a operate to the identical compute node is the quick path, it’s additionally how a sizzling spot varieties — the place one busy operate competes for assets with all the things else on that field. A service taking a big share of a area’s visitors pinned to a single node will saturate the node on the expense of different companies.
We steadiness this by enjoyable the stickiness. Over a sure threshold, we unfold the service throughout a slice of nodes. This lets us take in sudden spikes in visitors from a single buyer with out affecting different companies that hashed to the identical node.
Finally, as soon as a node is chosen, it pulls within the operate’s code. A compute node that has served the operate earlier than already has it. A node seeing it for the primary time fetches it as soon as and caches it, so solely the primary request pays that price.

Starting the MicroVM
Each operate runs in its personal Firecracker MicroVM. These are created in underneath a millisecond and begin in about 2ms at p99, as a result of the VM begins a stripped-down Linux surroundings moderately than a full working system. The edge operate’s recordsdata are mounted as an uncompressed EROFS picture after which memory-mapped, so the VM reads solely the elements of the bundle it truly makes use of as a substitute of loading all of it.
When the MicroVM boots up and the JavaScript server begins to pay attention on a port, we take a snapshot of the MicroVM. When the sting operate isn’t being invoked, the MicroVMs operating it scale to zero as a substitute of sitting idle. The subsequent time it’s invoked, we begin a brand new MicroVM from that snapshot. The snapshot is memory-mapped, so the VM can begin executing with out ready for your complete snapshot to be learn again into reminiscence.
The lifecycle of the VM — boot, snapshot, restore, and scale to zero — is the work of Unikraft’s product. We labored carefully with them all through the migration to verify it holds up underneath our request quantity and visitors patterns.
Running and response dealing with
After working this undertaking for a number of years, we already had learnings we included to maximise efficiency and the flexibility to debug. At scale, we’ve run into all types of points, from operating out of ports on digital switches to DNS (we had been stunned, however it wasn’t all the time DNS).
In this iteration, we made certain compute nodes run native DNS resolvers. We’ve additionally expanded the metrics we accumulate, recording issues like boot time, time to first port open, and time to start out person code. There are additionally a number of circuit breakers in place to make sure immediate rerouting and decommissioning of compute nodes.

That’s the entire path, and on a heat occasion it provides about 6ms. None of it leaves our community, and we’re accountable for the entire request cycle. Everything above occurs between the request arriving and the response going again out.
Designing resilient compute transport systems
When you construct a system just like the one described above, you’re optimizing for 2 issues without delay: the end-user expertise and rollout resiliency. We want to have the ability to roll out adjustments shortly however steadiness that with the flexibility to roll again simply as shortly.
The compute nodes are constructed from a base picture revealed by Unikraft and set up a set of packages. These nodes are constructed individually from our edge nodes for a few causes: it retains our edge nodes light-weight and quick, it lets us use totally different occasion varieties for our compute nodes, and it lets us scale these nodes independently.
A management aircraft retains observe of which compute nodes exist and that are wholesome, and the sting nodes ballot it for that listing. It’s additionally what drives a deploy — a brand new fleet comes up alongside the operating one, scales to match it, and takes over visitors solely as soon as it’s wholesome.
Building the compute transport systems has required shut collaboration with the Unikraft staff. Throughout the migration, we’ve labored with them on testing correctness, dealing with giant volumes of requests, and constructing capabilities particular to our platform.
It’s stay
The work to rebuild our edge compute structure is greater than only a velocity increase. It’s a sooner basis we are able to hold constructing on and have extra management over. The better part is that it’s already serving your manufacturing visitors in the present day, on the identical pricing, with no migration step and nothing to alter in any undertaking.
Running the compute ourselves means the ceiling on Edge Functions is ours to lift. Three issues this makes tractable that weren’t earlier than:
- npm package deal assist, out of beta. npm packages work in edge functions today, in beta, with caveats round native binaries and importing recordsdata at runtime. An actual VM with an actual filesystem removes a lot of the causes these caveats exist.
- Room to revisit the operation limits. The documented limits of 50ms of CPU per request, 512MB of reminiscence, and 20MB of compressed code got here from the isolate-based execution mannequin.
- Compute inside our personal community. Anything that depends upon controlling the community path, as a substitute of reaching a 3rd social gathering throughout the web, is now one thing we are able to construct.
We’re not completed right here. The limits and tough edges we couldn’t contact earlier than are those we’re engaged on now, so keep tuned.
