← All postsAugust 30, 2026

A failed deploy should never take your site down

We found the worst kind of bug in our own deploy pipeline, the kind you find because it happens to you.

A project was live and serving. A new deploy started, ground through its build for nine minutes, and failed. Fine — builds fail. Except the site went down with it, and stayed down for forty minutes until a later build happened to succeed.

What was wrong

Server apps ran in one container per project. When a new deploy arrived, the pipeline stopped that container first and then built the new code inside it. The reasoning even sounded right: a redeploy must boot the new code, so clear out the old.

The consequence was that the running app died the moment a build started. If the build succeeded, nobody noticed — a short gap, then the new version. If the build failed, there was nothing left to serve. The container sat there holding a broken build, answering every request with an error, until someone shipped again.

The property that fixes it

The fix is not "handle the failure better." It is a structural property:

Traffic follows the newest deployment that reached ready, and nothing else can touch it.

Every deployment now gets its own container. A new deploy builds beside the running one, not inside it. The router answers one question on every request: which deployment of this project most recently reached ready? A build that fails never reaches ready, so traffic never follows it — and the previous deployment keeps serving, untouched, because nothing destroys it anymore.

This is how deploys should work everywhere, and most mature platforms converged on the same shape for the same reason. We are writing it down because the failure mode is so easy to build by accident: one mutable "current version" that gets updated in place is the obvious first design, and it works perfectly right up until the first failed build at a bad time.

What you see as a user

Nothing, which is the point. A failed deploy is now a red row in your deployments list and an honest error in its logs. Your site does not participate in your build's problems. Deploying again is always safe — the message on a timed-out build says exactly that, because we can finally promise it.