NestJS course Β· Module 9: Deployment and Infrastructure
Blue-Green Deployment and Zero-Downtime - the art of seamless guard change
In this lesson6
Master of strategy! Senator Cicero says: "The Empire never sleeps, and the roads must always be open." Meanwhile every deployment means half a minute of 502 errors, because the old process has already gone out and the new one has not yet risen. The Roman legions changed the guard without leaving the walls undefended - applications must also update without interruptions.
Blue-Green Deployment is a technique in which we keep two identical production environments (blue and green). At any given moment only one serves traffic, and the other receives the new version. The switch happens instantly, with no downtime.
The Blue-Green Deployment concept
Imagine two city gates. Blue is open and serves travellers, while in green the guards install new fortifications. When green is ready, traffic moves over to it. The configuration of both slots is described by two interfaces:
1// src/deployment/blue-green.config.ts
2interface DeploymentSlot {
3 name: 'blue' | 'green';
4 port: number;
5 version: string;
6 status: 'active' | 'standby' | 'deploying';
7 healthCheckUrl: string;
8 startedAt: Date;
9}
10
11interface BlueGreenConfig {
12 activeSlot: 'blue' | 'green';
13 slots: {
14 blue: DeploymentSlot;
15 green: DeploymentSlot;
16 };
17 loadBalancerUrl: string;
18 healthCheckInterval: number; // ms
19 healthCheckRetries: number;
20 rollbackTimeout: number; // ms
21}
22
23const deploymentConfig: BlueGreenConfig = {
24 activeSlot: 'blue',
25 slots: {
26 blue: {
27 name: 'blue',
28 port: 3001,
29 version: '2.3.0',
30 status: 'active',
31 healthCheckUrl: 'http://blue.imperium.internal:3001/health',
32 startedAt: new Date(),
33 },
34 green: {
35 name: 'green',
36 port: 3002,
37 version: '2.4.0',
38 status: 'standby',
39 healthCheckUrl: 'http://green.imperium.internal:3002/health',
40 startedAt: new Date(),
41 },
42 },
43 loadBalancerUrl: 'http://lb.imperium.internal',
44 healthCheckInterval: 5000,
45 healthCheckRetries: 3,
46 rollbackTimeout: 30000,
47};activeSlot points to the gate with traffic, and the switch is a change of target in the load balancer. A rollback means switching back, because the old version is still running. The price is double resources, and long-lived connections such as WebSockets must be moved deliberately.
Rolling Updates
A rolling update gradually replaces old instances with new ones, like centurions changing the guard, so some instances always serve traffic:
1// src/deployment/rolling-update.service.ts
2import { Injectable, Logger } from '@nestjs/common';
3
4interface Instance {
5 id: string;
6 version: string;
7 status: 'running' | 'updating' | 'ready' | 'failed';
8 port: number;
9}
10
11@Injectable()
12export class RollingUpdateService {
13 private readonly logger = new Logger(RollingUpdateService.name);
14
15 async performRollingUpdate(
16 instances: Instance[],
17 newVersion: string,
18 maxUnavailable: number = 1,
19 ): Promise<void> {
20 this.logger.log(
21 `Starting rolling update to version ${newVersion}`
22 );
23 this.logger.log(
24 `Instances: ${instances.length}, max unavailable: ${maxUnavailable}`
25 );
26
27 // Update instances in batches - the whole batch at once
28 for (let i = 0; i < instances.length; i += maxUnavailable) {
29 const batch = instances.slice(i, i + maxUnavailable);
30 const results = await Promise.all(
31 batch.map((instance) => this.updateInstance(instance, newVersion)),
32 );
33
34 // A failed instance stops the whole rollout
35 if (results.includes(false)) {
36 throw new Error(`Rolling update to version ${newVersion} halted`);
37 }
38 }
39 }Promise.all updates the whole batch at once. The first version had a loop with await inside the batch, so despite the maxUnavailable parameter it always took down one instance at a time. A failed instance now stops the whole rollout instead of breaking the next ones - in a test with six instances the last two were left untouched.
Every instance goes through five steps in a fixed order:
1 private async updateInstance(instance: Instance, newVersion: string): Promise<boolean> {
2 instance.status = 'updating';
3 this.logger.log(`Updating instance ${instance.id}...`);
4
5 // 1. Remove the instance from the load balancer
6 await this.removeFromLoadBalancer(instance);
7
8 // 2. Wait for current requests to finish
9 await this.drainConnections(instance);
10
11 // 3. Update to the new version
12 await this.deployNewVersion(instance, newVersion);
13
14 // 4. Check the health check
15 const healthy = await this.waitForHealthy(instance);
16
17 if (healthy) {
18 instance.status = 'running';
19 instance.version = newVersion;
20 // 5. Add the instance back to the load balancer
21 await this.addToLoadBalancer(instance);
22 this.logger.log(`Instance ${instance.id} updated`);
23 return true;
24 }
25
26 instance.status = 'failed';
27 this.logger.error(
28 `Instance ${instance.id} - rollback!`
29 );
30 await this.rollback(instance);
31 return false;
32 }The instance leaves the load balancer, finishes ongoing requests, gets the new version, passes the health check and only then returns to the pool.
The helper methods are a skeleton to connect to your infrastructure:
1 private async removeFromLoadBalancer(instance: Instance) {
2 // Remove the instance from the load balancer pool
3 }
4
5 private async drainConnections(instance: Instance) {
6 // Wait for active connections to finish (graceful)
7 }
8
9 private async deployNewVersion(
10 instance: Instance, version: string
11 ) {
12 // Deploy the new version to the instance
13 }
14
15 private async waitForHealthy(instance: Instance): Promise<boolean> {
16 const maxRetries = 10;
17 for (let i = 0; i < maxRetries; i++) {
18 try {
19 // Check the /health/ready endpoint of the new version
20 const response = await fetch(`http://localhost:${instance.port}/health/ready`, {
21 signal: AbortSignal.timeout(2000),
22 });
23 if (response.ok) return true;
24 } catch {
25 // The instance is not responding yet
26 }
27 await new Promise(r => setTimeout(r, 2000)); // pause after every attempt
28 }
29 return false;
30 }
31
32 private async addToLoadBalancer(instance: Instance) {
33 // Add the instance back to the pool
34 }
35
36 private async rollback(instance: Instance) {
37 // Restore the previous version and add the instance back to the pool
38 }
39}waitForHealthy() polls /health/ready with a timeout and waits after every attempt. The first version waited only after an exception, so ten failed responses flew by in a fraction of a second.
Health Check-based Deployment
Health checks are the Empire's scouts - they check whether a new fortification is ready before we open the gates:
1// src/health/deployment-health.controller.ts
2import { Controller, Get } from '@nestjs/common';
3import {
4 HealthCheck, HealthCheckService,
5 MongooseHealthIndicator, MemoryHealthIndicator,
6 DiskHealthIndicator,
7} from '@nestjs/terminus';
8
9@Controller('health')
10export class DeploymentHealthController {
11 constructor(
12 private health: HealthCheckService,
13 private mongoose: MongooseHealthIndicator,
14 private memory: MemoryHealthIndicator,
15 private disk: DiskHealthIndicator,
16 ) {}
17
18 // Liveness - is the application alive?
19 @Get('live')
20 @HealthCheck()
21 checkLiveness() {
22 return this.health.check([
23 () => this.memory.checkHeap('memory_heap', 200 * 1024 * 1024),
24 ]);
25 }
26
27 // Readiness - is it ready for traffic?
28 @Get('ready')
29 @HealthCheck()
30 checkReadiness() {
31 return this.health.check([
32 () => this.mongoose.pingCheck('mongodb'),
33 () => this.memory.checkHeap('memory_heap', 200 * 1024 * 1024),
34 () => this.disk.checkStorage('disk', {
35 thresholdPercent: 0.9, path: '/',
36 }),
37 ]);
38 }
39
40 // Startup - did the application start correctly?
41 @Get('startup')
42 @HealthCheck()
43 checkStartup() {
44 return this.health.check([
45 () => this.mongoose.pingCheck('mongodb'),
46 ]);
47 }
48}Liveness says whether the process is alive, readiness whether it can take traffic, and startup whether the application started correctly. TerminusModule.forRoot({ gracefulShutdownTimeoutMs: 10000 }) completes zero-downtime: after SIGTERM readiness returns 503 with the status shutting_down, while the application keeps serving traffic for another 10 seconds. In the test ordinary endpoints answered 200 during that time, so Kubernetes has time to take the pod out of the pool.
Database migrations - zero-downtime
Migrations are the hardest part of a zero-downtime deployment, because old and new code coexist with the same database. The rule: a migration must be backward compatible. ALTER TABLE legions RENAME COLUMN name TO legion_name causes downtime, because the old code looks for name and crashes. A safe field rename takes three deployments. The first adds the new field:
1// Migration 1 (deployment 1): add the new legion_name field
2import { Db } from 'mongodb';
3
4export async function up(db: Db) {
5 await db.collection('legions').updateMany(
6 {},
7 [{ $set: { legion_name: '$name' } }]
8 );
9 // Old code still reads 'name' - it works!
10}The pipeline update copies the value of name into legion_name in every MongoDB document, and the old code notices nothing.
The second deployment changes the code:
1// Code v2 (deployment 2): write to both fields
2async updateLegion(id: string, name: string) {
3 await this.model.updateOne(
4 { _id: id },
5 { $set: { name, legion_name: name } }
6 );
7}
8
9// Read from the new field
10async getLegion(id: string) {
11 const doc = await this.model.findById(id);
12 return doc?.legion_name ?? doc?.name; // Fallback
13}Code v2 writes to both fields and reads the new one, falling back to the old one for documents the migration did not cover.
The third cleans up:
1// Migration 3 (after v2 is fully deployed): remove the old field
2import { Db } from 'mongodb';
3
4export async function up(db: Db) {
5 // First fill in documents written by old code during the rollout
6 await db.collection('legions').updateMany(
7 { legion_name: { $exists: false } },
8 [{ $set: { legion_name: '$name' } }]
9 );
10
11 await db.collection('legions').updateMany(
12 {},
13 { $unset: { name: '' } }
14 );
15}First we fill in documents written by the old code after the first migration. The first version removed name straight away, and in the test a document added during the rollout lost its name.
Feature Flags
Feature flags are the emperor's secret orders - they turn features on and off without deploying new code:
1// src/features/feature-flag.service.ts
2import { Injectable } from '@nestjs/common';
3
4interface FeatureFlag {
5 name: string;
6 enabled: boolean;
7 rolloutPercentage: number; // 0-100
8 allowedUsers: string[];
9}
10
11@Injectable()
12export class FeatureFlagService {
13 private flags: Map<string, FeatureFlag> = new Map();
14
15 setFlag(flag: FeatureFlag) {
16 this.flags.set(flag.name, flag);
17 }
18
19 isEnabled(flagName: string, userId?: string): boolean {
20 const flag = this.flags.get(flagName);
21 if (!flag) return false;
22
23 // Flag completely disabled
24 if (!flag.enabled) return false;
25
26 // User on the allow list (e.g. testers)
27 if (userId && flag.allowedUsers.includes(userId)) {
28 return true;
29 }
30
31 // Gradual rollout (percentage rollout)
32 if (flag.rolloutPercentage < 100) {
33 // Flag name in the hash - every flag picks different users
34 const hash = this.hashUserId(`${flagName}:${userId || 'anonymous'}`);
35 return (hash % 100) < flag.rolloutPercentage;
36 }
37
38 return true;
39 }
40
41 private hashUserId(userId: string): number {
42 let hash = 0;
43 for (let i = 0; i < userId.length; i++) {
44 hash = ((hash << 5) - hash) + userId.charCodeAt(i);
45 hash |= 0;
46 }
47 return Math.abs(hash);
48 }
49}I added setFlag(), because nothing filled the map before. The flag name in the hash makes every flag pick different users: without it, in the test the same 20% of users ended up in every rollout. In production flags are kept in a database or in a service such as Unleash.
The controller asks the service on every request:
1// Usage in a controller
2@Controller('legions')
3class LegionsController {
4 constructor(private features: FeatureFlagService) {}
5
6 @Get()
7 async getLegions(@Req() req) {
8 if (this.features.isEnabled('new-ranking-system', req.user?.id)) {
9 return this.getNewRanking();
10 }
11 return this.getOldRanking();
12 }
13}I left out the ranking methods, and req.user is set by the JWT guard.
PM2 Cluster Mode
PM2 in cluster mode is like deploying many garrisons - it uses all CPU cores, restarts processes and reloads them without downtime:
1// ecosystem.config.cjs - PM2 configuration (.cjs, because a NestJS 12 project is an ES module)
2module.exports = {
3 apps: [{
4 name: 'imperium-api',
5 script: 'dist/main.js',
6 instances: 'max', // Use all CPU cores
7 exec_mode: 'cluster', // Cluster mode
8 autorestart: true,
9 watch: false,
10 max_memory_restart: '500M',
11
12 // Zero-downtime reload
13 wait_ready: true, // Wait for process.send('ready')
14 listen_timeout: 10000, // Max time to become ready
15 kill_timeout: 5000, // Time for graceful shutdown
16
17 env_production: {
18 NODE_ENV: 'production',
19 PORT: 3000,
20 },
21 }],
22};The file has the .cjs extension, because module.exports does not work in a project with "type": "module". wait_ready makes PM2 wait for the readiness signal (3 s by default, here listen_timeout is 10 s), and kill_timeout replaces the default 1.6 s for shutting down.
The readiness signal is sent by main.ts:
1// In main.ts - signalling readiness
2async function bootstrap() {
3 const app = await NestFactory.create(AppModule);
4 app.enableShutdownHooks(); // PM2 stops the process with SIGINT
5 await app.listen(3000);
6
7 // Tell PM2 the application is ready
8 if (process.send) {
9 process.send('ready');
10 }
11}PM2 stops processes with SIGINT, which enableShutdownHooks() turns into an orderly retreat. process.send exists only when the process has an IPC channel to PM2.
PM2 commands for zero-downtime:
1# Start from the configuration file (a .cjs file must be named explicitly)
2pm2 start ecosystem.config.cjs --env production
3
4# Reload without downtime (graceful)
5pm2 reload imperium-api
6
7# Instance status
8pm2 status
9
10# Real-time monitoring
11pm2 monitIn cluster mode pm2 reload restarts processes one by one, so one of them always serves traffic.
Blue-green, rolling updates and feature flags are powerful tools in the arsenal of every architect of the Empire. I recommend starting with rolling updates and good probes and keeping blue-green for systems where a rollback must be instant.
Remember: the guard is changed so that the wall never stands empty, not even for a moment.
Code for this lesson: src/blue-green-deployment.ts
1// Blue-Green Deployment and Zero-Downtime
2// Deployment strategies without downtime
3
4// 1. Blue-Green configuration
5interface DeploymentSlot {
6 name: 'blue' | 'green';
7 port: number;
8 version: string;
9 status: 'active' | 'standby' | 'deploying';
10}
11
12const slots: Record<string, DeploymentSlot> = {
13 blue: {
14 name: 'blue',
15 port: 3001,
16 version: '2.3.0',
17 status: 'active',
18 },
19 green: {
20 name: 'green',
21 port: 3002,
22 version: '2.4.0',
23 status: 'standby',
24 },
25};
26
27// 2. Health Check Controller
28// @Controller('health')
29// class HealthController {
30// @Get('live')
31// liveness() { return { status: 'ok' }; }
32//
33// @Get('ready')
34// readiness() {
35// // TODO: Check MongoDB, memory, disk
36// }
37// }
38
39// 3. Feature Flag Service
40class FeatureFlagService {
41 private flags = new Map<string, { enabled: boolean; rollout: number }>();
42
43 isEnabled(flag: string, userId?: string): boolean {
44 const f = this.flags.get(flag);
45 if (!f || !f.enabled) return false;
46 if (f.rollout < 100 && userId) {
47 // TODO: Implement a canary release
48 // Use hash userId % 100 < rollout
49 }
50 return true;
51 }
52}
53
54// 4. PM2 Cluster Config
55// ecosystem.config.js:
56// {
57// name: 'imperium-api',
58// script: 'dist/main.js',
59// instances: 'max',
60// exec_mode: 'cluster',
61// wait_ready: true,
62// listen_timeout: 10000,
63// }
64
65// TODO: Implement a full rolling update service
66// TODO: Add a database migration strategy (3 steps)
67// TODO: Configure PM2 with graceful shutdown
68Spotted a mistake in this lesson?
Check yourself
Answer the questions from this lesson. Pick an answer to see right away whether it is correct.
1. Blue-Green Deployment consists of:
Hands-on tasks in the game
- Code editor
Create a FeatureFlagService with an isEnabled method supporting rolloutPercentage and blue-green deployment configuration
- Vertical ordering
Arrange the rolling update steps for a single instance:
- Code editor
Create a .github/workflows/deploy.yml file with jobs: test, build, deploy triggered on push to main
- Click in order
Arrange the correct order of steps in a complete CI/CD pipeline:
- Vertical ordering
Arrange the correct health check decorator in a NestJS controller:
- Code editor
Create a configuration with Joi schema validating DATABASE_URL, PORT, NODE_ENV, and JWT_SECRET
- Click in order
Arrange security middleware in the recommended order of adding to main.ts: