NestJS course Β· Module 8: Caching and Performance

Load Balancing - distributing legion load

11 min read
In this lesson5

Commander of the legions! Consul Caesar.js has noticed that one fort does all the work while the others sit idle in the garrison. When a wave of reports arrives, that single fort falls, and the whole application falls with it. Time to learn load balancing - the art of spreading the load across all units of the army.

What is load balancing in the legionaries' world?

Imagine an army of Roman legions:

  • one fort serves all clients = overload,
  • units stand idle = wasted resources,
  • uneven load = some forts collapse under their tasks,
  • no redundancy = when the main fort falls, the whole army stops.

A load balancer is a wise legate who distributes tasks among the forts. In practice this role is usually played by a ready-made server in front of the application, such as Nginx, HAProxy or a cloud load balancer, which often also takes over HTTPS and response compression. When nobody in front of the application compresses, in NestJS the app.use(compression()) middleware does it, reducing the size of the HTTP responses sent - we will come back to it in the lesson on response optimization. Below we build a balancer in NestJS itself to see how the algorithms work.

The instance side: the /health endpoint

The balancer has to know which forts are alive. So every instance exposes a simple health endpoint:

1// health.controller.ts - on every instance
2import { Controller, Get } from '@nestjs/common';
3
4@Controller('health')
5export class HealthController {
6  @Get()
7  check() {
8    const memory = process.memoryUsage();
9    return {
10      status: 'ok',
11      uptime: Math.round(process.uptime()),
12      heapUsedMB: Math.round(memory.heapUsed / 1024 / 1024),
13      rssMB: Math.round(memory.rss / 1024 / 1024),
14    };
15  }
16}

process.uptime() gives the seconds since the process started, and process.memoryUsage() the memory usage. The response time is measured by the balancer itself: if /health answers slowly or not at all, the fort drops out of the rotation.

Application Load Balancing

The balancer service keeps a list of servers described by the ServerInstance interface: address, weight, current load and health status:

1// load-balancer.service.ts
2import { Injectable, OnModuleInit } from '@nestjs/common';
3import { ConfigService } from '@nestjs/config';
4import { Cron } from '@nestjs/schedule';
5
6interface ServerInstance {
7  id: string;
8  host: string;
9  port: number;
10  weight: number;
11  currentLoad: number;
12  isHealthy: boolean;
13  lastHealthCheck: Date;
14}
15
16@Injectable()
17export class LoadBalancerService implements OnModuleInit {
18  private servers: ServerInstance[] = [];
19  private currentIndex = 0;
20
21  constructor(private configService: ConfigService) {
22    this.initializeServers();
23  }
24
25  async onModuleInit() {
26    await this.startHealthChecks(); // the first check right after startup
27  }
28
29  private initializeServers(): void {
30    // LEGION_SERVERS is JSON with the server list, e.g. [{"host":"...","port":3000,"weight":10}]
31    const fromEnv = this.configService.get<string>('LEGION_SERVERS');
32    const serverConfigs = fromEnv ? JSON.parse(fromEnv) : [
33      { host: 'cohort-1.legion.local', port: 3000, weight: 10 },
34      { host: 'cohort-2.legion.local', port: 3000, weight: 15 },
35      { host: 'cohort-3.legion.local', port: 3000, weight: 8 },
36    ];
37
38    this.servers = serverConfigs.map((config, index) => ({
39      id: `cohort-${index + 1}`,
40      ...config,
41      currentLoad: 0,
42      isHealthy: true,
43      lastHealthCheck: new Date(),
44    }));
45  }

We read the list from the LEGION_SERVERS variable, and environment variables are text, so JSON.parse() comes first. The first health check starts in onModuleInit(); the original version fired it in the constructor, which cannot wait for anything.

The two simplest selection algorithms:

1  // Round Robin - each fort in turn
2  getNextServerRoundRobin(): ServerInstance | null {
3    const healthyServers = this.servers.filter(s => s.isHealthy);
4    if (healthyServers.length === 0) return null;
5
6    const server = healthyServers[this.currentIndex % healthyServers.length];
7    this.currentIndex++;
8
9    console.log(`Selected fort: ${server.id} (Round Robin)`);
10    return server;
11  }
12
13  // Weighted random - a weighted draw: a stronger cohort is picked more often
14  getNextServerWeighted(): ServerInstance | null {
15    const healthyServers = this.servers.filter(s => s.isHealthy);
16    if (healthyServers.length === 0) return null;
17
18    const totalWeight = healthyServers.reduce((sum, s) => sum + s.weight, 0);
19    let randomWeight = Math.random() * totalWeight;
20
21    for (const server of healthyServers) {
22      randomWeight -= server.weight;
23      if (randomWeight <= 0) {
24        console.log(`Selected fort: ${server.id} (Weighted)`);
25        return server;
26      }
27    }
28
29    return healthyServers[0];
30  }

Round Robin takes the forts in turn. The second method is a weighted draw, not "Weighted Round Robin" as the old comment claimed: in a test of 2,500 draws, weights 10 and 15 produced 986 and 1,514 selections.

The next two look at the load:

1  // Least Connections - the least loaded fort
2  getNextServerLeastConnections(): ServerInstance | null {
3    const healthyServers = this.servers.filter(s => s.isHealthy);
4    if (healthyServers.length === 0) return null;
5
6    const leastLoaded = healthyServers.reduce((min, server) =>
7      server.currentLoad < min.currentLoad ? server : min
8    );
9
10    console.log(`Selected fort: ${leastLoaded.id} (Least Connections: ${leastLoaded.currentLoad})`);
11    return leastLoaded;
12  }
13
14  // Resource-based - based on system resources
15  async getNextServerResourceBased(): Promise<ServerInstance | null> {
16    const healthyServers = this.servers.filter(s => s.isHealthy);
17    if (healthyServers.length === 0) return null;
18
19    // Fetch resource metrics for every cohort
20    const serversWithMetrics = await Promise.all(
21      healthyServers.map(async (server) => {
22        const metrics = await this.getServerMetrics(server);
23        return {
24          ...server,
25          cpuUsage: metrics.cpu,
26          memoryUsage: metrics.memory,
27          responseTime: metrics.avgResponseTime,
28          score: this.calculateServerScore(metrics),
29        };
30      })
31    );
32
33    // Pick the fort with the lowest score (least loaded)
34    const bestServer = serversWithMetrics.reduce((best, server) =>
35      server.score < best.score ? server : best
36    );
37
38    console.log(`Selected fort: ${bestServer.id} (Resource-based, score: ${bestServer.score})`);
39    return bestServer;
40  }

Least Connections picks the fort with the fewest ongoing requests. The resource-based variant asks every server for metrics on every request, which is costly - in production metrics are refreshed in the background and kept in memory.

Scoring, counters and fetching metrics:

1  private calculateServerScore(metrics: any): number {
2    // A simple formula: higher value = worse server
3    return (
4      metrics.cpu * 0.4 + // 40% weight for CPU
5      metrics.memory * 0.3 + // 30% weight for memory
6      metrics.avgResponseTime * 0.2 + // 20% weight for response time
7      metrics.activeConnections * 0.1 // 10% weight for connections
8    );
9  }
10
11  async incrementLoad(serverId: string): Promise<void> {
12    const server = this.servers.find(s => s.id === serverId);
13    if (server) {
14      server.currentLoad++;
15    }
16  }
17
18  async decrementLoad(serverId: string): Promise<void> {
19    const server = this.servers.find(s => s.id === serverId);
20    if (server && server.currentLoad > 0) {
21      server.currentLoad--;
22    }
23  }
24
25  private async getServerMetrics(server: ServerInstance): Promise<any> {
26    try {
27      // HTTP call to the metrics endpoint, 2 seconds at most
28      const response = await fetch(`http://${server.host}:${server.port}/health/metrics`, {
29        signal: AbortSignal.timeout(2000),
30      });
31      return await response.json();
32    } catch (error) {
33      // No answer = the worst possible metrics, so this fort is picked last
34      return {
35        cpu: 100,
36        memory: 100,
37        avgResponseTime: 10000,
38        activeConnections: 1000,
39      };
40    }
41  }

The score mixes percentages with milliseconds, so the response time easily dominates the rest - it is worth bringing metrics to a common scale before weighting. When there is no answer we return the worst possible metrics; the first version drew them at random, so a dead fort could win. AbortSignal.timeout(2000) cancels a call that takes too long, because the native Node.js fetch does not know a timeout option.

The health check repeats every 30 seconds:

1  @Cron('*/30 * * * * *') // Every 30 seconds
2  private async startHealthChecks(): Promise<void> {
3    for (const server of this.servers) {
4      try {
5        const response = await fetch(`http://${server.host}:${server.port}/health`, {
6          signal: AbortSignal.timeout(5000),
7        });
8
9        server.isHealthy = response.ok;
10        server.lastHealthCheck = new Date();
11
12        if (!server.isHealthy) {
13          console.warn(`Cohort ${server.id} is not responding!`);
14        }
15      } catch (error) {
16        server.isHealthy = false;
17        server.lastHealthCheck = new Date();
18        console.error(`Cohort ${server.id} unreachable:`, error.message);
19      }
20    }
21  }
22
23  getLegionStatus() {
24    return {
25      totalServers: this.servers.length,
26      healthyServers: this.servers.filter(s => s.isHealthy).length,
27      totalLoad: this.servers.reduce((sum, s) => sum + s.currentLoad, 0),
28      servers: this.servers.map(s => ({
29        id: s.id,
30        host: s.host,
31        isHealthy: s.isHealthy,
32        currentLoad: s.currentLoad,
33        weight: s.weight,
34        lastHealthCheck: s.lastHealthCheck,
35      })),
36    };
37  }
38}

The @Cron decorator comes from @nestjs/schedule and works only once ScheduleModule.forRoot() is imported. The pattern has six fields and the first one is seconds, so */30 * * * * * means every 30 seconds.

Proxy Load Balancer Implementation

The controller accepts every request under /api/proxy and forwards it to the chosen fort through http-proxy-middleware:

1// proxy-load-balancer.controller.ts
2import { All, Controller, Req, Res } from '@nestjs/common';
3import type { Request, Response } from 'express';
4import { createProxyMiddleware, fixRequestBody } from 'http-proxy-middleware';
5import { LoadBalancerService } from './load-balancer.service';
6
7@Controller('/api/proxy')
8export class ProxyLoadBalancerController {
9  constructor(private loadBalancer: LoadBalancerService) {}
10
11  @All('*path')
12  async proxyRequest(@Req() req: Request, @Res() res: Response): Promise<void> {
13    // Pick a server with the load balancer
14    const targetServer = await this.loadBalancer.getNextServerResourceBased();
15
16    if (!targetServer) {
17      res.status(503).json({
18        error: 'All cohorts of the legion are unavailable!',
19        code: 'LEGION_UNAVAILABLE'
20      });
21      return;
22    }
23
24    // Increase the load counter
25    await this.loadBalancer.incrementLoad(targetServer.id);
26
27    // Create a proxy to the chosen server
28    const proxy = createProxyMiddleware<Request, Response>({
29      target: `http://${targetServer.host}:${targetServer.port}`,
30      changeOrigin: true,
31      pathRewrite: {
32        '^/api/proxy': '', // Remove the prefix
33      },
34      on: {
35        proxyReq: (proxyReq, req) => {
36          // Add load balancer information
37          proxyReq.setHeader('X-Proxy-Server', targetServer.id);
38          proxyReq.setHeader('X-Load-Balancer', 'legionariusze-legion-lb');
39          console.log(`Forwarding ${req.method} ${req.url} to ${targetServer.id}`);
40          // NestJS has already read the body - restore it last, after the headers
41          fixRequestBody(proxyReq, req);
42        },
43        proxyRes: async (proxyRes) => {
44          // Decrease the load counter when done
45          await this.loadBalancer.decrementLoad(targetServer.id);
46          console.log(`Response from ${targetServer.id}: ${proxyRes.statusCode}`);
47        },
48        error: async (err) => {
49          await this.loadBalancer.decrementLoad(targetServer.id);
50          console.error(`Proxy error for ${targetServer.id}:`, err.message);
51
52          if (!res.headersSent) {
53            res.status(502).json({
54              error: 'Error communicating with the fort',
55              server: targetServer.id,
56              details: err.message
57            });
58          }
59        },
60      },
61    });
62
63    proxy(req, res);
64  }
65}

The code is adapted to http-proxy-middleware 4: events go in the on object, and the old onProxyReq options and friends no longer work. fixRequestBody is necessary because NestJS has already read the request body - without it the POST in the test hung with no response. It is called last, because it writes the body, and after that headers can no longer be set. '*path' is the named wildcard required since Express 5, and we import the Express types with import type, which TypeScript requires when isolatedModules and emitDecoratorMetadata are enabled.

Creating a new proxy for every request is wasteful; the router option lets a single proxy choose the target for each request.

Circuit Breaker Pattern

Every server gets its own breaker, kept in a map:

1// circuit-breaker.service.ts
2@Injectable()
3export class CircuitBreakerService {
4  private breakers = new Map<string, CircuitBreaker>();
5
6  getOrCreateBreaker(serverId: string): CircuitBreaker {
7    if (!this.breakers.has(serverId)) {
8      this.breakers.set(serverId, new CircuitBreaker({
9        serverId,
10        failureThreshold: 5, // 5 errors
11        recoveryTimeout: 30000, // 30 seconds
12        monitoringPeriod: 10000, // 10 seconds
13      }));
14    }
15    return this.breakers.get(serverId)!;
16  }
17
18  async executeWithBreaker<T>(
19    serverId: string,
20    operation: () => Promise<T>
21  ): Promise<T> {
22    const breaker = this.getOrCreateBreaker(serverId);
23    return await breaker.execute(operation);
24  }
25
26  getBreakerStatus(serverId: string) {
27    const breaker = this.breakers.get(serverId);
28    return breaker ? breaker.getStatus() : null;
29  }
30
31  getAllBreakersStatus() {
32    const status = {};
33    this.breakers.forEach((breaker, serverId) => {
34      status[serverId] = breaker.getStatus();
35    });
36    return status;
37  }
38}

getOrCreateBreaker() creates a breaker on first use, and executeWithBreaker() passes the operation through it, so the failure of one fort does not cut off the others. The breaker itself counts the errors of a given server:

1class CircuitBreaker {
2  private state: 'CLOSED' | 'OPEN' | 'HALF_OPEN' = 'CLOSED';
3  private failures = 0;
4  private lastFailureTime?: Date;
5  private successCount = 0;
6
7  constructor(private config: {
8    serverId: string;
9    failureThreshold: number;
10    recoveryTimeout: number;
11    monitoringPeriod: number;
12  }) {}
13
14  async execute<T>(operation: () => Promise<T>): Promise<T> {
15    if (this.state === 'OPEN') {
16      if (this.shouldAttemptReset()) {
17        this.state = 'HALF_OPEN';
18        console.log(`Circuit breaker for ${this.config.serverId}: attempting reset`);
19      } else {
20        throw new Error(`Circuit breaker OPEN for server ${this.config.serverId}`);
21      }
22    }
23
24    try {
25      const result = await operation();
26      this.onSuccess();
27      return result;
28    } catch (error) {
29      this.onFailure();
30      throw error;
31    }
32  }
33
34  private onSuccess(): void {
35    this.failures = 0;
36    this.successCount++;
37
38    if (this.state === 'HALF_OPEN') {
39      this.state = 'CLOSED';
40      console.log(`Circuit breaker for ${this.config.serverId}: restored`);
41    }
42  }
43
44  private onFailure(): void {
45    this.failures++;
46    this.lastFailureTime = new Date();
47
48    if (this.failures >= this.config.failureThreshold) {
49      this.state = 'OPEN';
50      console.log(`Circuit breaker for ${this.config.serverId}: OPENED`);
51    }
52  }
53
54  private shouldAttemptReset(): boolean {
55    if (!this.lastFailureTime) return false;
56
57    const timeSinceLastFailure = Date.now() - this.lastFailureTime.getTime();
58    return timeSinceLastFailure >= this.config.recoveryTimeout;
59  }
60
61  getStatus() {
62    return {
63      serverId: this.config.serverId,
64      state: this.state,
65      failures: this.failures,
66      successCount: this.successCount,
67      lastFailureTime: this.lastFailureTime,
68    };
69  }
70}

After five errors in a row the circuit opens, and a success resets the counter. monitoringPeriod is waiting for a time window that this version does not compute yet.

I recommend a division of roles: let Nginx or a cloud balancer steer the traffic, and let the application answer /health honestly. In the next lesson we will look into the memory storehouses.

Remember: a good legate does not send every report to one fort - he spreads them according to strength and avoids the forts that stay silent.

Code for this lesson: src/load-balancing.ts
1// Load Balancing - Distributing the Legion's Load
2
3// 1. Nginx as a load balancer (configuration)
4const nginxConfig = `
5upstream roman_legion {
6    # Round Robin (default) - one after another
7    server app1:3000;
8    server app2:3000;
9    server app3:3000;
10
11    # Least Connections - to the least loaded one
12    # least_conn;
13
14    # IP Hash - same client -> same server
15    # ip_hash;
16
17    # Weighted - servers of differing power
18    # server app1:3000 weight=3;
19    # server app2:3000 weight=2;
20    # server app3:3000 weight=1;
21}
22
23server {
24    listen 80;
25    server_name roman-empire.com;
26
27    location / {
28        proxy_pass http://roman_legion;
29        proxy_set_header Host \$host;
30        proxy_set_header X-Real-IP \$remote_addr;
31    }
32}
33`;
34
35// 2. PM2 - cluster mode in Node.js
36const pm2Config = {
37  apps: [{
38    name: 'roman-empire-api',
39    script: 'dist/main.js',
40    instances: 'max',    // Use all CPUs
41    exec_mode: 'cluster',
42    env_production: {
43      NODE_ENV: 'production',
44      PORT: 3000,
45    },
46  }],
47};
48
49// 3. Kubernetes Horizontal Pod Autoscaler
50const hpaYaml = `
51apiVersion: autoscaling/v2
52kind: HorizontalPodAutoscaler
53metadata:
54  name: roman-api-hpa
55spec:
56  scaleTargetRef:
57    apiVersion: apps/v1
58    kind: Deployment
59    name: roman-api
60  minReplicas: 2
61  maxReplicas: 10
62  metrics:
63  - type: Resource
64    resource:
65      name: cpu
66      target:
67        type: Utilization
68        averageUtilization: 70
69`;
70
71// 4. Health Check for the Load Balancer
72import { Controller, Get } from '@nestjs/common';
73
74@Controller('health')
75export class HealthController {
76  @Get()
77  check() {
78    return {
79      status: 'healthy',
80      uptime: process.uptime(),
81      pid: process.pid,
82      memory: Math.round(process.memoryUsage().heapUsed / 1024 / 1024),
83    };
84  }
85
86  @Get('ready')
87  ready() {
88    return { status: 'ready' };
89  }
90}
91

Spotted a mistake in this lesson?

Check yourself

Answer the questions from this lesson. Pick an answer to see right away whether it is correct.

  1. 1. The compression module in NestJS (app.use(compression())) causes:

Hands-on tasks in the game

  • Code editor

    Create a /health endpoint with uptime, memory, and response time information

Useful articles