NestJS course Β· Module 9: Deployment and Infrastructure
Production Best Practices - the rules of a true centurion
In this lesson6
On production the code does not change. What changes is who is watching: nobody. The application runs at night, on Sundays and on holidays, and all you know about it are the signals it sends about itself. This lesson is about organising them.
A centurion did not inspect each legionary in person. He had three sources of knowledge: numerical reports on the state of the cohorts, watch records of individual incidents, and couriers' dispatches showing where an order had got stuck on the way. The same division holds to this day.
The three pillars of observability
The three pillars of observability are metrics, logs and traces. Not frontend, backend and database, for those are layers of an application. Not CPU, memory and disk, for those are resources - one kind of metric among many. And not development, testing and production, for those are environments.
Each pillar answers a different question, so none replaces the others:
- Metrics tell you how much and how often - requests per second, response time, memory use. They are aggregated and cheap to store, so you keep them for months and build your alerts on them.
- Logs tell you what exactly happened in one particular case - with the error's text and the request's identifier. They are expensive, because there are so many, but only they answer the question "why this request".
- Traces tell you where the time went - they show one request's journey across all the services. Without them, with three services along the way, you know only that the answer took two seconds; you do not know which of them consumed it.
A metric tells you that something is wrong. A log tells you what broke. A trace tells you where.
The health check
An orchestrator - Kubernetes, Docker Swarm, or a plain load balancer - must know whether traffic may be sent to this instance. It asks through a health endpoint:
1@Controller('health')
2export class HealthController {
3 constructor(
4 private health: HealthCheckService,
5 private db: MongooseHealthIndicator,
6 private memory: MemoryHealthIndicator,
7 ) {}
8
9 @Get()
10 @HealthCheck()
11 check() {
12 return this.health.check([
13 () => this.db.pingCheck('database'),
14 () => this.memory.checkHeap('memory_heap', 300 * 1024 * 1024),
15 ]);
16 }
17}The @nestjs/terminus package supplies the parts ready-made: HealthCheckService gathers the results and assembles the response, MongooseHealthIndicator checks the database connection, and MemoryHealthIndicator watches heap use. There are more indicators - for Redis, for the disk, for any external HTTP service.
A health check should check the availability of the database, Redis and dependent services. Not whether the source code is up to date - that is not its role. Not the number of logged-in users - that is a business metric, not a health signal. And certainly not only whether the HTTP server responds.
That last distinction is the heart of it. An application whose database has died still answers HTTP - the process is alive, the port is open, 200 OK comes back without delay. A health check testing only that tells the orchestrator "all is well" about an instance that cannot serve a single real request. Check the dependencies without which the application can do nothing anyway.
Configuring monitoring
A monitoring system is set up in four steps, in this order:
- Install the monitoring agent - it is what gathers data from the machine and the application.
- Configure metrics collection - decide what you measure and how often.
- Set up alerting rules - fix the thresholds beyond which somebody is to be notified.
- Create visualization dashboards - charts to look at once you know something is happening.
The order is often reversed, and that is the commonest mistake in setting monitoring up. Dashboards come last, because until metrics are flowing you do not know what you would draw on them. Alerts come before them, because an alert arrives by itself whereas a dashboard has to be looked at - and at three in the morning nobody is looking.
Alert escalation levels
Not every threshold crossing means the same thing. Alerts fall into four levels:
- Warning (80% capacity) - a notification. Nothing is broken yet, but you are nearing the limit; there is time to react calmly.
- Critical (95% capacity) - an urgent response. The margin is nearly gone; failure is a matter of hours or minutes.
- Emergency (above 100%) - immediate intervention. The limit has been passed and users can already see it.
- Post-incident - analysis and lessons learned. Once the situation is contained you establish the cause and what to change so it does not return.
Note that the first two thresholds lie below a hundred per cent. That is deliberate: an alert at 100% is no longer a warning but a notification of failure. Sensible monitoring gives time to react, not a commentary on the fire.
The fourth level is often skipped, and it is the only one that changes anything for the future. Without a post-incident analysis the same alert will ring again next month.
Graceful shutdown
The last rule concerns switching off. When a new version is deployed the old instance receives a SIGTERM signal - and what it does over the next few seconds decides whether anyone sees an error:
1async function bootstrap() {
2 const app = await NestFactory.create(AppModule);
3
4 app.enableShutdownHooks();
5
6 await app.listen(3000);
7}enableShutdownHooks() makes NestJS intercept the signal and, before closing, call the onModuleDestroy and beforeApplicationShutdown methods in the modules. The application thus has time to finish the requests in flight, close database connections and deregister from the service registry.
Without it a deployment cuts off the requests being served at that moment. At ten deployments a day that is ten bursts of errors nobody connects with the deployment - because in the logs they look like random dropped connections.
Summary
The centurion reads reports rather than inspecting every legionary:
- the three pillars of observability are metrics, logs and traces - not frontend/backend/database, not CPU/memory/disk, not environments,
- a metric tells you that something is wrong, a log what broke, a trace where the time went,
- a health check checks the availability of the database, Redis and dependent services - not the code's freshness, not the user count, and not merely whether HTTP responds,
- an application with a dead database still returns
200- which is why checking HTTP alone is worthless, @nestjs/terminus:HealthCheckServiceassembles the result,MongooseHealthIndicatortests the database,MemoryHealthIndicatorthe heap,- configuring monitoring: install the agent β configure metrics collection β set up alerting rules β create dashboards,
- escalation levels: Warning (80%) β Critical (95%) β Emergency (>100%) β Post-incident,
- thresholds below 100% leave time to react; an alert at 100% is already a report of failure,
enableShutdownHooks()lets in-flight requests finish when an instance is shut down.
That closes our survey of production practice. For now remember: an application in production is only as good as the signals it sends about itself - because nobody is going to guess what is happening to it.
Code for this lesson: src/production-best-practices.ts
1// Production Best Practices - Rules of a True Captain
2import { Injectable, Logger } from '@nestjs/common';
3
4// 1. Structured Logging (not console.log!)
5@Injectable()
6class ProductionLogger {
7 private readonly logger = new Logger('RomanAPI');
8
9 logRequest(method: string, url: string, duration: number) {
10 this.logger.log(JSON.stringify({
11 type: 'request',
12 method,
13 url,
14 duration,
15 timestamp: new Date().toISOString(),
16 }));
17 }
18
19 logError(error: Error, context?: string) {
20 this.logger.error(JSON.stringify({
21 type: 'error',
22 message: error.message,
23 stack: error.stack,
24 context,
25 timestamp: new Date().toISOString(),
26 }));
27 }
28}
29
30// 2. Error Handling - global error filter
31// @Catch()
32// class GlobalExceptionFilter implements ExceptionFilter {
33// catch(exception: any, host: ArgumentsHost) {
34// const ctx = host.switchToHttp();
35// const response = ctx.getResponse();
36// const status = exception.getStatus?.() || 500;
37//
38// // Do NOT expose internal errors!
39// response.status(status).json({
40// statusCode: status,
41// message: status === 500
42// ? 'Internal Server Error'
43// : exception.message,
44// timestamp: new Date().toISOString(),
45// });
46// }
47// }
48
49// 3. Security Checklist
50const securityConfig = {
51 helmet: true, // Secure HTTP headers
52 cors: {
53 origin: ['https://roman-empire.com'],
54 credentials: true,
55 },
56 rateLimit: {
57 ttl: 60,
58 limit: 100,
59 },
60 validation: {
61 whitelist: true,
62 forbidNonWhitelisted: true,
63 },
64 // NEVER in production:
65 // - synchronize: true (database)
66 // - debug: true
67 // - console.log with sensitive data
68 // - hardcoded secrets
69};
70
71// 4. Graceful Shutdown
72// app.enableShutdownHooks();
73// @Injectable()
74// class CleanupService implements OnModuleDestroy {
75// async onModuleDestroy() {
76// // Close database connections
77// // Close Redis connections
78// // Close open streams
79// console.log('Graceful shutdown complete');
80// }
81// }
82
83// 5. Monitoring Endpoints
84const monitoringEndpoints = {
85 '/health': 'Basic health check',
86 '/health/ready': 'Readiness to handle traffic',
87 '/health/live': 'Is the application alive',
88 '/metrics': 'Prometheus metrics',
89};
90Spotted a mistake in this lesson?
Check yourself
Answer the questions from this lesson. Pick an answer to see right away whether it is correct.
1. The three pillars of observability are:
2. A health check endpoint in production should check:
Hands-on tasks in the game
- Code editor
Create a HealthController checking the database and memory
- Click in order
Arrange the monitoring system configuration steps:
- Vertical ordering
Arrange the monitoring alert escalation levels: