NestJS course Β· Module 6: Error Handling and Monitoring
Monitoring and Alerting - early warning systems
In this lesson6
The application works. No exception has been thrown, the logs are silent, the health check shines green. And yet users write that "the site drags". You check - responses arrive after eight hundred milliseconds instead of a hundred. Since when? Nobody knows. Logs record events, and this is a trend: something that grew over weeks and is invisible in any single entry.
Rome set signal towers on its borders. They did not report single events - they measured traffic: how many riders passed, how long the crossing took, how many sentries failed to return. Only from those numbers could you see something going wrong before a gate fell. That is what metrics are.
Four types of metric
The standard for collecting metrics is Prometheus - a system that collects and stores numbers describing your application. It offers four kinds of measurement, from simplest to most elaborate:
- Counter - only goes up. Requests served, errors seen. It never decreases; on restart it begins at zero.
- Gauge - can rise and fall. Active connections, memory in use, queue length. It takes the current value, like a needle on a dial.
- Histogram - the distribution of values across buckets. Instead of one number it remembers how many measurements fell into 0-100 ms, how many into 100-500 ms, and so on.
- Summary - like a histogram, but with percentiles computed on the client side, that is inside your application rather than in Prometheus.
For HTTP request duration the right choice is a Histogram. A Counter would say only how many requests there were, a Gauge how long the last one took. A Histogram shows the shape: that nine out of ten requests finish under 200 ms while every tenth exceeds a second. That shape reveals a problem an average would hide.
Defining a metric
You create a metric once, describing it with three fields:
1import { Counter } from 'prom-client';
2
3@Injectable()
4export class PrometheusService {
5 private readonly httpRequestCounter = new Counter({
6 name: 'http_requests_total',
7 help: 'Total HTTP requests',
8 labelNames: ['method', 'route', 'status'],
9 });
10
11 recordRequest(method: string, route: string, status: number) {
12 this.httpRequestCounter.inc({ method, route, status: String(status) });
13 }
14}The order of the fields is conventional but always the same: name is the metric's identifier, help a human-readable description, labelNames the list of dimensions you will be able to slice the data by.
The labels are the interesting part. Thanks to them one counter answers many questions: how many POST requests there were, how many hit /legions, how many ended with a 500. Without labels you would need a separate counter for every combination.
Beware of one trap: a label with many possible values multiplies the number of data series. Putting a user id in a label creates as many series as you have users - and will bring Prometheus down. Labels must have a finite, small set of values.
System metrics for free
Before writing your own metrics, it is worth switching on the built-in ones:
1import { collectDefaultMetrics, register } from 'prom-client';
2
3collectDefaultMetrics();collectDefaultMetrics() collects the default system metrics - processor and memory usage and the Node.js event loop lag. That last one is especially valuable: a growing event loop lag means something is blocking the thread and the application is falling behind, though no endpoint has failed yet.
Note what this function does not do: it sends nothing, draws no charts and resets no counters. It only starts collecting.
The /metrics endpoint
Prometheus does not accept data pushed by an application - it comes for it itself. Your job is to expose it at an agreed address:
1@Controller()
2export class MetricsController {
3 @Get('/metrics')
4 async getMetrics(@Res() res: Response) {
5 res.set('Content-Type', register.contentType);
6 res.send(await register.metrics());
7 }
8}register is the registry of every defined metric, and register.metrics() returns them in the text format Prometheus understands. The Content-Type header must say plain text, not JSON - hence register.contentType, which sets the right value for you.
This model is called scraping: every dozen seconds or so Prometheus queries that address and records what it found. The application need not know who is watching it, or whether anyone is.
Five deployment stages
The whole road from nothing to a chart looks like this:
- Install the
prom-clientpackage. - Define the metrics - Counter, Histogram, as many as needed.
- Create the
/metricsendpoint. - Configure the Prometheus scraper so it knows where and how often to ask.
- Visualise the metrics in Grafana - and only here do charts and alerts appear.
The division of labour between the last two often gets blurred. Prometheus collects and stores, Grafana draws and alerts. That separation lets you replace one without the other - and means the application knows neither.
Summary
The signal towers stand and the traffic is measured:
- logs record events, metrics show trends - the latter is invisible in any single entry,
- Prometheus collects and stores metrics; it does not log errors and repairs nothing,
- four types from simplest: Counter (only rises), Gauge (rises and falls), Histogram (distribution across buckets), Summary (percentiles computed client-side),
- for HTTP request duration the right type is a Histogram - it shows the shape an average would hide,
- three fields define a metric:
name,help,labelNames, - labels let you slice the data but must have a small set of values - a user id as a label will bring Prometheus down,
collectDefaultMetrics()collects system metrics - CPU, memory, event loop lag; it sends nothing,- Prometheus comes for the data itself (scraping), so you expose a
/metricsendpoint withContent-Typeset to plain text viaregister.contentType, - five stages: install
prom-client, define metrics, the/metricsendpoint, configure the scraper, visualise in Grafana, - Prometheus collects, Grafana draws and alerts.
In the next lesson we will drop from trends down to a single bug - you will meet debugging techniques for when you know something is wrong but not where. For now remember: a log says what happened once; a metric says what keeps happening - and it is the one that warns you before a gate falls.
Code for this lesson: src/monitoring/monitoring-system.ts
1// Monitoring and Alerting - Early Warning Systems
2// Metrics, alerts and a dashboard for the Imperium
3import { Injectable, Logger } from '@nestjs/common';
4
5// ===========================================
6// 1. Metrics Service (Prometheus style)
7// ===========================================
8
9@Injectable()
10export class MetricsService {
11 private logger = new Logger('Metrics');
12
13 // Counters - only grow
14 private counters = new Map<string, number>();
15
16 // Histograms - distribution of values
17 private histograms = new Map<string, number[]>();
18
19 // Gauge - instant value (can grow and shrink)
20 private gauges = new Map<string, number>();
21
22 // Increment the counter
23 incrementCounter(name: string, value: number = 1) {
24 const current = this.counters.get(name) || 0;
25 this.counters.set(name, current + value);
26 }
27
28 // Save histogram value
29 observeHistogram(name: string, value: number) {
30 if (!this.histograms.has(name)) {
31 this.histograms.set(name, []);
32 }
33 this.histograms.get(name).push(value);
34 }
35
36 // Set gauge
37 setGauge(name: string, value: number) {
38 this.gauges.set(name, value);
39 }
40
41 // Get all metrics
42 getMetrics() {
43 const result: Record<string, any> = {};
44
45 // Counters
46 this.counters.forEach((v, k) => {
47 result[k + '_total'] = v;
48 });
49
50 // Histograms - average and percentiles
51 this.histograms.forEach((values, k) => {
52 const sorted = [...values].sort((a, b) => a - b);
53 const sum = values.reduce((a, b) => a + b, 0);
54 result[k + '_avg'] = Math.round(sum / values.length);
55 result[k + '_p95'] = sorted[Math.floor(sorted.length * 0.95)];
56 result[k + '_count'] = values.length;
57 });
58
59 // Gauges
60 this.gauges.forEach((v, k) => {
61 result[k] = v;
62 });
63
64 return result;
65 }
66}
67
68// ===========================================
69// 2. Alert Service
70// ===========================================
71
72@Injectable()
73export class AlertService {
74 private logger = new Logger('Alerting');
75 private alerts: Array<{
76 name: string;
77 severity: string;
78 message: string;
79 timestamp: Date;
80 }> = [];
81
82 // Alert threshold definitions
83 private thresholds = {
84 error_rate: { warn: 0.01, critical: 0.05 },
85 response_time_ms: { warn: 500, critical: 2000 },
86 memory_percent: { warn: 80, critical: 95 },
87 };
88
89 checkThreshold(metric: string, value: number) {
90 const threshold = this.thresholds[metric];
91 if (!threshold) return;
92
93 if (value >= threshold.critical) {
94 this.fireAlert(metric, 'CRITICAL', metric + ' = ' + value);
95 } else if (value >= threshold.warn) {
96 this.fireAlert(metric, 'WARNING', metric + ' = ' + value);
97 }
98 }
99
100 private fireAlert(name: string, severity: string, message: string) {
101 const alert = { name, severity, message, timestamp: new Date() };
102 this.alerts.push(alert);
103 this.logger.warn('ALERT [' + severity + ']: ' + message);
104 }
105
106 getActiveAlerts() {
107 return this.alerts.slice(-20);
108 }
109}
110
111console.log('=== Monitoring & Alerting ===');
112console.log('Counter - counts events (requests, errors)');
113console.log('Histogram - distribution of response times');
114console.log('Gauge - current value (memory, connections)');
115console.log('Alert thresholds: warn -> critical');
116Spotted a mistake in this lesson?
Check yourself
Answer the questions from this lesson. Pick an answer to see right away whether it is correct.
1. What is Prometheus used for in the context of monitoring NestJS applications?
2. Which Prometheus metric type is best suited for measuring HTTP request duration?
These are 2 of 3 questions for this lesson. Solve the rest in the game.
Hands-on tasks in the game
- Code editor
Complete the PrometheusService with an httpRequestCounter field (Counter with labelNames: ['method', 'route', 'status_code']) and httpRequestDuration (Histogram with buckets: [0.1, 0.5, 1, 2, 5, 10])
- Vertical ordering
Order Prometheus metric types from simplest to most complex
- Click in order
Arrange the elements of a Counter definition with labelNames in the correct order
- Code editor
Complete the MetricsController with @Get('/metrics') that sets Content-Type to 'text/plain' and returns register.metrics() from prom-client
- Vertical ordering
Order the steps for implementing Prometheus monitoring in a NestJS application from installation to visualization