NestJS course Β· Module 6: Error Handling and Monitoring

Monitoring and Alerting - early warning systems

6 min read
In this lesson6

The application works. No exception has been thrown, the logs are silent, the health check shines green. And yet users write that "the site drags". You check - responses arrive after eight hundred milliseconds instead of a hundred. Since when? Nobody knows. Logs record events, and this is a trend: something that grew over weeks and is invisible in any single entry.

Rome set signal towers on its borders. They did not report single events - they measured traffic: how many riders passed, how long the crossing took, how many sentries failed to return. Only from those numbers could you see something going wrong before a gate fell. That is what metrics are.

Four types of metric

The standard for collecting metrics is Prometheus - a system that collects and stores numbers describing your application. It offers four kinds of measurement, from simplest to most elaborate:

  1. Counter - only goes up. Requests served, errors seen. It never decreases; on restart it begins at zero.
  2. Gauge - can rise and fall. Active connections, memory in use, queue length. It takes the current value, like a needle on a dial.
  3. Histogram - the distribution of values across buckets. Instead of one number it remembers how many measurements fell into 0-100 ms, how many into 100-500 ms, and so on.
  4. Summary - like a histogram, but with percentiles computed on the client side, that is inside your application rather than in Prometheus.

For HTTP request duration the right choice is a Histogram. A Counter would say only how many requests there were, a Gauge how long the last one took. A Histogram shows the shape: that nine out of ten requests finish under 200 ms while every tenth exceeds a second. That shape reveals a problem an average would hide.

Defining a metric

You create a metric once, describing it with three fields:

1import { Counter } from 'prom-client';
2
3@Injectable()
4export class PrometheusService {
5  private readonly httpRequestCounter = new Counter({
6    name: 'http_requests_total',
7    help: 'Total HTTP requests',
8    labelNames: ['method', 'route', 'status'],
9  });
10
11  recordRequest(method: string, route: string, status: number) {
12    this.httpRequestCounter.inc({ method, route, status: String(status) });
13  }
14}

The order of the fields is conventional but always the same: name is the metric's identifier, help a human-readable description, labelNames the list of dimensions you will be able to slice the data by.

The labels are the interesting part. Thanks to them one counter answers many questions: how many POST requests there were, how many hit /legions, how many ended with a 500. Without labels you would need a separate counter for every combination.

Beware of one trap: a label with many possible values multiplies the number of data series. Putting a user id in a label creates as many series as you have users - and will bring Prometheus down. Labels must have a finite, small set of values.

System metrics for free

Before writing your own metrics, it is worth switching on the built-in ones:

1import { collectDefaultMetrics, register } from 'prom-client';
2
3collectDefaultMetrics();

collectDefaultMetrics() collects the default system metrics - processor and memory usage and the Node.js event loop lag. That last one is especially valuable: a growing event loop lag means something is blocking the thread and the application is falling behind, though no endpoint has failed yet.

Note what this function does not do: it sends nothing, draws no charts and resets no counters. It only starts collecting.

The /metrics endpoint

Prometheus does not accept data pushed by an application - it comes for it itself. Your job is to expose it at an agreed address:

1@Controller()
2export class MetricsController {
3  @Get('/metrics')
4  async getMetrics(@Res() res: Response) {
5    res.set('Content-Type', register.contentType);
6    res.send(await register.metrics());
7  }
8}

register is the registry of every defined metric, and register.metrics() returns them in the text format Prometheus understands. The Content-Type header must say plain text, not JSON - hence register.contentType, which sets the right value for you.

This model is called scraping: every dozen seconds or so Prometheus queries that address and records what it found. The application need not know who is watching it, or whether anyone is.

Five deployment stages

The whole road from nothing to a chart looks like this:

  1. Install the prom-client package.
  2. Define the metrics - Counter, Histogram, as many as needed.
  3. Create the /metrics endpoint.
  4. Configure the Prometheus scraper so it knows where and how often to ask.
  5. Visualise the metrics in Grafana - and only here do charts and alerts appear.

The division of labour between the last two often gets blurred. Prometheus collects and stores, Grafana draws and alerts. That separation lets you replace one without the other - and means the application knows neither.

Summary

The signal towers stand and the traffic is measured:

  • logs record events, metrics show trends - the latter is invisible in any single entry,
  • Prometheus collects and stores metrics; it does not log errors and repairs nothing,
  • four types from simplest: Counter (only rises), Gauge (rises and falls), Histogram (distribution across buckets), Summary (percentiles computed client-side),
  • for HTTP request duration the right type is a Histogram - it shows the shape an average would hide,
  • three fields define a metric: name, help, labelNames,
  • labels let you slice the data but must have a small set of values - a user id as a label will bring Prometheus down,
  • collectDefaultMetrics() collects system metrics - CPU, memory, event loop lag; it sends nothing,
  • Prometheus comes for the data itself (scraping), so you expose a /metrics endpoint with Content-Type set to plain text via register.contentType,
  • five stages: install prom-client, define metrics, the /metrics endpoint, configure the scraper, visualise in Grafana,
  • Prometheus collects, Grafana draws and alerts.

In the next lesson we will drop from trends down to a single bug - you will meet debugging techniques for when you know something is wrong but not where. For now remember: a log says what happened once; a metric says what keeps happening - and it is the one that warns you before a gate falls.

Code for this lesson: src/monitoring/monitoring-system.ts
1// Monitoring and Alerting - Early Warning Systems
2// Metrics, alerts and a dashboard for the Imperium
3import { Injectable, Logger } from '@nestjs/common';
4
5// ===========================================
6// 1. Metrics Service (Prometheus style)
7// ===========================================
8
9@Injectable()
10export class MetricsService {
11  private logger = new Logger('Metrics');
12
13  // Counters - only grow
14  private counters = new Map<string, number>();
15
16  // Histograms - distribution of values
17  private histograms = new Map<string, number[]>();
18
19  // Gauge - instant value (can grow and shrink)
20  private gauges = new Map<string, number>();
21
22  // Increment the counter
23  incrementCounter(name: string, value: number = 1) {
24    const current = this.counters.get(name) || 0;
25    this.counters.set(name, current + value);
26  }
27
28  // Save histogram value
29  observeHistogram(name: string, value: number) {
30    if (!this.histograms.has(name)) {
31      this.histograms.set(name, []);
32    }
33    this.histograms.get(name).push(value);
34  }
35
36  // Set gauge
37  setGauge(name: string, value: number) {
38    this.gauges.set(name, value);
39  }
40
41  // Get all metrics
42  getMetrics() {
43    const result: Record<string, any> = {};
44
45    // Counters
46    this.counters.forEach((v, k) => {
47      result[k + '_total'] = v;
48    });
49
50    // Histograms - average and percentiles
51    this.histograms.forEach((values, k) => {
52      const sorted = [...values].sort((a, b) => a - b);
53      const sum = values.reduce((a, b) => a + b, 0);
54      result[k + '_avg'] = Math.round(sum / values.length);
55      result[k + '_p95'] = sorted[Math.floor(sorted.length * 0.95)];
56      result[k + '_count'] = values.length;
57    });
58
59    // Gauges
60    this.gauges.forEach((v, k) => {
61      result[k] = v;
62    });
63
64    return result;
65  }
66}
67
68// ===========================================
69// 2. Alert Service
70// ===========================================
71
72@Injectable()
73export class AlertService {
74  private logger = new Logger('Alerting');
75  private alerts: Array<{
76    name: string;
77    severity: string;
78    message: string;
79    timestamp: Date;
80  }> = [];
81
82  // Alert threshold definitions
83  private thresholds = {
84    error_rate: { warn: 0.01, critical: 0.05 },
85    response_time_ms: { warn: 500, critical: 2000 },
86    memory_percent: { warn: 80, critical: 95 },
87  };
88
89  checkThreshold(metric: string, value: number) {
90    const threshold = this.thresholds[metric];
91    if (!threshold) return;
92
93    if (value >= threshold.critical) {
94      this.fireAlert(metric, 'CRITICAL', metric + ' = ' + value);
95    } else if (value >= threshold.warn) {
96      this.fireAlert(metric, 'WARNING', metric + ' = ' + value);
97    }
98  }
99
100  private fireAlert(name: string, severity: string, message: string) {
101    const alert = { name, severity, message, timestamp: new Date() };
102    this.alerts.push(alert);
103    this.logger.warn('ALERT [' + severity + ']: ' + message);
104  }
105
106  getActiveAlerts() {
107    return this.alerts.slice(-20);
108  }
109}
110
111console.log('=== Monitoring & Alerting ===');
112console.log('Counter - counts events (requests, errors)');
113console.log('Histogram - distribution of response times');
114console.log('Gauge - current value (memory, connections)');
115console.log('Alert thresholds: warn -> critical');
116

Spotted a mistake in this lesson?

Check yourself

Answer the questions from this lesson. Pick an answer to see right away whether it is correct.

  1. 1. What is Prometheus used for in the context of monitoring NestJS applications?

  2. 2. Which Prometheus metric type is best suited for measuring HTTP request duration?

These are 2 of 3 questions for this lesson. Solve the rest in the game.

Hands-on tasks in the game

  • Code editor

    Complete the PrometheusService with an httpRequestCounter field (Counter with labelNames: ['method', 'route', 'status_code']) and httpRequestDuration (Histogram with buckets: [0.1, 0.5, 1, 2, 5, 10])

  • Vertical ordering

    Order Prometheus metric types from simplest to most complex

  • Click in order

    Arrange the elements of a Counter definition with labelNames in the correct order

  • Code editor

    Complete the MetricsController with @Get('/metrics') that sets Content-Type to 'text/plain' and returns register.metrics() from prom-client

  • Vertical ordering

    Order the steps for implementing Prometheus monitoring in a NestJS application from installation to visualization

Useful articles