NestJS course Β· Module 9: Deployment and Infrastructure

Backup and Recovery - rescuing tributes after a disaster

5 min read
In this lesson4

Guardian of the tributes! One night a migration with a typo deletes a column, and in the morning it turns out the last working copy of the database is three weeks old, because the backup script had been failing for a month and nobody noticed. Even the best forts can fall into ruins during a siege, which is why every wise legionary has a plan for rescuing his tributes.

Two numbers where everything starts

A recovery plan starts with two goals. RTO (Recovery Time Objective) is the maximum time to restore the service after a failure. RPO (Recovery Point Objective) is the acceptable data loss, measured in time: a backup once a day means you lose up to 24 hours of work. Want less - you need more frequent copies or continuous archiving of the change log.

Keep copies according to the 3-2-1 rule: three copies, on two different media, one of them outside the main location.

Database Backup Strategy

The PostgreSQL script makes a dump, checks it, sends it to S3 and cleans up old copies:

1#!/bin/bash
2# backup-database.sh
3set -euo pipefail
4
5# Without a database address the script fails instead of dumping a default database
6: "${DATABASE_URL:?Set DATABASE_URL}"
7BACKUP_DIR="${BACKUP_DIR:-/backups}"
8DATE=$(date +%Y%m%d_%H%M%S)
9BACKUP_NAME="legionariusze_legion_backup_${DATE}.dump"
10
11# Create backup (the custom format is already compressed)
12echo "Creating database backup..."
13pg_dump --format=custom --file="$BACKUP_DIR/$BACKUP_NAME" "$DATABASE_URL"
14
15# Verify backup - pg_restore must be able to read the table of contents
16pg_restore --list "$BACKUP_DIR/$BACKUP_NAME" > /dev/null
17
18# Upload to S3
19aws s3 cp "$BACKUP_DIR/$BACKUP_NAME" s3://legionariusze-legion-backups/
20
21# Cleanup old backups (keep last 7 days)
22find "$BACKUP_DIR" -name "*.dump" -mtime +7 -delete
23
24echo "Backup completed: ${BACKUP_NAME}"

set -euo pipefail aborts the script on the first error, an unset variable or a failure in a pipeline. The :? line is a fix: in the first version an empty DATABASE_URL made pg_dump silently connect to the default local database. The custom format is already compressed, so the separate gzip is gone, and pg_restore --list checks that the file can be read.

In MongoDB, which many NestJS applications run on, the same strategy looks like this:

1#!/bin/bash
2# backup-mongo.sh - the same strategy for MongoDB
3set -euo pipefail
4
5: "${MONGODB_URI:?Set MONGODB_URI}"
6BACKUP_DIR="${BACKUP_DIR:-/backups/mongo}"
7TIMESTAMP=$(date +%Y%m%d_%H%M%S)
8ARCHIVE="$BACKUP_DIR/imperium_$TIMESTAMP.archive.gz"
9
10mkdir -p "$BACKUP_DIR"
11
12# Dump the database into a single compressed archive
13mongodump --uri="$MONGODB_URI" --archive="$ARCHIVE" --gzip
14
15# Verification: a trial restore without writing data
16mongorestore --uri="$MONGODB_URI" --archive="$ARCHIVE" --gzip --dryRun
17
18# Checksum for verifying the copy after transfer
19sha256sum "$ARCHIVE" > "$ARCHIVE.sha256"
20
21# Rotation: delete copies older than 7 days
22find "$BACKUP_DIR" -name "imperium_*" -mtime +7 -delete
23
24echo "Backup completed: $ARCHIVE"

mongodump writes the database into a single compressed archive, mongorestore --dryRun tries to read it without writing any data, and the checksum lets you verify the copy after transfer. A dry run is not a test, though: a copy you have never restored into a separate database is only hope.

Disaster Recovery Plan

It is worth writing the recovery plan as code, so that nobody skips a step under stress:

1// Recovery service
2@Injectable()
3export class RecoveryService {
4  async restoreFromBackup(backupFile: string) {
5    // 1. Download backup from S3
6    const backup = await this.downloadBackup(backupFile);
7
8    // 2. Create new database
9    await this.createRecoveryDatabase();
10
11    // 3. Restore data
12    await this.restoreData(backup);
13
14    // 4. Verify integrity
15    await this.verifyDataIntegrity();
16
17    // 5. Switch traffic
18    await this.switchTraffic();
19  }
20}

The service methods are a skeleton to fill in for your infrastructure. The order matters: you create the new database next to the damaged one, not on top of it, to preserve evidence for analysis, and you check integrity with record counts, a checksum and an application smoke test before switching traffic. The whole disaster recovery process goes like this:

  1. detect and assess the incident,
  2. activate the disaster recovery plan,
  3. restore services from backups,
  4. validate integrity and resume operations.

Infrastructure as code

Data is half of the recovery - the other half is infrastructure. Infrastructure as Code (IaC), e.g. Terraform or Pulumi, describes servers, networks and databases in files kept in a repository. It gives versioning, reproducibility and automation: after a disaster you recreate the environment with a single command, even in another region, instead of clicking through a console from memory.

The application image must be just as trustworthy. A secure Docker image build pipeline has these stages:

  1. an optimized Dockerfile with a multi-stage build,
  2. security and vulnerability scanning, e.g. Trivy,
  3. tagging the image and pushing it to a registry,
  4. image signing and attestation, e.g. with cosign.

The signature is created after the push, because it refers to a specific image digest in the registry; during recovery you deploy only images with a valid signature.

I recommend practising a full restore on a fresh environment once a month and measuring how long it takes - it is the only honest test of your RTO. In the project that closes the module you will combine backups with the whole deployment.

Remember: you can rebuild a fort from plans, but you cannot recover tributes without a copy that someone has actually checked.

Code for this lesson: scripts/backup.sh
1#!/bin/bash
2# Backup and Recovery - Saving the Tributes after a Disaster
3
4# === CONFIGURATION ===
5BACKUP_DIR="/backups/roman-empire"
6DB_NAME="roman_empire"
7DB_HOST="localhost"
8DB_PORT="27017"
9RETENTION_DAYS=7
10TIMESTAMP=$(date +%Y%m%d_%H%M%S)
11BACKUP_PATH="$BACKUP_DIR/$TIMESTAMP"
12
13echo "=== Roman Empire Database Backup ==="
14echo "Timestamp: $TIMESTAMP"
15echo "Database: $DB_NAME"
16
17# === 1. MONGODUMP ===
18mkdir -p "$BACKUP_PATH"
19
20echo "Starting mongodump..."
21mongodump \
22  --host "$DB_HOST" \
23  --port "$DB_PORT" \
24  --db "$DB_NAME" \
25  --out "$BACKUP_PATH"
26
27# Check whether the backup succeeded
28if [ $? -eq 0 ]; then
29  echo "Backup successful!"
30else
31  echo "ERROR: Backup FAILED!"
32  exit 1
33fi
34
35# === 2. COMPRESSION ===
36echo "Compressing backup..."
37tar -czf "$BACKUP_PATH.tar.gz" -C "$BACKUP_DIR" "$TIMESTAMP"
38rm -rf "$BACKUP_PATH"
39echo "Compressed: $BACKUP_PATH.tar.gz"
40
41# === 3. ROTATION (remove old backups) ===
42echo "Removing backups older than $RETENTION_DAYS days..."
43find "$BACKUP_DIR" -name "*.tar.gz" \
44  -mtime +$RETENTION_DAYS -delete
45
46# === 4. RESTORE (recovery) ===
47# mongorestore \
48#   --host "$DB_HOST" \
49#   --port "$DB_PORT" \
50#   --db "$DB_NAME" \
51#   --drop \
52#   "$BACKUP_PATH/$DB_NAME"
53
54# === 5. CRON - automatic backups ===
55# Add to crontab:
56# 0 2 * * * /scripts/backup.sh >> /var/log/backup.log 2>&1
57# (daily at 2:00 AM)
58
59# === 6. BACKUP STRATEGY ===
60# - Daily: full backup every 24h
61# - Incremental: changes every 6h
62# - Offsite: a copy on an external server
63# - Test restore: test recovery every month!
64
65echo "=== Backup Complete ==="
66echo "File: $BACKUP_PATH.tar.gz"
67echo "Size: $(du -h "$BACKUP_PATH.tar.gz" | cut -f1)"
68

Spotted a mistake in this lesson?

Check yourself

Answer the questions from this lesson. Pick an answer to see right away whether it is correct.

  1. 1. Infrastructure as Code (IaC) enables:

  2. 2. RTO (Recovery Time Objective) and RPO (Recovery Point Objective) define:

Hands-on tasks in the game

  • Code editor

    Write a backup script with mongodump, rotation, and integrity verification

  • Vertical ordering

    Arrange the disaster recovery process steps:

  • Vertical ordering

    Arrange the stages of a secure Docker image build pipeline:

Useful articles