NestJS course Β· Module 9: Deployment and Infrastructure
Backup and Recovery - rescuing tributes after a disaster
In this lesson4
Guardian of the tributes! One night a migration with a typo deletes a column, and in the morning it turns out the last working copy of the database is three weeks old, because the backup script had been failing for a month and nobody noticed. Even the best forts can fall into ruins during a siege, which is why every wise legionary has a plan for rescuing his tributes.
Two numbers where everything starts
A recovery plan starts with two goals. RTO (Recovery Time Objective) is the maximum time to restore the service after a failure. RPO (Recovery Point Objective) is the acceptable data loss, measured in time: a backup once a day means you lose up to 24 hours of work. Want less - you need more frequent copies or continuous archiving of the change log.
Keep copies according to the 3-2-1 rule: three copies, on two different media, one of them outside the main location.
Database Backup Strategy
The PostgreSQL script makes a dump, checks it, sends it to S3 and cleans up old copies:
1#!/bin/bash
2# backup-database.sh
3set -euo pipefail
4
5# Without a database address the script fails instead of dumping a default database
6: "${DATABASE_URL:?Set DATABASE_URL}"
7BACKUP_DIR="${BACKUP_DIR:-/backups}"
8DATE=$(date +%Y%m%d_%H%M%S)
9BACKUP_NAME="legionariusze_legion_backup_${DATE}.dump"
10
11# Create backup (the custom format is already compressed)
12echo "Creating database backup..."
13pg_dump --format=custom --file="$BACKUP_DIR/$BACKUP_NAME" "$DATABASE_URL"
14
15# Verify backup - pg_restore must be able to read the table of contents
16pg_restore --list "$BACKUP_DIR/$BACKUP_NAME" > /dev/null
17
18# Upload to S3
19aws s3 cp "$BACKUP_DIR/$BACKUP_NAME" s3://legionariusze-legion-backups/
20
21# Cleanup old backups (keep last 7 days)
22find "$BACKUP_DIR" -name "*.dump" -mtime +7 -delete
23
24echo "Backup completed: ${BACKUP_NAME}"set -euo pipefail aborts the script on the first error, an unset variable or a failure in a pipeline. The :? line is a fix: in the first version an empty DATABASE_URL made pg_dump silently connect to the default local database. The custom format is already compressed, so the separate gzip is gone, and pg_restore --list checks that the file can be read.
In MongoDB, which many NestJS applications run on, the same strategy looks like this:
1#!/bin/bash
2# backup-mongo.sh - the same strategy for MongoDB
3set -euo pipefail
4
5: "${MONGODB_URI:?Set MONGODB_URI}"
6BACKUP_DIR="${BACKUP_DIR:-/backups/mongo}"
7TIMESTAMP=$(date +%Y%m%d_%H%M%S)
8ARCHIVE="$BACKUP_DIR/imperium_$TIMESTAMP.archive.gz"
9
10mkdir -p "$BACKUP_DIR"
11
12# Dump the database into a single compressed archive
13mongodump --uri="$MONGODB_URI" --archive="$ARCHIVE" --gzip
14
15# Verification: a trial restore without writing data
16mongorestore --uri="$MONGODB_URI" --archive="$ARCHIVE" --gzip --dryRun
17
18# Checksum for verifying the copy after transfer
19sha256sum "$ARCHIVE" > "$ARCHIVE.sha256"
20
21# Rotation: delete copies older than 7 days
22find "$BACKUP_DIR" -name "imperium_*" -mtime +7 -delete
23
24echo "Backup completed: $ARCHIVE"mongodump writes the database into a single compressed archive, mongorestore --dryRun tries to read it without writing any data, and the checksum lets you verify the copy after transfer. A dry run is not a test, though: a copy you have never restored into a separate database is only hope.
Disaster Recovery Plan
It is worth writing the recovery plan as code, so that nobody skips a step under stress:
1// Recovery service
2@Injectable()
3export class RecoveryService {
4 async restoreFromBackup(backupFile: string) {
5 // 1. Download backup from S3
6 const backup = await this.downloadBackup(backupFile);
7
8 // 2. Create new database
9 await this.createRecoveryDatabase();
10
11 // 3. Restore data
12 await this.restoreData(backup);
13
14 // 4. Verify integrity
15 await this.verifyDataIntegrity();
16
17 // 5. Switch traffic
18 await this.switchTraffic();
19 }
20}The service methods are a skeleton to fill in for your infrastructure. The order matters: you create the new database next to the damaged one, not on top of it, to preserve evidence for analysis, and you check integrity with record counts, a checksum and an application smoke test before switching traffic. The whole disaster recovery process goes like this:
- detect and assess the incident,
- activate the disaster recovery plan,
- restore services from backups,
- validate integrity and resume operations.
Infrastructure as code
Data is half of the recovery - the other half is infrastructure. Infrastructure as Code (IaC), e.g. Terraform or Pulumi, describes servers, networks and databases in files kept in a repository. It gives versioning, reproducibility and automation: after a disaster you recreate the environment with a single command, even in another region, instead of clicking through a console from memory.
The application image must be just as trustworthy. A secure Docker image build pipeline has these stages:
- an optimized Dockerfile with a multi-stage build,
- security and vulnerability scanning, e.g. Trivy,
- tagging the image and pushing it to a registry,
- image signing and attestation, e.g. with cosign.
The signature is created after the push, because it refers to a specific image digest in the registry; during recovery you deploy only images with a valid signature.
I recommend practising a full restore on a fresh environment once a month and measuring how long it takes - it is the only honest test of your RTO. In the project that closes the module you will combine backups with the whole deployment.
Remember: you can rebuild a fort from plans, but you cannot recover tributes without a copy that someone has actually checked.
Code for this lesson: scripts/backup.sh
1#!/bin/bash
2# Backup and Recovery - Saving the Tributes after a Disaster
3
4# === CONFIGURATION ===
5BACKUP_DIR="/backups/roman-empire"
6DB_NAME="roman_empire"
7DB_HOST="localhost"
8DB_PORT="27017"
9RETENTION_DAYS=7
10TIMESTAMP=$(date +%Y%m%d_%H%M%S)
11BACKUP_PATH="$BACKUP_DIR/$TIMESTAMP"
12
13echo "=== Roman Empire Database Backup ==="
14echo "Timestamp: $TIMESTAMP"
15echo "Database: $DB_NAME"
16
17# === 1. MONGODUMP ===
18mkdir -p "$BACKUP_PATH"
19
20echo "Starting mongodump..."
21mongodump \
22 --host "$DB_HOST" \
23 --port "$DB_PORT" \
24 --db "$DB_NAME" \
25 --out "$BACKUP_PATH"
26
27# Check whether the backup succeeded
28if [ $? -eq 0 ]; then
29 echo "Backup successful!"
30else
31 echo "ERROR: Backup FAILED!"
32 exit 1
33fi
34
35# === 2. COMPRESSION ===
36echo "Compressing backup..."
37tar -czf "$BACKUP_PATH.tar.gz" -C "$BACKUP_DIR" "$TIMESTAMP"
38rm -rf "$BACKUP_PATH"
39echo "Compressed: $BACKUP_PATH.tar.gz"
40
41# === 3. ROTATION (remove old backups) ===
42echo "Removing backups older than $RETENTION_DAYS days..."
43find "$BACKUP_DIR" -name "*.tar.gz" \
44 -mtime +$RETENTION_DAYS -delete
45
46# === 4. RESTORE (recovery) ===
47# mongorestore \
48# --host "$DB_HOST" \
49# --port "$DB_PORT" \
50# --db "$DB_NAME" \
51# --drop \
52# "$BACKUP_PATH/$DB_NAME"
53
54# === 5. CRON - automatic backups ===
55# Add to crontab:
56# 0 2 * * * /scripts/backup.sh >> /var/log/backup.log 2>&1
57# (daily at 2:00 AM)
58
59# === 6. BACKUP STRATEGY ===
60# - Daily: full backup every 24h
61# - Incremental: changes every 6h
62# - Offsite: a copy on an external server
63# - Test restore: test recovery every month!
64
65echo "=== Backup Complete ==="
66echo "File: $BACKUP_PATH.tar.gz"
67echo "Size: $(du -h "$BACKUP_PATH.tar.gz" | cut -f1)"
68Spotted a mistake in this lesson?
Check yourself
Answer the questions from this lesson. Pick an answer to see right away whether it is correct.
1. Infrastructure as Code (IaC) enables:
2. RTO (Recovery Time Objective) and RPO (Recovery Point Objective) define:
Hands-on tasks in the game
- Code editor
Write a backup script with mongodump, rotation, and integrity verification
- Vertical ordering
Arrange the disaster recovery process steps:
- Vertical ordering
Arrange the stages of a secure Docker image build pipeline: