Backups that could not be restored
A backup script ran cleanly every night and stored empty files, until the first restore test. What the script lacked, and how restores are checked now.
- Company
- Language school with online course booking, 14 employees
- Team
- One in-house developer, an external administrator as needed
- Environment
- One VPS, Docker Compose: application, PostgreSQL, Nginx
- Backups
- Nightly via cron: pg_dump to the server's disk and a copy to Amazon S3
- Constraints
- A restore had never been tried; only site availability was monitored
01Starting point
The database runs in a container with its data in a Docker volume. Uploaded files (contracts, invoices, teacher photos) sit in a directory on the server mounted into the application container. A cron script ran pg_dump through docker exec every night, compressed the output with gzip and copied the file to an S3 bucket. The key on the server had full permissions on that bucket.
- Internet
- Nginx
- Application (Docker)
- PostgreSQL (Docker volume)
- cron: pg_dump | gzip
- Amazon S3: backups
02The problem
A system review ahead of expanding the booking system included a restore test: download the latest backup and restore it into a clean container. The latest backup was 20 bytes. So was every backup from the previous seven weeks.
Had a disk failed in that time, or had someone encrypted the server, the school would have lost seven weeks of registrations, payments and timetable changes. Uploaded files could not have been restored at all, because they had never been backed up.
A third problem on top: the key on the server was allowed to delete from the bucket. An attacker who took over the server could have deleted the database and the backups in one go.
03Technical root cause
Seven weeks earlier the database had moved to a new major version of PostgreSQL, and the container got a different name in the process. From then on docker exec in the script failed, because no container with the old name existed.
The script did not notice. Without set -o pipefail, the exit status of the pipeline docker exec ... | gzip > backup.gz is that of its last command, gzip, which succeeded: it compressed empty input. An empty gzip file is exactly 20 bytes.
The error message went to the cron job's error output. Cron mails that output, but no mail was configured on the server, so it was discarded. Uploaded files had been missing from the start; the script only backed up the database.
04Investigation
The file sizes in the bucket showed the exact day the backups became empty. It matched the day of the database version change in the history of the Docker Compose configuration repository.
Running the script by hand printed Error response from daemon: No such container. The system log confirmed that cron had been discarding the job's output because no mail agent was installed.
The last non-empty backup did restore, and the data since then was still in the running database. Nothing was lost. It was a risk, not a loss.
05Remediation
The same day: a manual backup of the database and the files, and confirmation that it restores. Only then the fixes.
The script was rewritten so that a failure in any step stops the whole run, it addresses the Docker Compose service rather than the container name, and it reports its own success.
#!/usr/bin/env bash
set -euo pipefail
source /etc/backup.env # HC_UUID and the AWS credentials, mode 600
STAMP=$(date -u +%Y%m%dT%H%M%SZ)
FILE="/var/backups/db/app-$STAMP.dump"
trap 'rm -f "$FILE"' ERR
cd /srv/app
# the service name survives a renamed container; -T: no pseudo-TTY
docker compose exec -T db pg_dump -U app -Fc app > "$FILE"
# an empty or unreadable archive fails here, not on the day of a restore
docker compose exec -T db pg_restore --list < "$FILE" > /dev/null
aws s3 cp "$FILE" "s3://skola-zalohy/db/app-$STAMP.dump"
aws s3 sync /srv/app/uploads "s3://skola-zalohy/uploads/" # no --delete
# only a run that got this far reports in; silence raises the alarm
curl -fsS -m 10 --retry 3 "https://hc-ping.com/$HC_UUID" > /dev/nullpg_dump -Fc produces a compressed archive in its own format, so gzip is not needed. pg_restore --list checks straight after the backup that the archive is readable. The -T flag turns off the pseudo-terminal that docker compose exec allocates by default, which must be off when binary output is redirected to a file.
Uploaded files are backed up with aws s3 sync without --delete, so deleting a file on the server does not delete its copy.
Protection against deletion: the bucket has versioning enabled and a lifecycle rule that removes old versions after 35 days. The key on the server may only upload (s3:PutObject) and list contents for the sync (s3:ListBucket). It has neither s3:DeleteObject nor s3:DeleteObjectVersion, nor any right to change versioning or rules.
The success report goes to Healthchecks.io. If it does not arrive on time, the service sends an email. Its open-source version on your own server works just as well.
A weekly restore test: a second script downloads the latest backup, restores it into a temporary PostgreSQL container and checks that the newest booking is less than a day old. It reports its result the same way.
06Why the fix works
set -e ends the script at the first error, pipefail makes a failure in the middle of a pipeline count, not just one in the last command, and set -u stops the script on a mistyped variable name. The ERR trap deletes the half-written file so it cannot pass for a backup.
The service name db does not change when the container name does, which was the original cause.
Success is monitored, not failure. The service speaks up when the report does not arrive, which also covers the script not running at all, for example after the cron entry was removed.
With versioning on, overwriting an object creates a new version and the original stays. A key without delete permissions can remove neither the object nor its older versions. An attacker on the server therefore cannot destroy the backups.
What it does not address: a backup is not high availability. When the server goes down, the site stays down until it is restored elsewhere.
07Validation
- A test copy of the script with a deliberately wrong service name failed, deleted its half-written file, and Healthchecks sent an alert.
- A full restore to a clean server following the written procedure, database and files, with the time measured. That time is now in the documentation as the expected downtime for a restore.
- An attempt to delete an object and one of its versions with the server's key ended in
AccessDenied.
08Results
Fixed
- Database backups have content, and a failure shows up within a day.
- Uploaded files are backed up.
Risk reduced
- A compromised server no longer means losing the backups.
- Restoring has been tried and written down rather than assumed.
Still open
- With a nightly backup, the school can lose up to one day of data.
- There is still one server. If it fails, the site is down until it is restored.
09Limitations
pg_restore --listchecks only the header and the list of objects in the archive, not all the data. Corrupt data is caught only by the weekly restore test.- The automated test covers the database. Restoring files is tried by hand once a quarter.
- The backups contain personal data. Access to the bucket and the retention period must match what the school states in its privacy policy.
- Old versions are removed after 35 days. A data error nobody notices for longer cannot be undone from backups.
10Lessons
- A backup that has not been restored is only an assumption.
- A shell script without
set -euo pipefailcan report success on failure. - Monitor success, not only errors. Silence is not good news.
- A server must not be able to delete its own backups.
- Back up everything that cannot be recreated, not just the database.
11Next steps for a small company
- 1.Write the restore procedure so that someone who does not know the server can follow it.
- 2.Add a daytime database backup if losing a day is not acceptable.
- 3.Turn on two-factor authentication for the AWS console and manage backups from a different account than production.
- 4.As the number of applications grows, consider a managed database with automatic backups.
Recognise your own company in this?
Tell us what it involves. We reply within one business day and say whether it is work for us, even when the answer is no.