Borg Backup
Another solo topic today as I’ve had “a week” work-wise and a weekend full of nor’easter prep, a sick spouse, and time to stand up the topic we’re covering today for my new self-hosted bare Git + cgit repo setup.
Safe-keeping My Bare Git Repos
I’ve been forced by some performative outrage of others into self-hosting Git repos (again) and have been documenting my AI scraper woes of the public cgit web setup on Mastodon (so I won’t bore you about that here).
An issue with/concern for self-hosting Git repos is the need to ensure there are decent offsite backups, and I have finally configured those thanks to the cheap, non-public-facing storage box service of Hetzner. Said storage boxes support Borg Backup, so today we’re going to talk about it and the setup I have chosen.
My cgit serves 13 public repos (more coming soon!) and (for now) two private ones. Hetzner offers system backups, but I’m more concerned with being able to quickly recover from me fat-fingering some commands and wreaking havoc in the bare git dir.
Hetzner storage boxes provide a very thin set of services, but they are a pretty solid set of services for the purpose this offering serves.
For four U.S. dollars a month, you get unlimited internal network traffic, 10 snapshots, up to 100 sub-accounts, and 10 concurrent connections. It has full support for all FTP, SFTP, SCP, Samba/CIFS, BorgBackup, Restic, Rclone, rsync via SSH, HTTPS, and WebDAV. You can also use it as a network drive with free email support.
With the box provisioned with the ssh key from the Git host server, it was time to be assimilated.
About the collective
Borg Backup (usually just Borg) is a deduplicating backup program with compression and authenticated encryption. It is BSD licensed and has been around since about 2015 and is a descendant of Attic.
You may want to keep the Borg site up to cross reference with some of the blather, below.
It ships as a single binary, and there is no daemon, no agent, no database, and no “server”, save for the linux install on Hetzner’s storage box. A borg repo is a directory that can live on a local disk, a USB drive, or any machine you can reach over SSH.
These repos hold archives, which contain a full view of the files named on the command line along with their metadata. Underneath all that are items (one per file, directory, or symlink) and data chunks (file contents cut into pieces).
Deduplication runs across the entire repo. This includes files, archives, and machines that write into the same repository. Nothing in the design depends on an incremental chain, so every archive restores on its own.
Borg cuts file content with a content-defined chunker, and the default setup uses the Buzhash rolling hash with parameters buzhash,19,23,21,4095. (You kind of need to read or at least skim that article for the following to make sense.) The cuts land at content-dependent boundaries instead of fixed offsets, so an insertion at the start of a file does not shift every later chunk and defeat deduplication. A fixed chunker exists for disk images and block devices.
Each chunk gets an ID from a keyed hash of its contents: HMAC-SHA256 for repokey and keyfile, and keyed BLAKE2b-256 for the -blake2 modes. Before writing, Borg looks the ID up in its chunks cache, which is a fancy shorthand for saying “a hash table of chunks that already exist.” If there’s a known ID, there’s no write. So, the second hourly archive of my bare Git repo over unchanged Git objects cost a mere extra 642 bytes. (More on the setup in a bit.)
Encryption runs on the client, with the repos only ever holding ciphertext, so the host holding it never needs your trust. There are a few modes for this:
| Mode | Encryption | Authentication |
|---|---|---|
repokey / keyfile | AES-256-CTR | HMAC-SHA256, encrypt-then-MAC |
repokey-blake2 / keyfile-blake2 | AES-256-CTR | BLAKE2b-256 |
authenticated / authenticated-blake2 | none | HMAC-SHA256 or BLAKE2b-256 |
none | none | none |
You should likely use either of the repokey options for encryption, which means you need to come up with a decent passphrase and ensure you have a backup of it (your password manager is a good choice for this).
NOTE: The chunk ID hash is keyed as well, so two repositories with different keys never share chunk IDs, even for identical file content.
As noted, the repo is just a filesystem key-value store built from segments of roughly 500 MB. Each segment is a log of PUT, DELETE, and COMMIT transaction entries, all of which are only durable only after a COMMIT lands, so a crash during a write leaves no half-written archive behind.
DELETE appends an entry, and the data stays on disk until you run a compaction command. Borg 1.2 stopped compacting automatically, so borg compact became a scheduled step that my client Git server runs after every prune.
The choices I made
I chose to deploy the storage box with no public internet access. That doesn’t mean it’s “safe” since Hetzner has many users and all the boxes have visibiltiy into the Hetzner network. It does limit the attack surface.
Encryption is repokey-blake2 because it’s modern and fast and solid. Plus, the data is useless without the passphrase.
I run all the backups as an unprivileged user. Said backup are hourly, and I went with the lz4 algorithm since Git packfiles are already zlib-compressed.
For retention, I went with --keep-hourly 48 --keep-daily 7 --keep-weekly 4 --keep-monthly 6 and always do a borg compact after each run.
I usually use systemd timers, but wanted to keep this separate from the rest of the services on the server, so it’s a good ol’ cron job.
I used flock -n on the hourly runs so there’s no accidental overlap (in the event of a slow network or the emergence of a large new repo). For the weekly check I do flock -w 1800, so it will never run over a backup. There’s also a timeout on this one (30m), which also triggers a ntfy.sh notification.
The out-of-band copies of the passphrase live in my password manager and OS keychain.
FIN
I chose not to bore you with detailed commands because the Borg docs are super spiffy. However, if you do decide to take on the Borg, hit me up on Mastodon, and we can slide over to a Signal chat where I’ll be more than happy to help with anything that’s confusing or odd.
FIN
Remember, you can follow and interact with the full text of The Daily Drop’s free posts on:
- Mastodon via
@dailydrop.hrbrmstr.dev@dailydrop.hrbrmstr.dev - Bluesky via
<https://bsky.app/profile/dailydrop.hrbrmstr.dev.web.brid.gy>
☮️
Leave a Reply