Backup & Restore¶
The operator backs up a Nextcloud instance in two halves, both written to the same S3 bucket under the same per-instance prefix:
| Half | What it holds | Mechanism | Guide |
|---|---|---|---|
| Database | The PostgreSQL cluster, continuously | pgBackRest repo2 (S3), managed by Percona |
Managed PostgreSQL |
| Configuration | config/, encryption keys, declarative state, operator-owned Secrets, custom apps, themes |
Encrypted bundle uploaded by the operator | This page |
You need both halves
A database dump restored into a fresh instance is unusable without
config/config.php's instanceid, secret and passwordsalt. Restore the
configuration bundle first, then the database.
Neither half covers the tenant's file data in the S3 data bucket, or local
data/ volumes. Those are separate concerns — see What is not backed
up.
<bucket>/
└── <instance-name>/
└── backups/
├── db/ ← pgBackRest repo2 (Percona writes this)
└── config/
├── 20260825T031500Z.tar.enc
└── 20260826T031500Z.tar.enc
The configuration bundle¶
What is in it¶
Priority-ordered — the first two are the ones no restore works without:
| # | Item | Why it matters |
|---|---|---|
| 1 | config/ — the whole directory, not just config.php (so every *.config.php drop-in comes too) |
Holds instanceid, secret and passwordsalt. Without them there is no restore, only a new instance. |
| 2 | data/files_encryption/, when the server-side encryption app is active |
Encrypted files are unrecoverable noise without the keys. |
| 3 | Declarative state — the Nextcloud CR, the NextcloudInstance spec, the applied profile, and the operator/chart/Nextcloud versions |
Rebuild the instance shell on a clean cluster; input for restore tooling. |
| 4 | Operator-owned Secrets — admin, mail, the S3 data-bucket credentials, the OIDC client secret | Reattach primary storage and accounts without credential archaeology. |
| 5 | custom_apps/ |
Store apps are re-installable, but version drift is real and genuinely custom apps are in no store. |
| 6 | themes/ |
Custom theming is not reproducible from the spec. |
| 7 | manifest.json — inventory, per-item SHA-256 checksums, versions, warnings |
How a restore drill verifies what it actually got. |
Inline credentials in the captured spec (item 3) are replaced with a redaction marker: the real values live in item 4, exactly once.
What is not backed up¶
| Not included | Where it belongs instead |
|---|---|
| The tenant's files in the S3 data bucket | Bucket versioning / replication, or the restic backups.data path |
Local data/ volumes (other than the encryption keys) |
Volume snapshots |
data/appdata* previews and caches |
Regenerated on demand |
| Redis contents | Cache — nothing to restore |
The HelmRelease and rendered values |
Regenerated from the spec |
| Database dumps | pgBackRest owns the database |
When it runs¶
- Daily, at a per-instance hour derived from a stable hash of the instance name, so a whole estate does not upload to one bucket in the same minute.
- Once immediately after pool assignment. A just-assigned instance is the one most likely to be restored soon, and its pre-assignment bundle describes a spare rather than a tenant.
It runs only when both a bucket and an encryption key resolve — see Configuration. No key means no bundle, never an unencrypted one: this artifact carries the config triad, the encryption keys and the operator-owned Secrets.
Configuration¶
The bucket and key resolve exactly as for the database
backups: the merged instance spec
first, then the operator's own backup.dbDefaultS3.* values. Configure the operator
once and every instance is covered.
Per-instance knobs:
spec:
backups:
config:
enabled: true # code default: on, whenever a bucket and a key resolve
retention: 14 # code default: keep 14 bundles, prune older after each upload
Size guard¶
The bundle streams out of the running pod in bounded chunks and is capped at roughly
200 MiB. Items 1 and 2 (config/, the encryption keys) and items 5 and 6
(custom_apps/, themes/) are streamed separately, so an oversized
custom_apps/:
- drops items 5 and 6 from that bundle,
- records the reason in
manifest.jsonunderwarnings, and repeats it on theConfigBackupCompletedevent, - still uploads a complete, restorable bundle of everything critical.
An oversized or missing config/ fails the run instead: uploading a bundle that
cannot restore, while the freshness metric reports success, would be worse than no
bundle at all.
Encryption format¶
The bundle is an OpenSSL enc container — Salted__ + an 8-byte random salt, key and
IV derived with PBKDF2-HMAC-SHA256 (10 000 iterations), AES-256-CBC with PKCS#7
padding. That is deliberately not a bespoke format: a restore happens on someone's
worst day, possibly without this operator and without Python, so decryption has to be
possible with stock tooling.
enc containers are unauthenticated by design, so integrity is verified against the
per-item SHA-256 checksums in manifest.json, not by the container itself.
Security model¶
Worth stating plainly, because both properties are accepted design, not oversights — if either is unacceptable for a given tenant, give that tenant its own bucket via the per-instance override.
Tenant separation inside a shared bucket is a convention, not enforcement¶
Every instance's backups live under its own key prefix — /<namespace>/<instance>/backups/db/ and
/<namespace>/<instance>/backups/config/ — and that prefix is what keeps one instance's data
distinguishable from another's. But with the operator-central default, all instances
share one bucket and one S3 credential. Nothing at the object-store layer stops that
credential from reading or writing another instance's prefix; separation rests on the
operator always constructing the right prefix, not on IAM.
What that does and does not buy you:
- Tenants themselves have no access. The bucket credential lives in the operator's
namespace and is rendered into each instance namespace only in pgBackRest's
s3.conf, for the database repo host. A Nextcloud tenant never receives it. - A compromise of the operator, or of the shared credential, exposes every instance's backups in that bucket — including the archive of instances unrelated to the breach.
- A prefix bug is a cross-tenant bug. This is why the operator emits an explicit
per-instance
repo2-path(Percona would otherwise use one fleet-wide path) and why bundle pruning refuses to act on any key outside the instance's own prefix.
Give a tenant a genuinely isolated bucket — with its own IAM credential scoped to it —
by setting database.postgres.backup.s3.bucket on that instance. It overrides the
central default for both halves.
Secret-reference confinement¶
A Secret named by an instance spec always resolves in that instance's own
namespace. credentialsFrom.secretRef and encryption.keyFrom.secretRef have no
namespace subfield, in the CRD schema or in the operator.
The operator holds cluster-wide read on Secrets while an instance spec is
tenant-writable, so honouring a spec-supplied namespace would be a confused deputy: an
ordinary CR write would make the operator read any Secret in any namespace — the
db-backup-cipher above included — and materialize its contents into a Secret in the
tenant's own namespace, where the tenant can read it. The only cross-namespace source
is the operator's own Helm values.
Restoring¶
1. Fetch and decrypt a bundle¶
INSTANCE=nc-happy-sun-a1b2c3
BUCKET=nextcloud-db-backups
# Newest bundle for this instance
aws s3 ls "s3://$BUCKET/$INSTANCE/backups/config/" --endpoint-url https://s3.example.com
aws s3 cp "s3://$BUCKET/$INSTANCE/backups/config/20260825T031500Z.tar.enc" . \
--endpoint-url https://s3.example.com
# The passphrase is the SAME key as the pgBackRest cipher-pass for this environment.
kubectl -n nextcloud-operator-system get secret db-backup-cipher \
-o jsonpath='{.data.cipher-pass}' | base64 -d > cipher.key
chmod 600 cipher.key
openssl enc -d -aes-256-cbc -pbkdf2 -pass file:cipher.key \
-in 20260825T031500Z.tar.enc -out bundle.tar
mkdir bundle && tar -xf bundle.tar -C bundle && ls bundle
2. Verify it before trusting it¶
cd bundle
python3 - <<'EOF'
import hashlib, json
manifest = json.load(open("manifest.json"))
print(manifest["instance"], manifest["timestamp"], manifest["versions"])
for item in manifest["items"]:
actual = hashlib.sha256(open(item["name"], "rb").read()).hexdigest()
print(("OK " if actual == item["sha256"] else "FAIL"), item["name"])
for warning in manifest["warnings"]:
print("WARNING:", warning)
EOF
Any FAIL, or an unexpected warnings entry, means stop and use another bundle.
3. Rebuild the instance shell¶
python3 -c "import json;print(json.load(open('bundle/declarative-state.json'))['nextcloudInstance'])"
Re-apply the captured NextcloudInstance (and Nextcloud) spec. Credentials in that
copy are redacted — restore them from bundle/secrets.json, which holds the real
admin, mail, S3 data-bucket and OIDC values.
Handle the extracted bundle like a live credential
secrets.json and files-critical.tar.gz contain plaintext credentials and the
passwordsalt/secret triad. Extract on an encrypted volume, and delete the
working directory when you are done.
4. Restore the configuration files¶
kubectl -n <namespace> cp bundle/files-critical.tar.gz <pod>:/tmp/
kubectl -n <namespace> exec <pod> -- tar -xzf /tmp/files-critical.tar.gz -C /var/www/html
kubectl -n <namespace> exec <pod> -- rm /tmp/files-critical.tar.gz
files-optional.tar.gz (custom apps, themes) restores the same way, if present.
5. Restore the database¶
The database is restored with pgBackRest through Percona's own restore path, from
repo2. Follow the Percona PG Operator restore procedure for the cluster
<instance>-pg, selecting repo2 and the backup set nearest the bundle's timestamp.
The repository is encrypted with the same passphrase as above.
6. Verify¶
kubectl -n <namespace> exec <pod> -- php occ status
kubectl -n <namespace> exec <pod> -- php occ user:list
Write the runbook from a drill, not from this page
Restore steps that have never been executed are a hypothesis. Run a restore drill against a scratch instance, record the real timings, and keep the RTO number that comes out of it — that is what an RPO/RTO commitment can be based on.
Monitoring¶
| Metric | Meaning |
|---|---|
nextcloud_operator_config_backup_timestamp_seconds{namespace,instance} |
Unix timestamp of the last successful bundle upload |
pgbackrest_last_backup_completion_timestamp_seconds{namespace,instance,repo,type} |
Last successful pgBackRest backup, per repository and type |
Once S3 is the durable copy, alert on repo="repo2":
# Config bundle older than two days
time() - nextcloud_operator_config_backup_timestamp_seconds > 172800
# No full database backup on S3 in the last 8 days
time() - pgbackrest_last_backup_completion_timestamp_seconds{repo="repo2",type="full"} > 691200
Both series are removed when an instance is deleted, including when a stuck deletion is force-completed by removing the finalizer — so a firing alert always refers to a resource that still exists.
Events¶
| Event | Meaning |
|---|---|
BackupScheduleApplied |
The effective database backup configuration changed and was applied |
BackupRepo2Bootstrapped |
The one-time initial full backup on the S3 repository was triggered |
ConfigBackupCompleted |
A bundle was uploaded; carries the object key, size and any warnings |
ConfigBackupFailed |
A bundle run failed; the reason is redacted of credentials and deduplicated |
Status fields¶
| Field | Meaning |
|---|---|
configBackup.lastSuccess |
Timestamp of the last successful upload |
configBackup.lastKey |
Object key of that bundle |
configBackup.lastFailure / lastFailureReason |
Last failure and its redacted reason |
configBackup.assignedRunDone |
The post-assignment bundle has been taken |
backupRepo2Bootstrapped |
The initial S3 full backup was triggered |
A failure never clears lastSuccess or lastKey: that bundle is still in the bucket,
and blanking the fields would make one bad day look like "no backups have ever run".
Key custody¶
Both halves are encrypted with the same passphrase per environment, held in a Secret in the operator's namespace.
- Generate with
openssl rand -base64 48. - Store it in the password manager as well as in the cluster. A backup encrypted with a key you have lost is noise.
- Record in your own runbook where it lives.
- Rotation is a known limitation of a shared key: rotating it makes existing backups readable only with the old key. Keep retired keys for at least as long as the backups they can open, and record which key covers which date range.
Rotating either central Secret re-syncs every instance namespace on the next
reconcile tick (the 30-second instance timer), because the derived
<instance>-pgbackrest-s3 Secret is compared by content hash — only instances whose
rendered s3.conf actually changed are rewritten. During that cutover window some
instances still hold the old values, so keep the old S3 key valid until the fleet has
converged; see Rotating the central
credentials for the convergence
check.