Skip to content

Backup & Restore

The operator backs up a Nextcloud instance in two halves, both written to the same S3 bucket under the same per-instance prefix:

Half What it holds Mechanism Guide
Database The PostgreSQL cluster, continuously pgBackRest repo2 (S3), managed by Percona Managed PostgreSQL
Configuration config/, encryption keys, declarative state, operator-owned Secrets, custom apps, themes Encrypted bundle uploaded by the operator This page

You need both halves

A database dump restored into a fresh instance is unusable without config/config.php's instanceid, secret and passwordsalt. Restore the configuration bundle first, then the database.

Neither half covers the tenant's file data in the S3 data bucket, or local data/ volumes. Those are separate concerns — see What is not backed up.

<bucket>/
└── <instance-name>/
    └── backups/
        ├── db/              ← pgBackRest repo2 (Percona writes this)
        └── config/
            ├── 20260825T031500Z.tar.enc
            └── 20260826T031500Z.tar.enc

The configuration bundle

What is in it

Priority-ordered — the first two are the ones no restore works without:

# Item Why it matters
1 config/ — the whole directory, not just config.php (so every *.config.php drop-in comes too) Holds instanceid, secret and passwordsalt. Without them there is no restore, only a new instance.
2 data/files_encryption/, when the server-side encryption app is active Encrypted files are unrecoverable noise without the keys.
3 Declarative state — the Nextcloud CR, the NextcloudInstance spec, the applied profile, and the operator/chart/Nextcloud versions Rebuild the instance shell on a clean cluster; input for restore tooling.
4 Operator-owned Secrets — admin, mail, the S3 data-bucket credentials, the OIDC client secret Reattach primary storage and accounts without credential archaeology.
5 custom_apps/ Store apps are re-installable, but version drift is real and genuinely custom apps are in no store.
6 themes/ Custom theming is not reproducible from the spec.
7 manifest.json — inventory, per-item SHA-256 checksums, versions, warnings How a restore drill verifies what it actually got.

Inline credentials in the captured spec (item 3) are replaced with a redaction marker: the real values live in item 4, exactly once.

What is not backed up

Not included Where it belongs instead
The tenant's files in the S3 data bucket Bucket versioning / replication, or the restic backups.data path
Local data/ volumes (other than the encryption keys) Volume snapshots
data/appdata* previews and caches Regenerated on demand
Redis contents Cache — nothing to restore
The HelmRelease and rendered values Regenerated from the spec
Database dumps pgBackRest owns the database

When it runs

  • Daily, at a per-instance hour derived from a stable hash of the instance name, so a whole estate does not upload to one bucket in the same minute.
  • Once immediately after pool assignment. A just-assigned instance is the one most likely to be restored soon, and its pre-assignment bundle describes a spare rather than a tenant.

It runs only when both a bucket and an encryption key resolve — see Configuration. No key means no bundle, never an unencrypted one: this artifact carries the config triad, the encryption keys and the operator-owned Secrets.

Configuration

The bucket and key resolve exactly as for the database backups: the merged instance spec first, then the operator's own backup.dbDefaultS3.* values. Configure the operator once and every instance is covered.

Per-instance knobs:

spec:
  backups:
    config:
      enabled: true    # code default: on, whenever a bucket and a key resolve
      retention: 14    # code default: keep 14 bundles, prune older after each upload

Size guard

The bundle streams out of the running pod in bounded chunks and is capped at roughly 200 MiB. Items 1 and 2 (config/, the encryption keys) and items 5 and 6 (custom_apps/, themes/) are streamed separately, so an oversized custom_apps/:

  • drops items 5 and 6 from that bundle,
  • records the reason in manifest.json under warnings, and repeats it on the ConfigBackupCompleted event,
  • still uploads a complete, restorable bundle of everything critical.

An oversized or missing config/ fails the run instead: uploading a bundle that cannot restore, while the freshness metric reports success, would be worse than no bundle at all.

Encryption format

The bundle is an OpenSSL enc container — Salted__ + an 8-byte random salt, key and IV derived with PBKDF2-HMAC-SHA256 (10 000 iterations), AES-256-CBC with PKCS#7 padding. That is deliberately not a bespoke format: a restore happens on someone's worst day, possibly without this operator and without Python, so decryption has to be possible with stock tooling.

enc containers are unauthenticated by design, so integrity is verified against the per-item SHA-256 checksums in manifest.json, not by the container itself.

Security model

Worth stating plainly, because both properties are accepted design, not oversights — if either is unacceptable for a given tenant, give that tenant its own bucket via the per-instance override.

Tenant separation inside a shared bucket is a convention, not enforcement

Every instance's backups live under its own key prefix — /<namespace>/<instance>/backups/db/ and /<namespace>/<instance>/backups/config/ — and that prefix is what keeps one instance's data distinguishable from another's. But with the operator-central default, all instances share one bucket and one S3 credential. Nothing at the object-store layer stops that credential from reading or writing another instance's prefix; separation rests on the operator always constructing the right prefix, not on IAM.

What that does and does not buy you:

  • Tenants themselves have no access. The bucket credential lives in the operator's namespace and is rendered into each instance namespace only in pgBackRest's s3.conf, for the database repo host. A Nextcloud tenant never receives it.
  • A compromise of the operator, or of the shared credential, exposes every instance's backups in that bucket — including the archive of instances unrelated to the breach.
  • A prefix bug is a cross-tenant bug. This is why the operator emits an explicit per-instance repo2-path (Percona would otherwise use one fleet-wide path) and why bundle pruning refuses to act on any key outside the instance's own prefix.

Give a tenant a genuinely isolated bucket — with its own IAM credential scoped to it — by setting database.postgres.backup.s3.bucket on that instance. It overrides the central default for both halves.

Secret-reference confinement

A Secret named by an instance spec always resolves in that instance's own namespace. credentialsFrom.secretRef and encryption.keyFrom.secretRef have no namespace subfield, in the CRD schema or in the operator.

The operator holds cluster-wide read on Secrets while an instance spec is tenant-writable, so honouring a spec-supplied namespace would be a confused deputy: an ordinary CR write would make the operator read any Secret in any namespace — the db-backup-cipher above included — and materialize its contents into a Secret in the tenant's own namespace, where the tenant can read it. The only cross-namespace source is the operator's own Helm values.

Restoring

1. Fetch and decrypt a bundle

INSTANCE=nc-happy-sun-a1b2c3
BUCKET=nextcloud-db-backups

# Newest bundle for this instance
aws s3 ls "s3://$BUCKET/$INSTANCE/backups/config/" --endpoint-url https://s3.example.com

aws s3 cp "s3://$BUCKET/$INSTANCE/backups/config/20260825T031500Z.tar.enc" . \
  --endpoint-url https://s3.example.com

# The passphrase is the SAME key as the pgBackRest cipher-pass for this environment.
kubectl -n nextcloud-operator-system get secret db-backup-cipher \
  -o jsonpath='{.data.cipher-pass}' | base64 -d > cipher.key
chmod 600 cipher.key

openssl enc -d -aes-256-cbc -pbkdf2 -pass file:cipher.key \
  -in 20260825T031500Z.tar.enc -out bundle.tar

mkdir bundle && tar -xf bundle.tar -C bundle && ls bundle

2. Verify it before trusting it

cd bundle
python3 - <<'EOF'
import hashlib, json
manifest = json.load(open("manifest.json"))
print(manifest["instance"], manifest["timestamp"], manifest["versions"])
for item in manifest["items"]:
    actual = hashlib.sha256(open(item["name"], "rb").read()).hexdigest()
    print(("OK  " if actual == item["sha256"] else "FAIL"), item["name"])
for warning in manifest["warnings"]:
    print("WARNING:", warning)
EOF

Any FAIL, or an unexpected warnings entry, means stop and use another bundle.

3. Rebuild the instance shell

python3 -c "import json;print(json.load(open('bundle/declarative-state.json'))['nextcloudInstance'])"

Re-apply the captured NextcloudInstance (and Nextcloud) spec. Credentials in that copy are redacted — restore them from bundle/secrets.json, which holds the real admin, mail, S3 data-bucket and OIDC values.

Handle the extracted bundle like a live credential

secrets.json and files-critical.tar.gz contain plaintext credentials and the passwordsalt/secret triad. Extract on an encrypted volume, and delete the working directory when you are done.

4. Restore the configuration files

kubectl -n <namespace> cp bundle/files-critical.tar.gz <pod>:/tmp/
kubectl -n <namespace> exec <pod> -- tar -xzf /tmp/files-critical.tar.gz -C /var/www/html
kubectl -n <namespace> exec <pod> -- rm /tmp/files-critical.tar.gz

files-optional.tar.gz (custom apps, themes) restores the same way, if present.

5. Restore the database

The database is restored with pgBackRest through Percona's own restore path, from repo2. Follow the Percona PG Operator restore procedure for the cluster <instance>-pg, selecting repo2 and the backup set nearest the bundle's timestamp. The repository is encrypted with the same passphrase as above.

6. Verify

kubectl -n <namespace> exec <pod> -- php occ status
kubectl -n <namespace> exec <pod> -- php occ user:list

Write the runbook from a drill, not from this page

Restore steps that have never been executed are a hypothesis. Run a restore drill against a scratch instance, record the real timings, and keep the RTO number that comes out of it — that is what an RPO/RTO commitment can be based on.

Monitoring

Metric Meaning
nextcloud_operator_config_backup_timestamp_seconds{namespace,instance} Unix timestamp of the last successful bundle upload
pgbackrest_last_backup_completion_timestamp_seconds{namespace,instance,repo,type} Last successful pgBackRest backup, per repository and type

Once S3 is the durable copy, alert on repo="repo2":

# Config bundle older than two days
time() - nextcloud_operator_config_backup_timestamp_seconds > 172800

# No full database backup on S3 in the last 8 days
time() - pgbackrest_last_backup_completion_timestamp_seconds{repo="repo2",type="full"} > 691200

Both series are removed when an instance is deleted, including when a stuck deletion is force-completed by removing the finalizer — so a firing alert always refers to a resource that still exists.

Events

Event Meaning
BackupScheduleApplied The effective database backup configuration changed and was applied
BackupRepo2Bootstrapped The one-time initial full backup on the S3 repository was triggered
ConfigBackupCompleted A bundle was uploaded; carries the object key, size and any warnings
ConfigBackupFailed A bundle run failed; the reason is redacted of credentials and deduplicated

Status fields

kubectl -n <namespace> get nci <instance> -o jsonpath='{.status.configBackup}' | jq
Field Meaning
configBackup.lastSuccess Timestamp of the last successful upload
configBackup.lastKey Object key of that bundle
configBackup.lastFailure / lastFailureReason Last failure and its redacted reason
configBackup.assignedRunDone The post-assignment bundle has been taken
backupRepo2Bootstrapped The initial S3 full backup was triggered

A failure never clears lastSuccess or lastKey: that bundle is still in the bucket, and blanking the fields would make one bad day look like "no backups have ever run".

Key custody

Both halves are encrypted with the same passphrase per environment, held in a Secret in the operator's namespace.

  • Generate with openssl rand -base64 48.
  • Store it in the password manager as well as in the cluster. A backup encrypted with a key you have lost is noise.
  • Record in your own runbook where it lives.
  • Rotation is a known limitation of a shared key: rotating it makes existing backups readable only with the old key. Keep retired keys for at least as long as the backups they can open, and record which key covers which date range.

Rotating either central Secret re-syncs every instance namespace on the next reconcile tick (the 30-second instance timer), because the derived <instance>-pgbackrest-s3 Secret is compared by content hash — only instances whose rendered s3.conf actually changed are rewritten. During that cutover window some instances still hold the old values, so keep the old S3 key valid until the fleet has converged; see Rotating the central credentials for the convergence check.