Managed PostgreSQL¶
Overview¶
The Nextcloud operator can automatically provision and manage PostgreSQL databases using Percona's PG Operator. This integration allows you to:
- Automatically provision HA PostgreSQL clusters when creating a Nextcloud instance
- Manage database lifecycle alongside Nextcloud
- Schedule recurring pgBackRest backups with retention — on by default, volume- or S3-backed
- Enable connection pooling via pgBouncer
Architecture¶
NextcloudInstance CR
↓
Operator detects spec.database.managed: true
↓
Creates PerconaPGCluster CR
↓
pg-operator provisions PostgreSQL
↓
Operator reads generated credentials secret
↓
Creates Nextcloud with database config
Prerequisites¶
- Percona PG Operator installed in the cluster
- Sufficient resources for PostgreSQL pods
- Storage class available
- (Optional) S3 credentials for backups
- Any egress-restricting
NetworkPolicyin the target namespace allows Patroni's API-server access — see Networking Requirements below
Networking Requirements¶
Percona's PG Operator manages HA and leader election via Patroni. On Kubernetes, Patroni's default DCS (Distributed Configuration Store) backend is the Kubernetes API server itself — every PostgreSQL pod must be able to reach the API server directly (not merely other in-cluster Services) to read and write leader-election state.
If your cluster or namespace enforces egress-restricting NetworkPolicy resources —
for example one delivered via a profile's or instance's helm.values.extraManifests
block — make sure the policy allows:
- TCP 6443 to the control-plane node(s) / API-server endpoints (the actual
API-server IPs, not just the in-cluster
kubernetesService) - TCP 443 to the API server's ClusterIP (
kubernetes.default.svc)
Without this, Patroni can never reach its DCS backend and the PostgreSQL cluster
never becomes healthy, no matter how long the operator's Database not ready yet
retry runs. See Troubleshooting → NetworkPolicy blocks Patroni's API-server
access
for the exact symptom and diagnostic commands.
The operator does not currently validate extraManifests content (it is an opaque
Helm-chart values passthrough — the operator has no NetworkPolicy-aware logic
anywhere today), so a missing egress rule surfaces only as a stuck/unhealthy
PostgreSQL cluster, not as an operator-side condition or event.
Install Percona PG Operator¶
helm repo add percona https://percona.github.io/percona-helm-charts/
helm repo update
helm install pgo percona/pg-operator \
--namespace pgo \
--create-namespace
# Verify installation
kubectl get pods -n pgo
Configuration¶
Managed PostgreSQL (Automatic Provisioning)¶
apiVersion: k8s.bnerd.com/v1alpha1
kind: NextcloudInstance
metadata:
name: my-nextcloud
spec:
profile: production
database:
managed: true
type: postgresql
postgres:
replicas: 3
version: "16"
resources:
requests: {cpu: "500m", memory: "1Gi"}
limits: {cpu: "2000m", memory: "4Gi"}
storage:
size: "20Gi"
storageClass: "fast-ssd"
backup:
# enabled defaults to true (see Backups below) -- shown explicitly here.
enabled: true
retention:
full: 4 # keep the last 4 full backups (default: 2)
s3:
bucket: my-nextcloud-db-backups
endpoint: s3.amazonaws.com
region: us-east-1
credentialsSecret: pgbackrest-s3-credentials
schedule:
full: "0 1 * * 0" # Weekly full backup
incremental: "0 */6 * * *" # Every 6 hours
Resources are reconciled, not create-only (0.23.5)¶
postgres.resources is written to PerconaPGCluster.spec.instances[0].resources at
create time and re-asserted on every reconcile since 0.23.5 (before that a later
edit changed the NextcloudInstance and nothing else; 0.23.4 shipped it broken and is withdrawn). The comparison is by Kubernetes
quantity value, so a converged cluster costs one read per reconcile and no write.
- A block naming only
requestskeeps the stock limits (2000m/4Gi), and vice versa. - Absent means don't touch. Omit
resourcesentirely and the create-time default (500m/1Girequests) stays; a cluster tuned by hand on the PerconaPGCluster is never reverted. - A change rolls the Postgres pod. On the default single-replica cluster that is a short database outage for the tenant; schedule fleet-wide changes accordingly.
External PostgreSQL (Bring Your Own)¶
apiVersion: k8s.bnerd.com/v1alpha1
kind: NextcloudInstance
metadata:
name: my-nextcloud
spec:
profile: production
database:
type: postgresql
credentialsSecret: my-db-credentials
managed: true always wins over a credentialsSecret (0.21.2)¶
database.managed: true and database.credentialsSecret describe two different
databases in intent — managed: true provisions and owns a real PerconaPGCluster;
credentialsSecret points at an external one. If a spec sets both (typically
left over from switching a previously-external instance to managed, or copy-pasted
from an external-DB example), the operator always uses its own just-provisioned
Percona connection info and ignores the credentialsSecret reference for connection
purposes — it never falls back to the foreign secret's host/user/password. This
is not a hard rejection: an existing spec with both fields set does not go Failed
at upgrade, it just gets a once-per-onset ConflictingDatabaseConfig Warning event
naming the ignored field (deduped via status.databaseConfigConflictWarned, so a
retried on_create doesn't re-fire it). See Troubleshooting → ConflictingDatabaseConfig
warning event
if you see this event and aren't sure which value is live.
Backups¶
What the birth replica-create backup is — and is NOT¶
Percona's PG Operator always takes one replica-create backup when a PerconaPGCluster
is first created, to seed the replica(s). This is a one-off snapshot, not a recurring
schedule — a fleet-wide finding (mid-2026) showed instances running for months with
their only backup still being that empty birth snapshot, because spec.database.postgres.backup
was never set. If you see exactly one replica-create backup and nothing since, scheduled
backups are not actually running for that instance — check enabled below and confirm the
schedule is applying (kubectl describe perconapgcluster <name>-pg, Events).
Backups are ON by default (0.21.3)¶
As of 0.21.3, database.postgres.backup.enabled defaults to true for every
managed PostgreSQL instance — an entirely absent backup key is no longer "disabled".
The safe baseline applied when nothing is configured:
- The standard cadence: weekly full + 6-hourly incremental. The exact minute (and, for
full, the hour and day-of-week) within that cadence is jittered per instance,
derived from a stable hash of the instance's cluster name — a fleet where every
instance used the literal same offset would start dozens of pgBackRest jobs in the same
minute every Sunday and again at every :00/:06/:12/:18 hour, a synchronized backup storm
against shared infrastructure. The same instance always gets the same offset (stable
across reconciles); different instances get different offsets. Set
schedule.full/schedule.incrementalexplicitly if you need a specific, predictable time — an explicit schedule always wins verbatim and is never jittered. - A 10Gi volume-backed repo (
repo1) — raised from the historical 1Gi default, which fills within days once a schedule is actually running - Retention: the last 2 full backups (
retention.full: 2)
Set enabled: false to explicitly opt out. This is a genuine per-instance/per-profile
decision now, not an omission.
Live reconciliation: unlike prior releases, toggling backup.enabled, schedule,
retention, or volumeSize on an already-running instance now takes effect on the
next reconcile tick — the instance's periodic 30s status timer, a
kubectl annotate ... k8s.bnerd.com/reconcile=... force-reconcile, or the next genuine
spec change — the operator patches the live PerconaPGCluster's
spec.backups.pgbackrest.repos[0] and retention global options in place. Prior to
0.21.3, backup config only ever applied at cluster creation; flipping it later was a
silent no-op. The reconciler also never shrinks an already-provisioned repo volume below
its current size, even when disabling backups or lowering volumeSize — Kubernetes
forbids shrinking a PVC below its provisioned capacity, so the larger size is always kept.
Retention¶
database:
managed: true
type: postgresql
postgres:
backup:
enabled: true
retention:
full: 4 # keep the last 4 full backups (pgBackRest repo1-retention-full)
differential: 8 # keep the last 8 differential backups (repo1-retention-diff)
retention.full defaults to 2 in code when backups are enabled and unset (no CRD
default). retention.differential is emitted only when set. retention.incremental is
accepted by the schema for forward compatibility but is not translated to a pgBackRest
option — pgBackRest has no repo-retention-incr setting; an incremental backup's expiry is
tied to its parent full/differential backup's own retention.
Volume sizing¶
database:
managed: true
type: postgresql
postgres:
backup:
enabled: true
volumeSize: "50Gi" # default: 10Gi when enabled, 1Gi when disabled
volumeSize only applies to the volume-backed repo (ignored once an s3 repo is
configured) and only when backups are enabled — a newly created disabled repo starts
at the cost-minimal 1Gi regardless of volumeSize. On an already-running instance, the
reconciler never shrinks a repo volume below its current live size, even when disabling
backups or setting a smaller volumeSize — Kubernetes rejects a PVC request below its
provisioned capacity, so a growth is permanent even if you later disable or reduce it.
Off-cluster backups to S3 (0.22.0)¶
Volume-backed backups are fine for smaller instances but do not survive PVC or node
loss. As of 0.22.0 the operator can add a second pgBackRest repository (repo2)
on S3 alongside the volume-backed repo1, for every managed PostgreSQL it looks
after — configured once on the operator, not per instance.
Where the bucket comes from¶
The operator resolves the backup target in this order, and stops at the first hit:
| # | Source | Use it for |
|---|---|---|
| 1 | The merged instance spec: database.postgres.backup.s3.bucket (including anything a profile merges in) |
one instance, or one profile, that needs its own bucket |
| 2 | The operator's own configuration: backup.dbDefaultS3.* in the operator Helm values |
the whole estate |
| 3 | Nothing resolves | volume-backed repo1 only — the pre-0.22.0 behaviour |
Enabling off-cluster backups estate-wide is therefore one change to the operator's Helm release plus one Secret, with no edit to any instance, profile, or CRD:
# values.yaml of the nextcloud-operator Helm release
backup:
dbDefaultS3:
bucket: nextcloud-db-backups
endpoint: s3.example.com
region: eu-central-1
secretRef: db-backup-s3 # S3 access keys
cipherSecretRef: db-backup-cipher # backup-encryption passphrase
# secretNamespace / cipherSecretNamespace default to the operator's namespace
# usePathStyle: true # MinIO and some on-prem gateways
The credentials Secret must live in secretNamespace (the operator's own namespace by
default) and carry the access keys under any one of these spellings:
apiVersion: v1
kind: Secret
metadata:
name: db-backup-s3
type: Opaque
stringData:
accessKey: "..." # or access-key, or AWS_ACCESS_KEY_ID
secretKey: "..." # or secret-key, or AWS_SECRET_ACCESS_KEY
Percona can only project a Secret that lives in the same namespace as the
PerconaPGCluster, so the operator renders this central Secret into a per-instance
<instance>-pgbackrest-s3 Secret (in pgBackRest's s3.conf format) in every instance
namespace. That derived Secret is content-hashed: a steady-state reconcile rewrites
nothing, and rotating the central Secret re-syncs every namespace on the next tick
(see Rotating the central credentials).
Instance-level Secret references are confined to the instance's namespace
An instance may name its own Secret instead
(backup.s3.credentialsFrom.secretRef.name, or credentialsSecret for a
ready-made s3.conf). Those references always resolve in that instance's own
namespace — neither has a namespace subfield.
That is a security boundary. The operator holds cluster-wide Secret read while an
instance spec is tenant-writable, so a spec-supplied namespace would let an
ordinary CR write steer the operator's read at any namespace — including this
one — and have the contents copied into a Secret the tenant can read. Only the
operator's own backup.dbDefaultS3.* values may point across namespaces.
Rotating the central credentials¶
Overwrite the Secret; every instance namespace re-syncs on its next reconcile tick
(the 30-second instance timer, so typically well under a minute, bounded by
TIMER_INSTANCE_STATUS_INTERVAL). The derived Secret is compared by content hash, so
only the instances whose rendered s3.conf actually changed are rewritten — no
unnecessary repo-host pod rolls.
During the cutover window some instances still hold the old credentials, so keep the old S3 key valid until every instance has re-synced; revoking it immediately fails in-flight backup jobs on instances that have not ticked yet. Confirm convergence with:
kubectl get secret -A -l app.kubernetes.io/component=pgbackrest-s3 \
-o jsonpath='{range .items[*]}{.metadata.namespace}{"\t"}{.metadata.annotations.k8s\.bnerd\.com/pgbackrest-s3-conf-hash}{"\n"}{end}' \
| sort | uniq -c -f1
One distinct hash across the fleet means the roll-out is complete.
Rotating the CIPHER key is different
Rotating cipherSecretRef does not re-encrypt anything. New backups use the new
key; existing backups still need the old one. Keep every retired key at least
as long as the backups it can open, and record which key covers which date range.
Per-instance paths (why this matters)¶
Percona renders the same repository path for every cluster it manages. Pointing a
whole fleet at one bucket without an explicit per-instance path would interleave every
instance's backups into a single prefix — which is not merely untidy, it is
unrestorable. The operator therefore always emits an explicit per-instance
repo2-path, defaulting to:
Namespace first, deliberately. An instance name is unique only within its
namespace; the namespace is unique cluster-wide. Keying on the name alone would let
two instances called nextcloud in different namespaces share one prefix in the shared
central bucket — the same unrestorable interleaving, one level up.
Override the template per instance with backup.s3.pathPrefix; the literals
{namespace} and {instance} are substituted. The configuration-files bundle uses the
sibling prefix /<namespace>/<instance-name>/backups/config/, so one bucket holds a
complete, per-instance-separated backup set.
Encryption¶
pgBackRest encrypts the repository with a key you hold. Set cipherSecretRef on the
operator (or backup.s3.encryption.keyFrom.secretRef on an instance) to a Secret with
a cipher-pass key:
The passphrase only ever exists inside the derived per-namespace Secret — it is
never written into the PerconaPGCluster CR, an event, or a log line. The cipher
type (aes-256-cbc) is what appears in the CR.
Store the key outside the cluster too
A backup encrypted with a key you have lost is noise. Put the passphrase in the password manager as well as in the cluster, and record where it lives in your restore runbook. If no key resolves, the operator emits no cipher configuration at all rather than a half-configured repository — and the configuration-files bundle does not run.
Retention, cadence and the first backup¶
repo2 gets its own retention, independent of repo1:
database:
managed: true
type: postgresql
postgres:
backup:
s3:
retention:
full: 4 # code default: 4 full backups on S3 (repo1 keeps 2)
differential: 8 # emitted only when set
Both repositories share the instance's schedule (weekly full + 6-hourly incremental,
jittered per instance). Because a scheduled full could otherwise be up to a week away —
and a repository with no full backup can restore nothing — the operator triggers one
manual full backup on repo2 as soon as it is first configured, records
status.backupRepo2Bootstrapped, and never repeats it. Watch for the
BackupRepo2Bootstrapped event.
alongside vs replace¶
backup:
s3:
bucket: nextcloud-db-backups
mode: alongside # default (operator code default, no CRD default)
alongside(default) keepsrepo1volume-backed and addsrepo2on S3. The local repository and its history survive, and WAL archiving goes to both.replaceis the pre-0.22.0 behaviour: the S3 target replacesrepo1in place. The volume repository and its history are orphaned and the WAL archive target flips in one step.
replace requires the instance's own bucket
Replace mode puts the S3 target on repo1, and the operator deliberately does
not emit a repo1-path — doing so would relocate the backup history of every
existing bring-your-own-bucket user. But Percona renders the same default repo1
path for every cluster it manages, so replace mode pointed at the shared
operator-central bucket would make the whole fleet write to one identical prefix.
The operator therefore refuses that combination: mode: replace with no
instance-provided s3.bucket resolves to no target at all (no repo change, no
credentials Secret, no config bundle), and logs why. Give the instance its own
bucket, or use the default alongside.
Watch pg_wal headroom
With two repositories, pgBackRest's archive-push writes each WAL segment to
both. If the S3 endpoint is unavailable, archiving can fall behind and
PostgreSQL retains WAL in pg_wal until it succeeds — on a busy instance that can
grow quickly, and a full pg_wal volume stops the database.
Verify your instances have pg_wal headroom before enabling this fleet-wide, and
watch the PVC during the first S3 outage. The safe magnitude has not yet been
measured on our own fleet; treat this as a known unknown rather than a solved
problem until it has been.
Migration notes for existing instances¶
0.21.3 (backups on by default). Deploying 0.21.3 required no spec change: every
existing managed-PostgreSQL instance that never set spec.database.postgres.backup
picked up schedules, the 10Gi volume, and retention.full: 2 on its next reconcile
tick. To preserve the old no-schedules behaviour for a specific instance, set
backup.enabled: false explicitly.
0.22.0 (off-cluster S3).
Action required before upgrading — only if you already set backup.s3
If an instance already sets database.postgres.backup.s3 today, set
s3.mode: replace on it to keep the current single-repository behaviour —
otherwise its backup history moves to a different object path. Details in point
2 below. An estate that has only ever used the volume-backed default is
unaffected, and needs no action.
Two things to know before upgrading:
- Nothing changes unless you configure a bucket. With
backup.dbDefaultS3.bucketunset and no instance-levelbackup.s3, the emitted configuration is byte-identical to 0.21.x. -
modenow defaults toalongside. If an instance already setsdatabase.postgres.backup.s3today, its S3 target currently replacesrepo1in place. On 0.22.0 that instance switches torepo1(volume) plusrepo2(S3) at a different object path, sorepo1is re-provisioned as a fresh volume repository and the existing S3 history stays where it is under the old path rather than being continued.To keep exactly today's behaviour for such an instance, set the mode explicitly before upgrading:
Check whether you are affected:
Disabling backups (backup.enabled: false) still wins over everything above: a
centrally-configured default bucket never resurrects backups on an instance that
deliberately opted out, and no credentials Secret is written into its namespace.
Examples¶
Minimal Managed Database¶
Uses defaults: 1 replica, 10Gi storage, PostgreSQL 16, and — as of 0.21.3 — scheduled backups (weekly full + 6-hourly incremental, 10Gi backup repo, retention.full=2).
Production with HA and Backups¶
database:
managed: true
type: postgresql
postgres:
replicas: 3
version: "16"
resources:
requests: {cpu: "1000m", memory: "2Gi"}
limits: {cpu: "4000m", memory: "8Gi"}
storage:
size: "100Gi"
storageClass: "ssd"
backup:
# enabled: true is the default (0.21.3) -- shown explicitly here.
enabled: true
retention:
full: 4
s3:
# Most estates set this ONCE on the operator (backup.dbDefaultS3.*)
# instead -- shown per instance here to illustrate the override.
bucket: prod-nc-db-backups
endpoint: s3.example.com
region: eu-central-1
# A ready-made Secret in THIS namespace, in pgBackRest s3.conf format.
# Drop it to have the operator derive one from the central Secret.
credentialsSecret: s3-backup-creds
# mode: alongside is the default -- repo1 (volume) is kept and this
# bucket becomes repo2 under /<instance>/backups/db.
retention:
full: 4
schedule:
full: "0 2 * * 0"
incremental: "0 */4 * * *"
With Connection Pooling (pgBouncer)¶
pgBouncer is off by default. Enable it explicitly for workloads with high
connection churn. The production built-in profile sets proxy.enabled: true.
database:
managed: true
type: postgresql
postgres:
replicas: 2
proxy:
enabled: true
replicas: 2
poolMode: transaction
poolSize: 25
See Redis Topology & pgBouncer for the full toggle runbook and connection-pool sizing guidance.
Status Tracking¶
The Nextcloud CR status includes database provisioning progress:
status:
phase: Creating
conditions:
- type: DatabaseReady
status: "False"
reason: Provisioning
message: "Waiting for PerconaPGCluster to be ready"
Once ready:
status:
phase: Ready
conditions:
- type: DatabaseReady
status: "True"
reason: Ready
message: "PostgreSQL cluster is ready"
database:
managed: true
clusterName: my-nextcloud-pg
host: my-nextcloud-pg-pgbouncer.default.svc.cluster.local
port: 5432
name: nextcloud
Cleanup Behavior¶
When you delete a Nextcloud instance with a managed database:
- The operator deletes the HelmRelease and Nextcloud secrets
- The PerconaPGCluster is kept by default (data safety)
To also delete the database:
Or manually:
Note
When the instance owns its namespace, namespace deletion removes the database regardless of
deleteOnCleanup. See Deletion & Cleanup for the full teardown order.
Best Practices¶
- Use S3 for production backups — backups are on by default (0.21.3), but the
default volume-backed repo is bounded by
volumeSize; switch to S3 for anything long-lived (see Off-cluster backups to S3) - Review retention for your compliance/cost needs — the code default
(
retention.full: 2) is a safe minimum, not a policy - Use at least 3 replicas for HA
- Provision adequate resources (PostgreSQL is memory-hungry)
- Use fast storage (SSD recommended)
- Test restores regularly
- Monitor database metrics
- Keep pg-operator updated
Troubleshooting¶
# Check PerconaPGCluster status
kubectl get perconapgcluster
kubectl describe perconapgcluster <name>-pg
# Check PostgreSQL pods
kubectl get pods -l postgres-operator.crunchydata.com/cluster=<name>-pg
# Check pg-operator logs
kubectl logs -n pgo -l app.kubernetes.io/name=percona-postgresql-operator
# Check database credentials secret
kubectl get secret <name>-pguser-nextcloud -o jsonpath='{.data}' | jq
# Check backup pods
kubectl get pods -l postgres-operator.crunchydata.com/pgbackrest=<name>-pg
# Trigger manual reconcile after database change
kubectl annotate nci <name> k8s.bnerd.com/reconcile=$(date +%s) --overwrite
See the Operations & Annotations guide for the full list of operational annotations.