Skip to content

Managed PostgreSQL

Overview

The Nextcloud operator can automatically provision and manage PostgreSQL databases using Percona's PG Operator. This integration allows you to:

  1. Automatically provision HA PostgreSQL clusters when creating a Nextcloud instance
  2. Manage database lifecycle alongside Nextcloud
  3. Schedule recurring pgBackRest backups with retention — on by default, volume- or S3-backed
  4. Enable connection pooling via pgBouncer

Architecture

NextcloudInstance CR
    ↓
Operator detects spec.database.managed: true
    ↓
Creates PerconaPGCluster CR
    ↓
pg-operator provisions PostgreSQL
    ↓
Operator reads generated credentials secret
    ↓
Creates Nextcloud with database config

Prerequisites

  • Percona PG Operator installed in the cluster
  • Sufficient resources for PostgreSQL pods
  • Storage class available
  • (Optional) S3 credentials for backups
  • Any egress-restricting NetworkPolicy in the target namespace allows Patroni's API-server access — see Networking Requirements below

Networking Requirements

Percona's PG Operator manages HA and leader election via Patroni. On Kubernetes, Patroni's default DCS (Distributed Configuration Store) backend is the Kubernetes API server itself — every PostgreSQL pod must be able to reach the API server directly (not merely other in-cluster Services) to read and write leader-election state.

If your cluster or namespace enforces egress-restricting NetworkPolicy resources — for example one delivered via a profile's or instance's helm.values.extraManifests block — make sure the policy allows:

  • TCP 6443 to the control-plane node(s) / API-server endpoints (the actual API-server IPs, not just the in-cluster kubernetes Service)
  • TCP 443 to the API server's ClusterIP (kubernetes.default.svc)

Without this, Patroni can never reach its DCS backend and the PostgreSQL cluster never becomes healthy, no matter how long the operator's Database not ready yet retry runs. See Troubleshooting → NetworkPolicy blocks Patroni's API-server access for the exact symptom and diagnostic commands.

The operator does not currently validate extraManifests content (it is an opaque Helm-chart values passthrough — the operator has no NetworkPolicy-aware logic anywhere today), so a missing egress rule surfaces only as a stuck/unhealthy PostgreSQL cluster, not as an operator-side condition or event.

Install Percona PG Operator

helm repo add percona https://percona.github.io/percona-helm-charts/
helm repo update

helm install pgo percona/pg-operator \
  --namespace pgo \
  --create-namespace

# Verify installation
kubectl get pods -n pgo

Configuration

Managed PostgreSQL (Automatic Provisioning)

apiVersion: k8s.bnerd.com/v1alpha1
kind: NextcloudInstance
metadata:
  name: my-nextcloud
spec:
  profile: production
  database:
    managed: true
    type: postgresql
    postgres:
      replicas: 3
      version: "16"
      resources:
        requests: {cpu: "500m", memory: "1Gi"}
        limits: {cpu: "2000m", memory: "4Gi"}
      storage:
        size: "20Gi"
        storageClass: "fast-ssd"
      backup:
        # enabled defaults to true (see Backups below) -- shown explicitly here.
        enabled: true
        retention:
          full: 4                    # keep the last 4 full backups (default: 2)
        s3:
          bucket: my-nextcloud-db-backups
          endpoint: s3.amazonaws.com
          region: us-east-1
          credentialsSecret: pgbackrest-s3-credentials
        schedule:
          full: "0 1 * * 0"           # Weekly full backup
          incremental: "0 */6 * * *"   # Every 6 hours

Resources are reconciled, not create-only (0.23.5)

postgres.resources is written to PerconaPGCluster.spec.instances[0].resources at create time and re-asserted on every reconcile since 0.23.5 (before that a later edit changed the NextcloudInstance and nothing else; 0.23.4 shipped it broken and is withdrawn). The comparison is by Kubernetes quantity value, so a converged cluster costs one read per reconcile and no write.

  • A block naming only requests keeps the stock limits (2000m / 4Gi), and vice versa.
  • Absent means don't touch. Omit resources entirely and the create-time default (500m / 1Gi requests) stays; a cluster tuned by hand on the PerconaPGCluster is never reverted.
  • A change rolls the Postgres pod. On the default single-replica cluster that is a short database outage for the tenant; schedule fleet-wide changes accordingly.

External PostgreSQL (Bring Your Own)

apiVersion: k8s.bnerd.com/v1alpha1
kind: NextcloudInstance
metadata:
  name: my-nextcloud
spec:
  profile: production
  database:
    type: postgresql
    credentialsSecret: my-db-credentials

managed: true always wins over a credentialsSecret (0.21.2)

database.managed: true and database.credentialsSecret describe two different databases in intent — managed: true provisions and owns a real PerconaPGCluster; credentialsSecret points at an external one. If a spec sets both (typically left over from switching a previously-external instance to managed, or copy-pasted from an external-DB example), the operator always uses its own just-provisioned Percona connection info and ignores the credentialsSecret reference for connection purposes — it never falls back to the foreign secret's host/user/password. This is not a hard rejection: an existing spec with both fields set does not go Failed at upgrade, it just gets a once-per-onset ConflictingDatabaseConfig Warning event naming the ignored field (deduped via status.databaseConfigConflictWarned, so a retried on_create doesn't re-fire it). See Troubleshooting → ConflictingDatabaseConfig warning event if you see this event and aren't sure which value is live.

Backups

What the birth replica-create backup is — and is NOT

Percona's PG Operator always takes one replica-create backup when a PerconaPGCluster is first created, to seed the replica(s). This is a one-off snapshot, not a recurring schedule — a fleet-wide finding (mid-2026) showed instances running for months with their only backup still being that empty birth snapshot, because spec.database.postgres.backup was never set. If you see exactly one replica-create backup and nothing since, scheduled backups are not actually running for that instance — check enabled below and confirm the schedule is applying (kubectl describe perconapgcluster <name>-pg, Events).

Backups are ON by default (0.21.3)

As of 0.21.3, database.postgres.backup.enabled defaults to true for every managed PostgreSQL instance — an entirely absent backup key is no longer "disabled". The safe baseline applied when nothing is configured:

  • The standard cadence: weekly full + 6-hourly incremental. The exact minute (and, for full, the hour and day-of-week) within that cadence is jittered per instance, derived from a stable hash of the instance's cluster name — a fleet where every instance used the literal same offset would start dozens of pgBackRest jobs in the same minute every Sunday and again at every :00/:06/:12/:18 hour, a synchronized backup storm against shared infrastructure. The same instance always gets the same offset (stable across reconciles); different instances get different offsets. Set schedule.full/ schedule.incremental explicitly if you need a specific, predictable time — an explicit schedule always wins verbatim and is never jittered.
  • A 10Gi volume-backed repo (repo1) — raised from the historical 1Gi default, which fills within days once a schedule is actually running
  • Retention: the last 2 full backups (retention.full: 2)

Set enabled: false to explicitly opt out. This is a genuine per-instance/per-profile decision now, not an omission.

Live reconciliation: unlike prior releases, toggling backup.enabled, schedule, retention, or volumeSize on an already-running instance now takes effect on the next reconcile tick — the instance's periodic 30s status timer, a kubectl annotate ... k8s.bnerd.com/reconcile=... force-reconcile, or the next genuine spec change — the operator patches the live PerconaPGCluster's spec.backups.pgbackrest.repos[0] and retention global options in place. Prior to 0.21.3, backup config only ever applied at cluster creation; flipping it later was a silent no-op. The reconciler also never shrinks an already-provisioned repo volume below its current size, even when disabling backups or lowering volumeSize — Kubernetes forbids shrinking a PVC below its provisioned capacity, so the larger size is always kept.

Retention

database:
  managed: true
  type: postgresql
  postgres:
    backup:
      enabled: true
      retention:
        full: 4           # keep the last 4 full backups (pgBackRest repo1-retention-full)
        differential: 8   # keep the last 8 differential backups (repo1-retention-diff)

retention.full defaults to 2 in code when backups are enabled and unset (no CRD default). retention.differential is emitted only when set. retention.incremental is accepted by the schema for forward compatibility but is not translated to a pgBackRest option — pgBackRest has no repo-retention-incr setting; an incremental backup's expiry is tied to its parent full/differential backup's own retention.

Volume sizing

database:
  managed: true
  type: postgresql
  postgres:
    backup:
      enabled: true
      volumeSize: "50Gi"   # default: 10Gi when enabled, 1Gi when disabled

volumeSize only applies to the volume-backed repo (ignored once an s3 repo is configured) and only when backups are enabled — a newly created disabled repo starts at the cost-minimal 1Gi regardless of volumeSize. On an already-running instance, the reconciler never shrinks a repo volume below its current live size, even when disabling backups or setting a smaller volumeSize — Kubernetes rejects a PVC request below its provisioned capacity, so a growth is permanent even if you later disable or reduce it.

Off-cluster backups to S3 (0.22.0)

Volume-backed backups are fine for smaller instances but do not survive PVC or node loss. As of 0.22.0 the operator can add a second pgBackRest repository (repo2) on S3 alongside the volume-backed repo1, for every managed PostgreSQL it looks after — configured once on the operator, not per instance.

Where the bucket comes from

The operator resolves the backup target in this order, and stops at the first hit:

# Source Use it for
1 The merged instance spec: database.postgres.backup.s3.bucket (including anything a profile merges in) one instance, or one profile, that needs its own bucket
2 The operator's own configuration: backup.dbDefaultS3.* in the operator Helm values the whole estate
3 Nothing resolves volume-backed repo1 only — the pre-0.22.0 behaviour

Enabling off-cluster backups estate-wide is therefore one change to the operator's Helm release plus one Secret, with no edit to any instance, profile, or CRD:

# values.yaml of the nextcloud-operator Helm release
backup:
  dbDefaultS3:
    bucket: nextcloud-db-backups
    endpoint: s3.example.com
    region: eu-central-1
    secretRef: db-backup-s3                # S3 access keys
    cipherSecretRef: db-backup-cipher      # backup-encryption passphrase
    # secretNamespace / cipherSecretNamespace default to the operator's namespace
    # usePathStyle: true                   # MinIO and some on-prem gateways

The credentials Secret must live in secretNamespace (the operator's own namespace by default) and carry the access keys under any one of these spellings:

apiVersion: v1
kind: Secret
metadata:
  name: db-backup-s3
type: Opaque
stringData:
  accessKey: "..."        # or access-key, or AWS_ACCESS_KEY_ID
  secretKey: "..."        # or secret-key, or AWS_SECRET_ACCESS_KEY

Percona can only project a Secret that lives in the same namespace as the PerconaPGCluster, so the operator renders this central Secret into a per-instance <instance>-pgbackrest-s3 Secret (in pgBackRest's s3.conf format) in every instance namespace. That derived Secret is content-hashed: a steady-state reconcile rewrites nothing, and rotating the central Secret re-syncs every namespace on the next tick (see Rotating the central credentials).

Instance-level Secret references are confined to the instance's namespace

An instance may name its own Secret instead (backup.s3.credentialsFrom.secretRef.name, or credentialsSecret for a ready-made s3.conf). Those references always resolve in that instance's own namespace — neither has a namespace subfield.

That is a security boundary. The operator holds cluster-wide Secret read while an instance spec is tenant-writable, so a spec-supplied namespace would let an ordinary CR write steer the operator's read at any namespace — including this one — and have the contents copied into a Secret the tenant can read. Only the operator's own backup.dbDefaultS3.* values may point across namespaces.

Rotating the central credentials

Overwrite the Secret; every instance namespace re-syncs on its next reconcile tick (the 30-second instance timer, so typically well under a minute, bounded by TIMER_INSTANCE_STATUS_INTERVAL). The derived Secret is compared by content hash, so only the instances whose rendered s3.conf actually changed are rewritten — no unnecessary repo-host pod rolls.

During the cutover window some instances still hold the old credentials, so keep the old S3 key valid until every instance has re-synced; revoking it immediately fails in-flight backup jobs on instances that have not ticked yet. Confirm convergence with:

kubectl get secret -A -l app.kubernetes.io/component=pgbackrest-s3 \
  -o jsonpath='{range .items[*]}{.metadata.namespace}{"\t"}{.metadata.annotations.k8s\.bnerd\.com/pgbackrest-s3-conf-hash}{"\n"}{end}' \
  | sort | uniq -c -f1

One distinct hash across the fleet means the roll-out is complete.

Rotating the CIPHER key is different

Rotating cipherSecretRef does not re-encrypt anything. New backups use the new key; existing backups still need the old one. Keep every retired key at least as long as the backups it can open, and record which key covers which date range.

Per-instance paths (why this matters)

Percona renders the same repository path for every cluster it manages. Pointing a whole fleet at one bucket without an explicit per-instance path would interleave every instance's backups into a single prefix — which is not merely untidy, it is unrestorable. The operator therefore always emits an explicit per-instance repo2-path, defaulting to:

/<namespace>/<instance-name>/backups/db

Namespace first, deliberately. An instance name is unique only within its namespace; the namespace is unique cluster-wide. Keying on the name alone would let two instances called nextcloud in different namespaces share one prefix in the shared central bucket — the same unrestorable interleaving, one level up.

Override the template per instance with backup.s3.pathPrefix; the literals {namespace} and {instance} are substituted. The configuration-files bundle uses the sibling prefix /<namespace>/<instance-name>/backups/config/, so one bucket holds a complete, per-instance-separated backup set.

Encryption

pgBackRest encrypts the repository with a key you hold. Set cipherSecretRef on the operator (or backup.s3.encryption.keyFrom.secretRef on an instance) to a Secret with a cipher-pass key:

openssl rand -base64 48

The passphrase only ever exists inside the derived per-namespace Secret — it is never written into the PerconaPGCluster CR, an event, or a log line. The cipher type (aes-256-cbc) is what appears in the CR.

Store the key outside the cluster too

A backup encrypted with a key you have lost is noise. Put the passphrase in the password manager as well as in the cluster, and record where it lives in your restore runbook. If no key resolves, the operator emits no cipher configuration at all rather than a half-configured repository — and the configuration-files bundle does not run.

Retention, cadence and the first backup

repo2 gets its own retention, independent of repo1:

database:
  managed: true
  type: postgresql
  postgres:
    backup:
      s3:
        retention:
          full: 4          # code default: 4 full backups on S3 (repo1 keeps 2)
          differential: 8  # emitted only when set

Both repositories share the instance's schedule (weekly full + 6-hourly incremental, jittered per instance). Because a scheduled full could otherwise be up to a week away — and a repository with no full backup can restore nothing — the operator triggers one manual full backup on repo2 as soon as it is first configured, records status.backupRepo2Bootstrapped, and never repeats it. Watch for the BackupRepo2Bootstrapped event.

alongside vs replace

backup:
  s3:
    bucket: nextcloud-db-backups
    mode: alongside     # default (operator code default, no CRD default)
  • alongside (default) keeps repo1 volume-backed and adds repo2 on S3. The local repository and its history survive, and WAL archiving goes to both.
  • replace is the pre-0.22.0 behaviour: the S3 target replaces repo1 in place. The volume repository and its history are orphaned and the WAL archive target flips in one step.

replace requires the instance's own bucket

Replace mode puts the S3 target on repo1, and the operator deliberately does not emit a repo1-path — doing so would relocate the backup history of every existing bring-your-own-bucket user. But Percona renders the same default repo1 path for every cluster it manages, so replace mode pointed at the shared operator-central bucket would make the whole fleet write to one identical prefix.

The operator therefore refuses that combination: mode: replace with no instance-provided s3.bucket resolves to no target at all (no repo change, no credentials Secret, no config bundle), and logs why. Give the instance its own bucket, or use the default alongside.

Watch pg_wal headroom

With two repositories, pgBackRest's archive-push writes each WAL segment to both. If the S3 endpoint is unavailable, archiving can fall behind and PostgreSQL retains WAL in pg_wal until it succeeds — on a busy instance that can grow quickly, and a full pg_wal volume stops the database.

Verify your instances have pg_wal headroom before enabling this fleet-wide, and watch the PVC during the first S3 outage. The safe magnitude has not yet been measured on our own fleet; treat this as a known unknown rather than a solved problem until it has been.

Migration notes for existing instances

0.21.3 (backups on by default). Deploying 0.21.3 required no spec change: every existing managed-PostgreSQL instance that never set spec.database.postgres.backup picked up schedules, the 10Gi volume, and retention.full: 2 on its next reconcile tick. To preserve the old no-schedules behaviour for a specific instance, set backup.enabled: false explicitly.

0.22.0 (off-cluster S3).

Action required before upgrading — only if you already set backup.s3

If an instance already sets database.postgres.backup.s3 today, set s3.mode: replace on it to keep the current single-repository behaviour — otherwise its backup history moves to a different object path. Details in point 2 below. An estate that has only ever used the volume-backed default is unaffected, and needs no action.

Two things to know before upgrading:

  1. Nothing changes unless you configure a bucket. With backup.dbDefaultS3.bucket unset and no instance-level backup.s3, the emitted configuration is byte-identical to 0.21.x.
  2. mode now defaults to alongside. If an instance already sets database.postgres.backup.s3 today, its S3 target currently replaces repo1 in place. On 0.22.0 that instance switches to repo1 (volume) plus repo2 (S3) at a different object path, so repo1 is re-provisioned as a fresh volume repository and the existing S3 history stays where it is under the old path rather than being continued.

    To keep exactly today's behaviour for such an instance, set the mode explicitly before upgrading:

    database:
      postgres:
        backup:
          s3:
            bucket: my-existing-bucket
            mode: replace
    

    Check whether you are affected:

    kubectl get nextcloudinstances -A \
      -o jsonpath='{range .items[?(@.spec.database.postgres.backup.s3.bucket)]}{.metadata.namespace}/{.metadata.name}{"\n"}{end}'
    

Disabling backups (backup.enabled: false) still wins over everything above: a centrally-configured default bucket never resurrects backups on an instance that deliberately opted out, and no credentials Secret is written into its namespace.

Examples

Minimal Managed Database

database:
  managed: true
  type: postgresql

Uses defaults: 1 replica, 10Gi storage, PostgreSQL 16, and — as of 0.21.3 — scheduled backups (weekly full + 6-hourly incremental, 10Gi backup repo, retention.full=2).

Production with HA and Backups

database:
  managed: true
  type: postgresql
  postgres:
    replicas: 3
    version: "16"
    resources:
      requests: {cpu: "1000m", memory: "2Gi"}
      limits: {cpu: "4000m", memory: "8Gi"}
    storage:
      size: "100Gi"
      storageClass: "ssd"
    backup:
      # enabled: true is the default (0.21.3) -- shown explicitly here.
      enabled: true
      retention:
        full: 4
      s3:
        # Most estates set this ONCE on the operator (backup.dbDefaultS3.*)
        # instead -- shown per instance here to illustrate the override.
        bucket: prod-nc-db-backups
        endpoint: s3.example.com
        region: eu-central-1
        # A ready-made Secret in THIS namespace, in pgBackRest s3.conf format.
        # Drop it to have the operator derive one from the central Secret.
        credentialsSecret: s3-backup-creds
        # mode: alongside is the default -- repo1 (volume) is kept and this
        # bucket becomes repo2 under /<instance>/backups/db.
        retention:
          full: 4
      schedule:
        full: "0 2 * * 0"
        incremental: "0 */4 * * *"

With Connection Pooling (pgBouncer)

pgBouncer is off by default. Enable it explicitly for workloads with high connection churn. The production built-in profile sets proxy.enabled: true.

database:
  managed: true
  type: postgresql
  postgres:
    replicas: 2
    proxy:
      enabled: true
      replicas: 2
      poolMode: transaction
      poolSize: 25

See Redis Topology & pgBouncer for the full toggle runbook and connection-pool sizing guidance.

Status Tracking

The Nextcloud CR status includes database provisioning progress:

status:
  phase: Creating
  conditions:
    - type: DatabaseReady
      status: "False"
      reason: Provisioning
      message: "Waiting for PerconaPGCluster to be ready"

Once ready:

status:
  phase: Ready
  conditions:
    - type: DatabaseReady
      status: "True"
      reason: Ready
      message: "PostgreSQL cluster is ready"
  database:
    managed: true
    clusterName: my-nextcloud-pg
    host: my-nextcloud-pg-pgbouncer.default.svc.cluster.local
    port: 5432
    name: nextcloud

Cleanup Behavior

When you delete a Nextcloud instance with a managed database:

  • The operator deletes the HelmRelease and Nextcloud secrets
  • The PerconaPGCluster is kept by default (data safety)

To also delete the database:

spec:
  database:
    managed: true
    deleteOnCleanup: true  # Database will be deleted with the instance

Or manually:

kubectl delete perconapgcluster my-nextcloud-pg

Note

When the instance owns its namespace, namespace deletion removes the database regardless of deleteOnCleanup. See Deletion & Cleanup for the full teardown order.

Best Practices

  1. Use S3 for production backups — backups are on by default (0.21.3), but the default volume-backed repo is bounded by volumeSize; switch to S3 for anything long-lived (see Off-cluster backups to S3)
  2. Review retention for your compliance/cost needs — the code default (retention.full: 2) is a safe minimum, not a policy
  3. Use at least 3 replicas for HA
  4. Provision adequate resources (PostgreSQL is memory-hungry)
  5. Use fast storage (SSD recommended)
  6. Test restores regularly
  7. Monitor database metrics
  8. Keep pg-operator updated

Troubleshooting

# Check PerconaPGCluster status
kubectl get perconapgcluster
kubectl describe perconapgcluster <name>-pg

# Check PostgreSQL pods
kubectl get pods -l postgres-operator.crunchydata.com/cluster=<name>-pg

# Check pg-operator logs
kubectl logs -n pgo -l app.kubernetes.io/name=percona-postgresql-operator

# Check database credentials secret
kubectl get secret <name>-pguser-nextcloud -o jsonpath='{.data}' | jq

# Check backup pods
kubectl get pods -l postgres-operator.crunchydata.com/pgbackrest=<name>-pg

# Trigger manual reconcile after database change
kubectl annotate nci <name> k8s.bnerd.com/reconcile=$(date +%s) --overwrite

See the Operations & Annotations guide for the full list of operational annotations.