Skip to content

Standard Upgrade Procedure

A single, top-to-bottom checklist for upgrading a Nextcloud instance safely. The Upgrades, Operations & Annotations, Running occ Commands, and Version Management guides remain the canonical reference for why each piece behaves the way it does — this page assembles them into the procedure you actually run, in order, with nothing left to piece together yourself.

0. Before any upgrade

Verify backup freshness

  • Managed database (spec.database.managed: true): check the pgbackrest_last_backup_completion_timestamp_seconds{namespace,instance,repo,type} Prometheus gauge, or read the underlying PerconaPGBackup CRs directly if you don't have the metric pipeline handy:
# Via the metric (per repo/type, Succeeded backups only)
curl -s http://<operator-pod>:9090/metrics | grep pgbackrest_last_backup

# Via the CRs directly
kubectl get perconapgbackups.pgv2.percona.com -n $NS -l postgres-operator.crunchydata.com/cluster=<name>-pg

Absence of the metric series is itself the signal that no Succeeded backup has ever completed for that instance — there is no zero/placeholder value. See Monitoring → Backup freshness (pgBackRest) for the full source/cadence detail and a ready-to-use PgBackRestBackupStale alert.

As of 0.21.3, scheduled backups default to ON for managed PostgreSQL (see Managed PostgreSQL → Backups) — an instance created or reconciled since then should show a real, recurring Succeeded backup here, not just the one-off replica-create snapshot taken at cluster birth. Still check the metric before every upgrade regardless: default-on protects an instance going forward from the moment it reconciles, it does not retroactively manufacture backup history for an instance that's been running unreconciled, and an explicit backup.enabled: false remains a valid opt-out you may have set deliberately.

  • S3 data backup (spec.backups.data): confirm it's configured and, if restoreTest is set, that the most recent automated restore test succeeded. Do not proceed on the strength of an unverified backup.

Confirm the instance is healthy before you touch it

kubectl get nci $NAME -n $NS -o jsonpath='{.status.conditions[?(@.type=="Ready")]}' | jq
kubectl get nci $NAME -n $NS -o jsonpath='{.status.conditions[?(@.type=="AppsHealthy")]}' | jq

Don't start an upgrade against an instance that isn't Ready, or whose AppsHealthy condition already reports AppsMissing — fix that first, so a genuine post-upgrade regression isn't confused with a pre-existing problem you carried into the upgrade.

Also check that no upgrade is already in flight — starting a second one while the first is still converging is how you end up debugging two things at once:

kubectl get nci $NAME -n $NS -o jsonpath='{.status.upgrade.phase}'   # empty or "Completed"

Know the version-resolution model before you touch spec.version

spec.version is only ever set explicitly at creation (or by you, later) — the operator never changes it on its own. What it resolves to can move, though:

  1. spec.helm.version (escape hatch) bypasses resolution entirely if set.
  2. Otherwise, spec.version (or the profile's defaults.version if the instance doesn't set its own — instance always wins) resolves through the NextcloudVersionMap CRD: an alias (latest, stable), a prefix ("32" → newest known 32.x.y), or an exact version.
  3. Every reconcile that changes the resolved version — whether from an edit to spec.version or from an alias/prefix re-resolving against an unchanged spec — is checked by validate_upgrade_path against the currently-deployed status.versionResolution.resolvedVersion: downgrades are always blocked, and so is skipping more than one major in a single step. An explicit spec.version edit that violates this fails loud (status.phase: Failed); a version-map re-resolution that violates it instead holds the instance on its current version and logs a warning, without failing anything (protects the fleet from a bad NextcloudVersionMap edit). Full detail: Upgrades → Upgrade-path validation.

1. Patch and minor upgrades

Same-major patch (31.0.8 → 31.0.9) and minor (31.0.x → 31.1.x) upgrades. Pick one of three ways to apply, per instance:

Path Trigger Respects the maintenance window? Use when
Explicit spec.version bump You edit the spec N/A — applies as soon as the reconcile runs One-off, deliberate change; strict change control
k8s.bnerd.com/upgrade-now annotation You annotate No — applies immediately, bypassing the window and upgradePolicy.mode entirely An already-resolved update needs to land now, regardless of mode
spec.upgradePolicy: {mode: auto, patchUpgrades, minorUpgrades} Automatic, on drift Yes — applies at spec.maintenance.windowStart Fleet-wide automation with a controlled blast-radius window

spec.maintenance.autoUpdate: true still works but is a deprecated alias for upgradePolicy: {mode: auto, patchUpgrades: true, minorUpgrades: false} — prefer upgradePolicy directly for new instances; if both are set on the same instance, upgradePolicy wins and a reason: UpgradePolicyAliasIgnored warning event fires.

Procedure for a manual (explicit-bump or upgrade-now) patch/minor upgrade:

  1. Confirm a green backup exists (step 0 above) — cheap insurance even for a patch.
  2. Apply the change: edit spec.version, or annotate:
    kubectl annotate nci $NAME -n $NS \
      k8s.bnerd.com/upgrade-now=$(date -u +"%Y-%m-%dT%H:%M:%SZ") --overwrite
    
  3. Watch status.phase go Ready → Deploying → Ready.
  4. Confirm post-upgrade tasks ran (see Post-upgrade verification below).

Full detail and the fleet-automation options: Upgrades → Patch and minor upgrades and Upgrades → Fleet strategy.

2. Major upgrades

Major hops are never automated, by design. upgradePolicy has no majorUpgrades field — there is no configuration that makes a major version change happen on its own; it always requires you to explicitly edit spec.version, one major at a time. Nextcloud itself doesn't support skipping majors, and the operator enforces the same rule (see Upgrade-path validation above) so a two-major jump fails before anything is touched.

Checklist per major hop (going from 30 → 32 means running this sequence twice: 30 → 31, then 31 → 32):

  1. Back up first (step 0 above) — non-negotiable for a major hop.
  2. Validate on a test instance first if at all possible — see Dry-run / canary testing today below for the currently-supported pattern.
  3. Set spec.version to the next major only (e.g. "31" or an exact "31.0.9"), never the final target — the operator rejects a skip.
  4. Wait for status.phase: Ready. Don't queue the next hop while this one is still Deploying.
  5. Wait for post-upgrade tasks to finish for this hop before starting the next — running repair/index tasks against a schema about to migrate again is wasted work.
  6. Spot-check the app (step 4, Post-upgrade verification).
  7. Repeat from step 3 for each remaining major.

Full checklist detail: Upgrades → Major upgrades.

3. What the operator handles automatically

Mechanism
Rolling the HelmRelease Flux applies the new chart/image once the operator writes the resolved version's values.
Maintenance window across the whole upgrade (0.23.0) The operator turns maintenance mode ON before patching the HelmRelease and lifts it only after the instance's declared apps are verified enabled again. The image's own toggle only spans the schema migration; Nextcloud's updater leaves a pre-existing maintenance flag alone, so the outer window survives occ upgrade. See Upgrades → Post-upgrade app convergence.
Running occ upgrade / DB migration The official Nextcloud image's own entrypoint, on pod start, inside the operator's window — the operator never execs this itself.
Post-upgrade app convergence (0.23.0) After the roll, once occ status reports the target version: app:update --all (which updates DISABLED apps — the store-lag healer), re-assert every declared app including user_oidc, then verify with app:list. Retried on the instance timer, with no attempt cap, until it converges. Results per step in status.upgrade.steps[], with exit codes.
Held-not-rolled-back failure semantics A stuck upgrade (spec.timeout elapsed, or the release unhealthy) is held on the failed revision, never auto-rolled-back (Flux upgrade remediation is disabled on operator-generated HelmReleases). Rolling old code back onto a schema a failed release already migrated forward crash-loops permanently — see Upgrades → A failed upgrade holds. Surfaced via status.conditions[type=UpgradeFailed] (reason: HelmUpgradeFailed) and a deduped HelmUpgradeStuck Warning event. Default Helm timeout is 15 minutes; override per instance with spec.helm.timeout for large migrations.
UpdateAvailable visibility status.conditions[type=UpdateAvailable] reports whether a newer version currently resolves, independent of upgradePolicy.mode and whether a maintenance window is even configured — a mode: manual fleet still gets the signal.
Post-upgrade occ repair tasks Two passes, deliberately. A FAST db:add-missing-indices + maintenance:repair inside the upgrade window (0.23.0) — --include-expensive can run for hours and is not something to hold an instance closed for. And the existing expensive set, unchanged: from the update handler's tail on a spec.version edit (which lands inside the window, against the old pod), and from the maintenance timer once it notices the version changed — for every instance since 0.23.0, not only those with a windowStart, and never while an upgrade is in flight. Seeing maintenance:repair twice per bump is expected — see Upgrades → Post-upgrade app convergence for how to tell the two apart in status.
Per-instance occ serialization See occ command parallelism below.

occ command parallelism during an upgrade

All operator-initiated occ execs against a given instance — NextcloudCommand runs, the admin-credential apply step, periodic/post-upgrade maintenance tasks, and the apps-health check — acquire a shared per-instance lock (operator/utils/occ.py, shipped 0.20.0) before running, so two concurrent operator-triggered execs against the same instance queue rather than race, even across different trigger sources. Execs against different instances still run fully in parallel — this is per-instance, not a global lock. A request that can't acquire the lock within OCC_LOCK_TIMEOUT (default 300s) fails with a retryable error rather than blocking forever. On top of that, Nextcloud's own maintenance mode (set by the image entrypoint during occ upgrade) independently blocks most occ subcommands for the duration of the migration itself.

Caveat — this lock only covers occ execs the operator itself issues. A manual kubectl exec ... occ ... bypasses the lock entirely; it will happily race a migration or an in-flight NextcloudCommand. During an upgrade window (or any time, really), run ad-hoc occ commands through NextcloudCommand CRs instead of raw kubectl exec — they queue through the same lock as everything else and their result lands in .status for you. See Running occ Commands and Upgrades → occ during an upgrade for the full mechanism (including why in-flight commands aren't cancelled when an upgrade starts).

4. Post-upgrade verification

# Version actually applied
kubectl get nci $NAME -n $NS -o jsonpath='{.status.versionResolution}' | jq

# The upgrade flow's own record: phase, pending apps, per-step results (0.23.0)
kubectl get nci $NAME -n $NS -o jsonpath='{.status.upgrade}' | jq
# expect: phase Completed, pendingApps [], maintenanceHeld false

# Post-upgrade maintenance ran for THIS version
kubectl get nci $NAME -n $NS -o jsonpath='{.status.maintenance}' | jq
# expect: lastRunTrigger: post-upgrade, lastVersion matching the new resolved version

# occ's own view, via NextcloudCommand rather than kubectl exec
kubectl apply -f - <<'EOF'
apiVersion: k8s.bnerd.com/v1alpha1
kind: NextcloudCommand
metadata:
  name: post-upgrade-status-check
  namespace: NAMESPACE
spec:
  targetRef: {kind: NextcloudInstance, name: NAME}
  commands: [["status", "--output=json"]]
EOF

Conditions to watch:

Condition What it tells you
Ready HelmRelease, Deployment, Endpoints, Ingress all ready — basic liveness.
UpgradeFailed True/HelmUpgradeFailed means the release is stuck, not just slow — see Rollback below. Absent (never written) on a healthy upgrade.
AppsHealthy False/PostUpgradeAppsDisabled means the upgrade left a declared app disabled and the operator is still retrying; False/AppsMissing is the older, steady-state signal from the apps-health check. Either way the app state is the thing to look at, not the upgrade — see App compatibility below.
UpgradeInProgress True while the flow is running (reason = the phase). False/Completed when every declared app is verified enabled. If a human used the accept annotation, the reason says which kind of accept it was: AcceptedWithPendingApps (verified, some apps knowingly left disabled), AcceptedAfterVerification (verified clean, window lifted by hand), or AcceptedBeforeConvergence (the apps were never checked for this upgrade).
UpgradeStuck True/NoProgress means a phase has not advanced for two hours — including an AppsPending that is still holding users out. Absent on a healthy upgrade.
UpdateAvailable Should flip to False/UpToDate once the resolved target matches what's deployed. Still True means the upgrade you meant to apply didn't actually land as the new baseline (or another update is already pending).

5. Dry-run / canary testing today

Blessed pattern: a test pool pinned to an older version, upgraded first. There is no built-in clone/dry-run feature (see Roadmap notes below) — until one exists, the currently-supported equivalent is to keep a NextcloudInstance (or a small NextcloudPool) on the same profile as production, pinned a major or two behind, and always run each major hop against it first:

  1. Create (or keep standing) a test instance using the same profile as the production instance/cohort you're planning to upgrade.
  2. Restore a recent production backup into it (same restore procedure as Rollback below, minus the allow-unsafe-version-change annotation, since the test instance starts from the pre-upgrade version already).
  3. Run the upgrade — or the next major hop — against the test instance first, following the Major upgrades checklist above.
  4. Validate via NextcloudCommand — occ status, occ app:list --output=json, and any app-specific sanity checks — before repeating the change against production.
  5. Keep the test instance a version (or a major) behind production going forward, so it's always the next hop's rehearsal target, not a one-off.

This is a genuine, supported way to de-risk an upgrade today — it's just a manual procedure built from the primitives above (profiles, pools, restore, NextcloudCommand), not a dedicated feature. See Upgrades → Staging / dry-run workaround for the underlying steps.

App compatibility today

Two separate things, and it is worth keeping them apart.

Creation-time installs are still best-effort. spec.apps entries are installed by occ app:install <app> || true in the chart's before-starting hook, which runs on pod start. A failed install there is still silent; it surfaces after the fact via the AppsHealthy condition, not as a gate.

Upgrades are no longer best-effort. As of 0.23.0 the operator owns app state across a version change: after the roll it runs app:update --all, re-asserts every declared app (including user_oidc when spec.oidc.enabled), verifies with app:list, and records every step's exit code in status.upgrade.steps[]. An app it cannot enable keeps the instance in AppsPending and gets retried — indefinitely, because the usual cause is an app-store release that has not been published yet, which resolves on its own in days. See Upgrades → Post-upgrade app convergence.

You still decide. The pre-flight app-store check the operator runs before an upgrade is a Warning event (AppCompatWarning), never a gate, and it cannot tell a bundled app apart from an incompatible one. Before a major upgrade, check each installed app's compatibility against the Nextcloud app store yourself and validate via the test-pool pattern above — knowing in advance that user_oidc has nothing published for the target major is worth more than watching the operator retry after the fact.

6. Rollback

Nextcloud has no in-place downgrade. "Rollback" always means restoring pre-upgrade backups — never reverting spec.version on a live, already-migrated instance. Never manually roll the image back once DB migrations may have run: if a failed release already ran occ upgrade far enough to migrate the schema forward before crashing, pointing the old Nextcloud code at the new, migrated schema produces a permanent crash-loop, not a working rollback. This is exactly why 0.21.0 made a stuck upgrade hold instead of auto-rolling-back (see What the operator handles automatically above) — the operator has no way to know from the outside whether a given failure happened before or after the migration point, so it never guesses.

If you hit a stuck upgrade, start with the stuck-upgrade runbook — diagnose via the UpgradeFailed condition and Flux/Helm history, fix the underlying cause, and let it retry or re-trigger with k8s.bnerd.com/upgrade-now. Only fall back to a genuine rollback if the underlying cause can't be fixed forward:

  1. Stop traffic to the instance if it's serving users (outside the operator's scope).
  2. Restore the pre-upgrade S3 data backup and the pre-upgrade database backup (pgBackRest restore for managed Postgres, or your external DB's own procedure) to the point captured before the upgrade.
  3. Set k8s.bnerd.com/allow-unsafe-version-change: "true" on the NextcloudInstance — required because setting spec.version back to the pre-upgrade value is, from the operator's point of view, a downgrade.
  4. Set spec.version back to the exact pre-upgrade resolvedVersion you noted in step 0 — not a prefix or alias, so the resolved version matches what you restored.
  5. Wait for status.phase: Ready and spot-check the app.
  6. Remove the allow-unsafe-version-change annotation once stable — leaving it set disables downgrade/major-skip protection for every future reconcile, not just this one.

There is no automated rollback — this is a manual procedure end to end. Practice it against the test-pool instance (above) before you need it in production. Full detail: Upgrades → Backup & rollback.

7. Roadmap notes — what isn't built yet

Honest status on the two gaps this procedure works around manually:

  • Clone/canary upgrade dry-run — designed-but-not-yet-implemented. The roadmap describes cloning an instance from a snapshot, running the upgrade against the clone, reporting the result, and discarding it — deferred from 0.20.0, no committed release yet (docs/roadmap/release-0.21.md). The test-pool pattern in step 5 above is the supported stand-in until it lands.
  • A BLOCKING app-compatibility gate — still not implemented. 0.23.0 ships the advisory half: a pre-flight AppCompatWarning event before the roll, and a convergence loop that repairs app state afterwards (Upgrades → Post-upgrade app convergence). What does not exist is an upgradePolicy.appCompatGate that REFUSES an upgrade because an app is not ready. The advisory pre-flight is deliberately not a suitable basis for one: the app store's platform listing cannot distinguish a bundled app from an incompatible one, so gating on it would refuse upgrades for apps that are perfectly fine. Check app compatibility manually against the Nextcloud app store before a major hop.

See also

  • Upgrades — the full mechanism behind every step above: upgrade-path validation internals, the autoUpdate alias's exact behavior change history, fleet strategy patterns, and the complete stuck-upgrade runbook.
  • Version Management — how spec.version resolves to a chart version, the NextcloudVersionMap CRD, image resolution, profile-level pinning.
  • Operations & Annotations — the full annotation reference, including k8s.bnerd.com/reconcile, k8s.bnerd.com/run-maintenance, and admin credential re-apply behavior on upgrade.
  • Running occ Commands — the NextcloudCommand CRD used throughout this procedure for validation, spot-checks, and ad-hoc commands during an upgrade window.
  • Managed PostgreSQL — pgBackRest backup configuration for database.managed: true.
  • Monitoring — the pgbackrest_last_backup_completion_timestamp_seconds metric and a ready-to-use staleness alert.