Standard Upgrade Procedure¶
A single, top-to-bottom checklist for upgrading a Nextcloud instance safely. The
Upgrades, Operations & Annotations,
Running occ Commands, and Version Management
guides remain the canonical reference for why each piece behaves the way it does —
this page assembles them into the procedure you actually run, in order, with nothing
left to piece together yourself.
0. Before any upgrade¶
Verify backup freshness¶
- Managed database (
spec.database.managed: true): check thepgbackrest_last_backup_completion_timestamp_seconds{namespace,instance,repo,type}Prometheus gauge, or read the underlyingPerconaPGBackupCRs directly if you don't have the metric pipeline handy:
# Via the metric (per repo/type, Succeeded backups only)
curl -s http://<operator-pod>:9090/metrics | grep pgbackrest_last_backup
# Via the CRs directly
kubectl get perconapgbackups.pgv2.percona.com -n $NS -l postgres-operator.crunchydata.com/cluster=<name>-pg
Absence of the metric series is itself the signal that no Succeeded backup has
ever completed for that instance — there is no zero/placeholder value. See
Monitoring → Backup freshness (pgBackRest)
for the full source/cadence detail and a ready-to-use PgBackRestBackupStale alert.
As of 0.21.3, scheduled backups default to ON for managed PostgreSQL (see
Managed PostgreSQL → Backups) — an instance created or
reconciled since then should show a real, recurring Succeeded backup here, not just the
one-off replica-create snapshot taken at cluster birth. Still check the metric before
every upgrade regardless: default-on protects an instance going forward from the moment
it reconciles, it does not retroactively manufacture backup history for an instance
that's been running unreconciled, and an explicit backup.enabled: false remains a valid
opt-out you may have set deliberately.
- S3 data backup (
spec.backups.data): confirm it's configured and, ifrestoreTestis set, that the most recent automated restore test succeeded. Do not proceed on the strength of an unverified backup.
Confirm the instance is healthy before you touch it¶
kubectl get nci $NAME -n $NS -o jsonpath='{.status.conditions[?(@.type=="Ready")]}' | jq
kubectl get nci $NAME -n $NS -o jsonpath='{.status.conditions[?(@.type=="AppsHealthy")]}' | jq
Don't start an upgrade against an instance that isn't Ready, or whose AppsHealthy
condition already reports AppsMissing — fix that first, so a genuine post-upgrade
regression isn't confused with a pre-existing problem you carried into the upgrade.
Also check that no upgrade is already in flight — starting a second one while the first is still converging is how you end up debugging two things at once:
Know the version-resolution model before you touch spec.version¶
spec.version is only ever set explicitly at creation (or by you, later) — the
operator never changes it on its own. What it resolves to can move, though:
spec.helm.version(escape hatch) bypasses resolution entirely if set.- Otherwise,
spec.version(or the profile'sdefaults.versionif the instance doesn't set its own — instance always wins) resolves through theNextcloudVersionMapCRD: an alias (latest,stable), a prefix ("32"→ newest known32.x.y), or an exact version. - Every reconcile that changes the resolved version — whether from an edit to
spec.versionor from an alias/prefix re-resolving against an unchanged spec — is checked byvalidate_upgrade_pathagainst the currently-deployedstatus.versionResolution.resolvedVersion: downgrades are always blocked, and so is skipping more than one major in a single step. An explicitspec.versionedit that violates this fails loud (status.phase: Failed); a version-map re-resolution that violates it instead holds the instance on its current version and logs a warning, without failing anything (protects the fleet from a badNextcloudVersionMapedit). Full detail: Upgrades → Upgrade-path validation.
1. Patch and minor upgrades¶
Same-major patch (31.0.8 → 31.0.9) and minor (31.0.x → 31.1.x) upgrades. Pick
one of three ways to apply, per instance:
| Path | Trigger | Respects the maintenance window? | Use when |
|---|---|---|---|
Explicit spec.version bump |
You edit the spec | N/A — applies as soon as the reconcile runs | One-off, deliberate change; strict change control |
k8s.bnerd.com/upgrade-now annotation |
You annotate | No — applies immediately, bypassing the window and upgradePolicy.mode entirely |
An already-resolved update needs to land now, regardless of mode |
spec.upgradePolicy: {mode: auto, patchUpgrades, minorUpgrades} |
Automatic, on drift | Yes — applies at spec.maintenance.windowStart |
Fleet-wide automation with a controlled blast-radius window |
spec.maintenance.autoUpdate: true still works but is a deprecated alias for
upgradePolicy: {mode: auto, patchUpgrades: true, minorUpgrades: false} — prefer
upgradePolicy directly for new instances; if both are set on the same instance,
upgradePolicy wins and a reason: UpgradePolicyAliasIgnored warning event fires.
Procedure for a manual (explicit-bump or upgrade-now) patch/minor upgrade:
- Confirm a green backup exists (step 0 above) — cheap insurance even for a patch.
- Apply the change: edit
spec.version, or annotate: - Watch
status.phasegoReady → Deploying → Ready. - Confirm post-upgrade tasks ran (see Post-upgrade verification below).
Full detail and the fleet-automation options: Upgrades → Patch and minor upgrades and Upgrades → Fleet strategy.
2. Major upgrades¶
Major hops are never automated, by design. upgradePolicy has no majorUpgrades
field — there is no configuration that makes a major version change happen on its
own; it always requires you to explicitly edit spec.version, one major at a time.
Nextcloud itself doesn't support skipping majors, and the operator enforces the same
rule (see Upgrade-path validation above) so a two-major jump
fails before anything is touched.
Checklist per major hop (going from 30 → 32 means running this sequence twice: 30 → 31, then 31 → 32):
- Back up first (step 0 above) — non-negotiable for a major hop.
- Validate on a test instance first if at all possible — see Dry-run / canary testing today below for the currently-supported pattern.
- Set
spec.versionto the next major only (e.g."31"or an exact"31.0.9"), never the final target — the operator rejects a skip. - Wait for
status.phase: Ready. Don't queue the next hop while this one is stillDeploying. - Wait for post-upgrade tasks to finish for this hop before starting the next — running repair/index tasks against a schema about to migrate again is wasted work.
- Spot-check the app (step 4, Post-upgrade verification).
- Repeat from step 3 for each remaining major.
Full checklist detail: Upgrades → Major upgrades.
3. What the operator handles automatically¶
| Mechanism | |
|---|---|
Rolling the HelmRelease |
Flux applies the new chart/image once the operator writes the resolved version's values. |
| Maintenance window across the whole upgrade | (0.23.0) The operator turns maintenance mode ON before patching the HelmRelease and lifts it only after the instance's declared apps are verified enabled again. The image's own toggle only spans the schema migration; Nextcloud's updater leaves a pre-existing maintenance flag alone, so the outer window survives occ upgrade. See Upgrades → Post-upgrade app convergence. |
Running occ upgrade / DB migration |
The official Nextcloud image's own entrypoint, on pod start, inside the operator's window — the operator never execs this itself. |
| Post-upgrade app convergence | (0.23.0) After the roll, once occ status reports the target version: app:update --all (which updates DISABLED apps — the store-lag healer), re-assert every declared app including user_oidc, then verify with app:list. Retried on the instance timer, with no attempt cap, until it converges. Results per step in status.upgrade.steps[], with exit codes. |
| Held-not-rolled-back failure semantics | A stuck upgrade (spec.timeout elapsed, or the release unhealthy) is held on the failed revision, never auto-rolled-back (Flux upgrade remediation is disabled on operator-generated HelmReleases). Rolling old code back onto a schema a failed release already migrated forward crash-loops permanently — see Upgrades → A failed upgrade holds. Surfaced via status.conditions[type=UpgradeFailed] (reason: HelmUpgradeFailed) and a deduped HelmUpgradeStuck Warning event. Default Helm timeout is 15 minutes; override per instance with spec.helm.timeout for large migrations. |
UpdateAvailable visibility |
status.conditions[type=UpdateAvailable] reports whether a newer version currently resolves, independent of upgradePolicy.mode and whether a maintenance window is even configured — a mode: manual fleet still gets the signal. |
Post-upgrade occ repair tasks |
Two passes, deliberately. A FAST db:add-missing-indices + maintenance:repair inside the upgrade window (0.23.0) — --include-expensive can run for hours and is not something to hold an instance closed for. And the existing expensive set, unchanged: from the update handler's tail on a spec.version edit (which lands inside the window, against the old pod), and from the maintenance timer once it notices the version changed — for every instance since 0.23.0, not only those with a windowStart, and never while an upgrade is in flight. Seeing maintenance:repair twice per bump is expected — see Upgrades → Post-upgrade app convergence for how to tell the two apart in status. |
Per-instance occ serialization |
See occ command parallelism below. |
occ command parallelism during an upgrade¶
All operator-initiated occ execs against a given instance — NextcloudCommand
runs, the admin-credential apply step, periodic/post-upgrade maintenance tasks, and
the apps-health check — acquire a shared per-instance lock
(operator/utils/occ.py, shipped 0.20.0) before running, so two concurrent
operator-triggered execs against the same instance queue rather than race, even
across different trigger sources. Execs against different instances still run fully
in parallel — this is per-instance, not a global lock. A request that can't acquire
the lock within OCC_LOCK_TIMEOUT (default 300s) fails with a retryable error rather
than blocking forever. On top of that, Nextcloud's own maintenance mode (set by the
image entrypoint during occ upgrade) independently blocks most occ subcommands for
the duration of the migration itself.
Caveat — this lock only covers occ execs the operator itself issues. A manual
kubectl exec ... occ ... bypasses the lock entirely; it will happily race a
migration or an in-flight NextcloudCommand. During an upgrade window (or any time,
really), run ad-hoc occ commands through NextcloudCommand CRs instead of raw
kubectl exec — they queue through the same lock as everything else and their result
lands in .status for you. See Running occ Commands and
Upgrades → occ during an upgrade for the full
mechanism (including why in-flight commands aren't cancelled when an upgrade starts).
4. Post-upgrade verification¶
# Version actually applied
kubectl get nci $NAME -n $NS -o jsonpath='{.status.versionResolution}' | jq
# The upgrade flow's own record: phase, pending apps, per-step results (0.23.0)
kubectl get nci $NAME -n $NS -o jsonpath='{.status.upgrade}' | jq
# expect: phase Completed, pendingApps [], maintenanceHeld false
# Post-upgrade maintenance ran for THIS version
kubectl get nci $NAME -n $NS -o jsonpath='{.status.maintenance}' | jq
# expect: lastRunTrigger: post-upgrade, lastVersion matching the new resolved version
# occ's own view, via NextcloudCommand rather than kubectl exec
kubectl apply -f - <<'EOF'
apiVersion: k8s.bnerd.com/v1alpha1
kind: NextcloudCommand
metadata:
name: post-upgrade-status-check
namespace: NAMESPACE
spec:
targetRef: {kind: NextcloudInstance, name: NAME}
commands: [["status", "--output=json"]]
EOF
Conditions to watch:
| Condition | What it tells you |
|---|---|
Ready |
HelmRelease, Deployment, Endpoints, Ingress all ready — basic liveness. |
UpgradeFailed |
True/HelmUpgradeFailed means the release is stuck, not just slow — see Rollback below. Absent (never written) on a healthy upgrade. |
AppsHealthy |
False/PostUpgradeAppsDisabled means the upgrade left a declared app disabled and the operator is still retrying; False/AppsMissing is the older, steady-state signal from the apps-health check. Either way the app state is the thing to look at, not the upgrade — see App compatibility below. |
UpgradeInProgress |
True while the flow is running (reason = the phase). False/Completed when every declared app is verified enabled. If a human used the accept annotation, the reason says which kind of accept it was: AcceptedWithPendingApps (verified, some apps knowingly left disabled), AcceptedAfterVerification (verified clean, window lifted by hand), or AcceptedBeforeConvergence (the apps were never checked for this upgrade). |
UpgradeStuck |
True/NoProgress means a phase has not advanced for two hours — including an AppsPending that is still holding users out. Absent on a healthy upgrade. |
UpdateAvailable |
Should flip to False/UpToDate once the resolved target matches what's deployed. Still True means the upgrade you meant to apply didn't actually land as the new baseline (or another update is already pending). |
5. Dry-run / canary testing today¶
Blessed pattern: a test pool pinned to an older version, upgraded first. There is
no built-in clone/dry-run feature (see Roadmap notes
below) — until one exists, the currently-supported equivalent is to keep a
NextcloudInstance (or a small NextcloudPool) on the same profile as production,
pinned a major or two behind, and always run each major hop against it first:
- Create (or keep standing) a test instance using the same profile as the production instance/cohort you're planning to upgrade.
- Restore a recent production backup into it (same restore procedure as
Rollback below, minus the
allow-unsafe-version-changeannotation, since the test instance starts from the pre-upgrade version already). - Run the upgrade — or the next major hop — against the test instance first, following the Major upgrades checklist above.
- Validate via
NextcloudCommand—occ status,occ app:list --output=json, and any app-specific sanity checks — before repeating the change against production. - Keep the test instance a version (or a major) behind production going forward, so it's always the next hop's rehearsal target, not a one-off.
This is a genuine, supported way to de-risk an upgrade today — it's just a manual
procedure built from the primitives above (profiles, pools, restore, NextcloudCommand),
not a dedicated feature. See Upgrades → Staging / dry-run
workaround for the underlying steps.
App compatibility today¶
Two separate things, and it is worth keeping them apart.
Creation-time installs are still best-effort. spec.apps entries are installed by
occ app:install <app> || true in the chart's before-starting hook, which runs on pod
start. A failed install there is still silent; it surfaces after the fact via the
AppsHealthy condition, not as a gate.
Upgrades are no longer best-effort. As of 0.23.0 the operator owns app state
across a version change: after the roll it runs app:update --all, re-asserts every
declared app (including user_oidc when spec.oidc.enabled), verifies with
app:list, and records every step's exit code in status.upgrade.steps[]. An app it
cannot enable keeps the instance in AppsPending and gets retried — indefinitely,
because the usual cause is an app-store release that has not been published yet, which
resolves on its own in days. See
Upgrades → Post-upgrade app convergence.
You still decide. The pre-flight app-store check the operator runs before an
upgrade is a Warning event (AppCompatWarning), never a gate, and it cannot tell a
bundled app apart from an incompatible one. Before a major upgrade, check each
installed app's compatibility against the Nextcloud app store
yourself and validate via the test-pool pattern above — knowing in advance that
user_oidc has nothing published for the target major is worth more than watching the
operator retry after the fact.
6. Rollback¶
Nextcloud has no in-place downgrade. "Rollback" always means restoring
pre-upgrade backups — never reverting spec.version on a live, already-migrated
instance. Never manually roll the image back once DB migrations may have run:
if a failed release already ran occ upgrade far enough to migrate the schema
forward before crashing, pointing the old Nextcloud code at the new, migrated schema
produces a permanent crash-loop, not a working rollback. This is exactly why 0.21.0
made a stuck upgrade hold instead of auto-rolling-back (see
What the operator handles automatically
above) — the operator has no way to know from the outside whether a given failure
happened before or after the migration point, so it never guesses.
If you hit a stuck upgrade, start with the stuck-upgrade
runbook — diagnose via the UpgradeFailed
condition and Flux/Helm history, fix the underlying cause, and let it retry or
re-trigger with k8s.bnerd.com/upgrade-now. Only fall back to a genuine rollback if
the underlying cause can't be fixed forward:
- Stop traffic to the instance if it's serving users (outside the operator's scope).
- Restore the pre-upgrade S3 data backup and the pre-upgrade database backup
(
pgBackRestrestore for managed Postgres, or your external DB's own procedure) to the point captured before the upgrade. - Set
k8s.bnerd.com/allow-unsafe-version-change: "true"on theNextcloudInstance— required because settingspec.versionback to the pre-upgrade value is, from the operator's point of view, a downgrade. - Set
spec.versionback to the exact pre-upgraderesolvedVersionyou noted in step 0 — not a prefix or alias, so the resolved version matches what you restored. - Wait for
status.phase: Readyand spot-check the app. - Remove the
allow-unsafe-version-changeannotation once stable — leaving it set disables downgrade/major-skip protection for every future reconcile, not just this one.
There is no automated rollback — this is a manual procedure end to end. Practice it against the test-pool instance (above) before you need it in production. Full detail: Upgrades → Backup & rollback.
7. Roadmap notes — what isn't built yet¶
Honest status on the two gaps this procedure works around manually:
- Clone/canary upgrade dry-run — designed-but-not-yet-implemented. The roadmap
describes cloning an instance from a snapshot, running the upgrade against the
clone, reporting the result, and discarding it — deferred from 0.20.0, no committed
release yet (
docs/roadmap/release-0.21.md). The test-pool pattern in step 5 above is the supported stand-in until it lands. - A BLOCKING app-compatibility gate — still not implemented. 0.23.0 ships the
advisory half: a pre-flight
AppCompatWarningevent before the roll, and a convergence loop that repairs app state afterwards (Upgrades → Post-upgrade app convergence). What does not exist is anupgradePolicy.appCompatGatethat REFUSES an upgrade because an app is not ready. The advisory pre-flight is deliberately not a suitable basis for one: the app store's platform listing cannot distinguish a bundled app from an incompatible one, so gating on it would refuse upgrades for apps that are perfectly fine. Check app compatibility manually against the Nextcloud app store before a major hop.
See also¶
- Upgrades — the full mechanism behind every step above: upgrade-path
validation internals, the
autoUpdatealias's exact behavior change history, fleet strategy patterns, and the complete stuck-upgrade runbook. - Version Management — how
spec.versionresolves to a chart version, theNextcloudVersionMapCRD, image resolution, profile-level pinning. - Operations & Annotations — the full annotation reference,
including
k8s.bnerd.com/reconcile,k8s.bnerd.com/run-maintenance, and admin credential re-apply behavior on upgrade. - Running occ Commands — the
NextcloudCommandCRD used throughout this procedure for validation, spot-checks, and ad-hoc commands during an upgrade window. - Managed PostgreSQL —
pgBackRestbackup configuration fordatabase.managed: true. - Monitoring — the
pgbackrest_last_backup_completion_timestamp_secondsmetric and a ready-to-use staleness alert.