Upgrades¶
Upgrading Nextcloud is a version bump on the CR, but doing it safely across a fleet is a platform-team procedure: back up first, step through majors one at a time, and validate before moving on. This guide is the production runbook — the mechanism behind every step. For a single top-to-bottom checklist that assembles this guide with Operations, occ Commands, and Version Management into one procedure, see Standard Upgrade Procedure. For how spec.version resolves to a chart version in the first place, see Version Management.
What the operator automates vs. what you do¶
| Operator | You | |
|---|---|---|
Resolve spec.version to a chart + image |
✅ via NextcloudVersionMap |
— |
| Roll the Helm release | ✅ Flux applies the new HelmRelease |
— |
| Hold maintenance mode across the WHOLE window | ✅ 0.23.0 — opened before the roll, lifted only after apps reconverge | — |
Run occ upgrade / DB migration |
✅ official image entrypoint, on pod start, inside the operator's window | — |
| Re-enable apps the upgrade disabled | ✅ 0.23.0 — app:update --all + re-assert, retried until it converges |
— |
Run post-upgrade occ repair tasks |
✅ in-window (fast repair) + maintenance timer (expensive repair) | — |
| Block unsafe major-skips and downgrades | ✅ upgrade-path validation | — |
Decide when to change spec.version |
— | ✅ you |
| Pre-upgrade backup + restore test | — | ✅ you |
| Step through intermediate majors | — | ✅ you |
| App compatibility check | ⚠️ 0.23.0 — advisory pre-flight warning only, never a gate | ✅ you (still your call, see below) |
| Rollback | — | ✅ you (restore from backup) |
The operator still does not exec occ upgrade itself — it never touches the pod for the migration. It changes the HelmRelease; the Nextcloud container image's own entrypoint detects the version bump and runs the DB migration before serving traffic again. What changed in 0.23.0 is everything around that step: the operator now opens Nextcloud's maintenance mode before the roll and holds it until the instance's declared apps are verified enabled again — see Post-upgrade app convergence. The image's own maintenance toggle only ever spanned the schema migration, and Nextcloud's updater leaves a maintenance flag it did not set alone, so the operator's outer window survives occ upgrade intact.
A failed upgrade holds — it does not roll back (0.21.0+)¶
As of 0.21.0, a Helm upgrade that fails partway (spec.timeout elapses, or the release comes back unhealthy) is held on the failed revision, not automatically rolled back. This is a deliberate change from earlier versions, and it matters because Nextcloud's own DB migration is not reversible: if the failed release already ran occ upgrade far enough to migrate the schema forward before crashing, rolling the Helm release back to the previous chart/image is not a rollback at all — it re-deploys the old Nextcloud code against the new, already-migrated schema, which crash-loops permanently. There is no way for Flux or the operator to know from the outside whether a given failure happened before or after the migration point, so as of 0.21.0 the operator never lets Flux guess: it disables Flux's automatic upgrade remediation entirely (spec.upgrade.remediation.retries: 0 on the generated HelmRelease) and instead surfaces the stuck state for a human to diagnose. Install remediation is unchanged (retries: 3) — a failed first install has no migrated data at risk, so self-healing there is still safe.
Two related changes ship alongside this:
- Default Helm timeout raised to 15 minutes. Flux's own default is 5 minutes, which large-migration upgrades (a big Nextcloud instance moving a major version) routinely exceed — a slow-but-otherwise-healthy migration hitting the old 5m ceiling looked identical to a genuinely broken upgrade. Override per instance with
spec.helm.timeout(a Go duration string, e.g."30m","1h") if even 15 minutes isn't enough for a particular instance's data size.
UpgradeFailedstatus condition. The operator's status-sync timer compares theHelmRelease'slastAttemptedRevisionagainstlastAppliedRevision: whenever the release is not ready and the two differ, it setsstatus.conditions[type=UpgradeFailed]toTruewithreason: HelmUpgradeFailed, and the message carries both revisions plus Flux's own failure message. It also emits aHelmUpgradeStuckWarning event — once per stuck revision, not once per 30-second reconcile tick, so a human investigating isn't flooded with duplicate events for the same still-unresolved failure. Once the revisions converge again (the upgrade is fixed and completes, or a correctedupgrade-nowre-applies successfully), the condition flips toFalsewithreason: Recovered. An instance that was never stuck gets noUpgradeFailedwrite at all —kubectl describestays clean for healthy instances, it doesn't carry a permanentFalsecondition nobody asked for.
Where the applied revision comes from depends on your Flux version, and the operator handles both. Flux ≥ 2.3 serves helm.toolkit.fluxcd.io/v2, which dropped status.lastAppliedRevision in favour of status.history — so the operator reads the chart version of the newest history snapshot Helm still reports as deployed (a failed upgrade leaves the previous release deployed and the new one failed, which is exactly the comparison this condition wants). On Flux 2.0–2.2, where a beta API was the storage version, it reads status.lastAppliedRevision directly. An instance whose first install never succeeded has no deployed snapshot at all, so it is never reported as a stuck upgrade.
kubectl get helmrelease my-nextcloud-nextcloud \
-o jsonpath='{range .status.history[*]}{.chartVersion}{"\t"}{.status}{"\n"}{end}'
Existing HelmReleases converge automatically — the disabled remediation and raised timeout apply the next time the operator writes to an instance's HelmRelease (any on_create/on_update/force-reconcile patch), not via a one-time migration step. A fleet upgrading the operator to 0.21.0 doesn't need to touch every NextcloudInstance by hand; the next normal reconcile (a spec edit, a version bump, or a k8s.bnerd.com/reconcile annotation) picks up the new remediation/timeout settings.
Runbook: a stuck upgrade¶
- Diagnose. Start with the condition itself — it already names both revisions and carries Flux's failure message:
kubectl describe nci my-nextcloud
# look for: Type: UpgradeFailed, Status: True, Reason: HelmUpgradeFailed
Then pull the full Helm history and the Flux controller's own logs for more detail than fits in the condition message:
helm history my-nextcloud-nextcloud -n my-namespace
kubectl logs -n flux-system deploy/helm-controller --since=1h | grep my-nextcloud-nextcloud
-
Fix the underlying cause. This is almost always outside the operator's control — a chart value that doesn't validate, a resource limit too low for the migration workload, an image pull failure, or the migration itself failing inside the pod (check
kubectl logson the Nextcloud pod for theocc upgradeoutput). Fix whateverhelm history/the pod logs point at. -
Let it retry, or re-trigger explicitly. Once the underlying cause is fixed, either wait for Flux's own reconcile interval to retry the same revision, or force it immediately with the on-demand upgrade annotation:
kubectl annotate nci my-nextcloud \
k8s.bnerd.com/upgrade-now=$(date -u +"%Y-%m-%dT%H:%M:%SZ") \
--overwrite -n my-namespace
UpgradeFailed flips to False/Recovered once lastAttemptedRevision and lastAppliedRevision converge again.
Never manually roll back the image once DB migrations may have run. Editing spec.version back to the previous value, or otherwise forcing the chart/image backward, is exactly the crash-loop scenario this feature exists to prevent — see the explanation above. If you genuinely need to undo a bad upgrade, that's a data restore, not a version edit: follow Backup & rollback below, which restores from a pre-upgrade backup rather than pointing a migrated DB at old code.
How an upgrade happens¶
- You change
spec.version(or a pinning profile'sdefaults.versionchanges and the instance re-resolves — see Version Management). - The operator resolves the new value through the
NextcloudVersionMapto a Nextcloud version + Helm chart version + image. - Upgrade-path validation runs against the resolved current and target versions (see below). An unsafe transition fails the reconcile with a permanent error before anything is touched.
- The operator turns Nextcloud's maintenance mode ON and records
status.upgrade(phase: MaintenanceOn,from,to) plus theUpgradeInProgresscondition. An advisory app-store pre-flight runs here too (below). (0.23.0) - The operator updates the
HelmRelease.status.phasemovesReady → Deploying,status.upgrade.phasemoves toRolling. - Flux reconciles the
HelmRelease; the new pod starts and the Nextcloud entrypoint runsocc upgrade. Because maintenance mode was already on, Nextcloud's updater leaves it on when it finishes (it only clears a flag it set itself) — so the instance stays closed to users. - The operator waits for
occ statusto report the target version, then converges the apps:app:update --all, re-assert every declared app, the in-window post-upgrade tasks, then verify withapp:list. (0.23.0 —status.upgrade.phase: Converging) - Maintenance mode is lifted and
status.upgrade.phasebecomesCompleted, with anUpgradeCompletedevent — orAppsPendingif declared apps are still not enabled (see below). status.versionResolutionrecordsrequestedVersion,resolvedVersion, andchartVersionfor the completed change.- The expensive post-upgrade task set (
maintenance:repair --include-expensive,db:convert-filecache-bigint, …) runs too, unchanged from earlier releases, as a separate pass from step 7's fast in-window repair. Which trigger you used decides when: aspec.versionedit runs it from the update handler's own tail, right after step 5 and therefore inside the maintenance window against the still-running old pod; an auto-update, versionmap drift,upgrade-nowor a reconcile annotation has no such tail. Either way the maintenance timer picks it up on its next tick oncestatus.maintenance.lastVersiondiffers — for every instance, with or without aspec.maintenance.windowStart, and it is skipped while an upgrade is still in flight. See the note below on seeingmaintenance:repairtwice.
kubectl get nci my-nextcloud -o jsonpath='{.status.phase}'
kubectl get nci my-nextcloud -o jsonpath='{.status.upgrade}' | jq
kubectl get nci my-nextcloud -o jsonpath='{.status.versionResolution}' | jq
kubectl get nci my-nextcloud -o jsonpath='{.status.maintenance}' | jq
Post-upgrade app convergence (0.23.0)¶
The failure this closes¶
Upgrading Nextcloud across a major disables every app that has no release compatible with the new version. That is the image entrypoint doing the right thing — an incompatible app would break the instance. It then tries to update those apps from the app store and re-enable them, which usually works, because the store usually already serves a compatible release by the time you upgrade.
When it does not, nothing ever tries again:
- the lifecycle hooks that install and enable apps run only on pod start;
occ app:installon an already-installed app reports "already installed" and does not update it;occ app:enablerefuses the installed, incompatible version;- every one of those hook commands ends in
|| true, so all of it is silent.
The app stays disabled until a human notices. For user_oidc — an instance whose only login path is SSO — "until a human notices" is an authentication outage.
Two states reach it: the store has not published a compatible release yet (normal in the days after a major Nextcloud release), or the store is unreachable from the cluster at all (egress policy, DNS, an outage).
What the operator does instead¶
The recovery is a convergence loop owned by the operator's timer, not a hook firing once at the wrong moment. Every version change goes through it — a manual spec.version edit, a version-map alias moving, an upgradePolicy auto-update, or k8s.bnerd.com/upgrade-now.
Phase (status.upgrade.phase) |
What is happening |
|---|---|
MaintenanceOn |
Maintenance mode enabled; the HelmRelease has not been patched yet. |
Rolling |
HelmRelease patched. Waiting for the pod to report the target version via occ status — Flux reporting Ready is not enough, because converging against the old pod would prove nothing. |
Converging |
occ app:update --all → re-assert each declared app (app:install, then app:enable) → db:add-missing-indices + a fast maintenance:repair → verify with occ app:list. |
AppsPending |
Declared apps are still not enabled. The loop retries; see below. |
Completed |
Every declared app verified enabled, maintenance lifted. |
app:update --all runs first on every pass, and that ordering is the entire fix: unlike app:install, it updates apps that are currently disabled — which is exactly the state the upgrade left them in. Re-asserting before updating would just fail against the same incompatible version again.
The declared set is spec.apps (every entry not explicitly enabled: false) plus user_oidc whenever spec.oidc.enabled is set. user_oidc is reserved in spec.apps because it is managed through spec.oidc, which is precisely why the older AppsHealthy check never saw it — and it is the app whose disablement caused the outage this feature exists for.
Each step's result is recorded in status.upgrade.steps[] with its exit code. Nothing is swallowed.
maintenance:repair runs twice per upgrade — and only one of them is in the window
The in-window pass above runs the fast maintenance:repair, so an instance is not
held closed for a repair that can take hours. The pre-existing
maintenance:repair --include-expensive set still runs as well, but since 0.23.1 it
never runs while an upgrade is in flight: the maintenance timer picks it up once the
upgrade completes, at a per-instance jittered time.
Earlier releases ran it inline from the update handler the moment spec.version
changed — inside the maintenance window, against the pod Flux was in the middle of
replacing. That is not merely wasteful: the exec is killed when the pod goes away
(exit 143), and if the upgrade then stalls, nothing retries it. 0.23.0's docs called
the double-run "expected, not a bug", which was true of the double-run and wrong about
the kill. Fixed in 0.23.1.
Telling the two apart: the fast in-window pass is recorded in
status.upgrade.steps[] as maintenanceRepair; the expensive pass is recorded in
status.maintenance.tasks as repair with lastRunTrigger: post-upgrade. Both
honour spec.maintenance.tasks.repairAfterUpgrade: false.
kubectl get nci my-nextcloud -o jsonpath='{.status.upgrade}' | jq
kubectl get nci my-nextcloud -o jsonpath='{.status.upgrade.steps}' | jq '.[] | select(.success == false)'
The maintenance window, and when it is not lifted¶
The window opens before the roll and closes only when the apps are verified. If some declared app is still disabled at the end of a pass, what happens next depends on which app:
- A non-critical app (calendar, contacts, groupfolders, …): maintenance is lifted. Availability beats completeness — a missing calendar app must not keep an entire instance dark. The instance serves; the affected apps are unavailable;
AppsHealthyreportsPostUpgradeAppsDisabled. - A critical app: maintenance stays on. The critical set is
{user_oidc}whenspec.oidc.enabledis set, and empty otherwise. An instance whose only login path is dead is better presented as "in maintenance", with a clear page, than as "up" and unusable to everyone who tries to log in.
An instance with OIDC disabled — including every pool spare — therefore has an empty critical set and can never be held dark by this rule.
If the operator cannot reach the pod to lift the window, it does not report Completed: it holds in Converging and retries. "Completed" is terminal, and claiming it on a locked-out instance would be a lie the loop could no longer fix.
The retry loop¶
While AppsPending, every instance-status tick re-runs the update / re-assert / verify steps (not the migrations — those already ran):
- no attempt cap. Store lag resolves in days; giving up would restore exactly the terminal state this feature exists to end.
- at least 15 minutes apart, plus a deterministic per-instance jitter, so a fleet upgraded in lockstep does not poll the store in lockstep. Tune with
UPGRADE_RETRY_INTERVAL/UPGRADE_RETRY_JITTERon the operator deployment. - it heals by itself. The moment the store serves a compatible release, the next pass's
app:update --allpulls it,app:enablesucceeds, and the instance goesCompleted— no human action, no re-upgrade. - a lifted window is never re-closed by a retry. Once users have the instance back, a routine retry does not take it away again.
While an upgrade is in flight, the upgrade flow owns the AppsHealthy condition: the 15-minute maintenance timer's own apps-health check keeps refreshing status.apps (pure observation) but stops writing the condition until the upgrade reaches a terminal phase. Without that, the timer would flip AppsHealthy back to True every 15 minutes on exactly the instance this feature exists for — it reads spec.apps, which excludes user_oidc, so "SSO dead, everything in spec.apps fine" looks healthy to it.
The alertable signals are the AppsHealthy=False/PostUpgradeAppsDisabled condition and the nextcloud_operator_upgrade_apps_pending{namespace,instance} gauge (see Monitoring). The Warning event fires when the pending set is first seen or changes — not once per retry for the days a store lag can last.
Pin exact versions, or know what a floating one costs¶
spec.version accepts an exact version ("34.0.2"), a bare major or prefix ("34"), or
an alias (latest, stable). All three are supported. What changed in 0.23.1 is that
the rendered container tag is always the resolved version — so "34" produces
image.tag: 34.0.2, not a floating 34.
That matters even when you are not upgrading. A floating tag means the running version is whatever the registry served when each pod started: two pods of the same instance, started either side of an upstream release, run different Nextcloud versions, and a node reschedule is enough to cause it. Pinning the tag to the resolved version makes an instance reproducible.
Recommendation: set an exact version for anything you care about reproducing, and use
a prefix or alias deliberately, understanding that the resolved version — and therefore
the instance — moves when the NextcloudVersionMap moves. That movement is the
upgradePolicy drift path, not an accident.
Migration note (0.23.1)
An instance whose spec.version is a bare major or alias currently runs a
floating tag. On its first reconcile under 0.23.1 the tag re-renders to the pinned,
resolved version — a normal rolling restart of that instance's pods, even though
no version changed. One restart per affected instance, once.
An instance with an exact spec.version, or with none at all, renders exactly
what it rendered before and does not restart.
If your spec.image.repository points at a mirror, check that the pinned tag
exists there before rolling the fleet. A floating 34 may have been the only tag
your mirror published; 34.0.2 has to be there too. spec.image.tag remains the
escape hatch and still overrides everything.
When the running version is not the version you asked for¶
The upgrade flow waits for the pod to report a version at or beyond the target, with
migrations finished. If what actually comes up is newer than the target — a mirror that
resolved differently, a hand-pinned spec.image.tag, an upstream release published
between the version-map snapshot and the roll — the operator adopts the running version as
the target, emits a Normal UpgradeTargetDrifted event naming both, and carries on:
It does not accept a pod more than one major beyond the target — that is not drift, it is a mistake, and converging apps on top of it would bury it. Nor does it accept a version it cannot parse (a nightly or custom build), because there is no ordering to compare.
A pod older than the target is simply still rolling, and the flow keeps waiting. Since
0.23.1 that wait is announced at INFO when it starts and whenever the observation changes,
so a stalled roll is visible in the logs immediately rather than at the two-hour
UpgradeStuck bound.
Resync: when status and reality disagree¶
Re-derives status.upgrade from what the instance actually reports (its version, whether
migrations are pending, and whether maintenance mode is on). Use it when the operator's
bookkeeping has come adrift from the instance — the case it was built for is an upgrade
that really completed while status.upgrade still describes a roll in progress.
| What it finds | What it does |
|---|---|
| The instance is at or beyond the window's target, migrations done | Adopts the running version as the target and hands the window back to app convergence (UpgradeResynced). It does not declare the upgrade complete — the apps have not been verified, and that check is what finishing an upgrade means. |
| The instance is behind the target, or still migrating | Nothing. The upgrade genuinely has not landed (UpgradeResyncInconclusive). |
| An active window describing nothing | Clears the window (UpgradeResynced). A completed window is left alone — it is the record of an upgrade that really happened. |
| Maintenance mode on, and one of the operator's OWN windows is recorded as holding it | Lifts it (UpgradeResynced). |
| Maintenance mode on, with no such record | Leaves it alone and reports MaintenanceModeNotOurs. Clear it yourself with occ maintenance:mode --off if it is stale. |
| Nothing wrong | Says so (UpgradeResyncNoop) rather than reporting a fix that did not happen. |
If the instance cannot be read at all — no pod, occ unreachable — the resync changes
nothing and reports UpgradeResyncFailed. Cleaning up real state on the strength of
not knowing would be worse than the state being wrong.
Adopted state converges on the next timer tick (within ~30s), not inside the annotation.
The operator spends the annotation itself once it has acted — it is a request, not a mode, and leaving it set would otherwise re-run this every 30 seconds. You do not need to remove it. (It is deliberately not cleared when the instance could not be read at all: that outcome should retry.)
MaintenanceModeOrphaned: closed, with nothing to explain it¶
The maintenance timer checks whether an instance is in maintenance mode while no upgrade is in flight, and reports it:
kubectl get nci my-nextcloud -o jsonpath='{.status.conditions[?(@.type=="MaintenanceModeOrphaned")]}' | jq
It never lifts the flag by itself, and that asymmetry is deliberate: you may have run
occ maintenance:mode --on yourself to work on the instance, and an operator that undoes
that mid-task is worse than one that leaves a stale flag.
Nor will the resync annotation lift this particular flag. The operator only lifts
maintenance mode when one of its own upgrade windows is recorded as holding it
(status.upgrade.maintenanceHeld), and by the time this condition can fire, no such
window exists — an upgrade that completed had already lifted its own flag, and one that
could not lift it is still recorded as in progress. So a flag reported here is, by
construction, not one the operator set.
The recovery is manual, and that is the honest answer rather than a limitation to work around:
apiVersion: k8s.bnerd.com/v1alpha1
kind: NextcloudCommand
metadata:
name: clear-stale-maintenance
namespace: my-namespace
spec:
targetRef: {kind: NextcloudInstance, name: my-nextcloud}
commands: [["maintenance:mode", "--off"]]
Use a NextcloudCommand rather than kubectl exec so it queues through the same
per-instance lock as everything else the operator runs.
The condition is not raised while an upgrade is in flight, because the flag is correct for most of every upgrade.
If the flow stalls: UpgradeStuck¶
A phase that has not progressed for two hours (UPGRADE_STUCK_TIMEOUT) raises status.conditions[type=UpgradeStuck] and one Warning event per onset. This covers MaintenanceOn, Rolling, Converging — and AppsPending when the window is still held, because that is the one case where users stay locked out indefinitely. A long-lived AppsPending with maintenance lifted is degraded, not stuck, and stays quiet.
Instances with the app store disabled¶
An instance with appstoreenabled=false (or no egress to the store) cannot update or install store apps at all. The convergence loop behaves exactly as it does for store lag: the steps fail, the failures are recorded with their exit codes, and the instance lands in AppsPending — held or lifted by the critical-set rule. There is no special-casing. Either re-enable the store, install the app by another route, or accept the state with the annotation below.
Escape hatch: accept the current state¶
kubectl annotate nci my-nextcloud \
k8s.bnerd.com/upgrade-apps-accept=true --overwrite -n my-namespace
Lifts the maintenance window, stops the convergence retry loop, and records
status.upgrade.accepted: true.
AppsHealthy deliberately stays False. Accepting means "stop holding this
instance", never "this is healthy" — the degraded app state stays visible until it
actually changes.
Use it when you have decided to run without the affected app for now. The alternative is
to stop declaring it: remove it from spec.apps (or set enabled: false), which is the
honest fix if you are not coming back to it.
It works in every phase — including before anything was checked¶
The annotation is honoured in every phase of an in-flight upgrade, not only
AppsPending: a convergence pass that fails repeatedly holds the window open too, and
that state needs an exit as much as a store-lagged one does.
That means "accepted" covers three different situations, and the UpgradeInProgress
reason tells you which one you are looking at:
| Reason | What it means |
|---|---|
AcceptedWithPendingApps |
A pass verified the apps; these ones are still disabled and you accepted that. The message names them. |
AcceptedAfterVerification |
A pass verified every declared app as enabled; the annotation only lifted the window (typically because the automatic lift kept failing). Nothing is known to be wrong. |
AcceptedBeforeConvergence |
You accepted while the upgrade was still rolling, so the declared apps were never checked for this upgrade and the operator will not re-assert them. status.upgrade.pendingApps is empty here because nothing looked — not because nothing is wrong. Check the instance's apps yourself. |
status.upgrade.verifiedAt is the underlying signal: absent means no pass ever read the
live app state for this upgrade.
Remove it when you are done with it¶
The annotation applies to the upgrade that was in flight when you applied it. An
annotation that is already on the CR when a later upgrade opens its window is treated
as left over and ignored, with a StaleUpgradeAcceptIgnored warning event — otherwise a
forgotten annotation would abort the next upgrade's orchestration before it converged,
lifting the maintenance window mid-roll and leaving the apps unverified.
So a left-behind annotation is no longer dangerous, but it is not harmless either: while it stays on the CR you cannot accept a later upgrade with it (the operator has no way to tell a fresh decision from the old one). Remove it once the situation is resolved; re-apply it if and when you want to accept a different upgrade.
Turning maintenance mode off by hand does not stop the operator
While an upgrade is in flight and the window is held, every convergence pass
re-asserts occ maintenance:mode --on. So an occ maintenance:mode --off you run
yourself is undone within one retry interval (~15 minutes) — the operator is not
fighting you deliberately, it simply has no way to distinguish your deliberate
override from the flag being lost to a pod restart. Use the annotation above
instead: it is the supported way to say "stop holding this instance", and it
records that decision in status where the next person can see it.
Who can do this — and the condition under which that changes¶
The annotation is an ordinary CR metadata write, so the authorisation boundary is
Kubernetes RBAC: it needs patch on nextcloudinstances in the instance's namespace,
and nothing weaker. That makes it a platform-operator action in any deployment where
tenants do not hold credentials against the cluster API.
Note what the annotation actually overrides, though: the critical-app hold is a protection applied to the tenant — it keeps an instance closed when its only login path is dead. So if a control-plane surface (a self-service portal, a support tool, any automation that writes CR annotations on a tenant's behalf) ever proxies this write, that protection becomes defeasible by whoever can reach that surface, and the RBAC boundary above stops being the real boundary. If you are building such a surface, treat proxying this particular annotation as a privilege decision, not a convenience feature — the failure mode it unlocks is "instance reachable, nobody can log in", presented as if it were healthy.
Advisory pre-flight (AppCompatWarning)¶
Before the roll, the operator asks the app store whether each declared app has a release for the target major, and emits a Warning event listing any that do not. It never blocks the upgrade: the store is a third party that can be slow, down, or firewalled off, and none of that is a reason to refuse a version bump you asked for.
AppCompatWarning— checked, and these apps have no compatible release published.AppCompatUnknown— the store could not be consulted at all. Deliberately a different event: "could not check" must never read as a clean bill of health.
Note the imprecision: the store's platform listing includes apps that have a release for that platform, so an app that ships bundled with the server (not published in the store) is indistinguishable from one with no compatible release, and will appear in the warning. That is acceptable for an advisory signal; it is why this is not a gate.
Disable it entirely with UPGRADE_APP_COMPAT_PREFLIGHT=false (air-gapped landscapes), or point it at a mirror with APP_STORE_URL.
Runbook: apps still disabled after an upgrade¶
- See what is pending and why.
kubectl get nci $NAME -n $NS -o jsonpath='{.status.upgrade.pendingApps}' | jq
kubectl get nci $NAME -n $NS -o jsonpath='{.status.upgrade.steps}' | jq '.[] | select(.success == false)'
kubectl get nci $NAME -n $NS -o jsonpath='{.status.upgrade.maintenanceHeld}'
The failed step's message carries occ's own stderr — an incompatible version reads differently from a store that could not be reached.
-
Decide whether to wait. If the app simply has no compatible release yet, waiting is the correct action: the loop is already retrying and will heal the instance without you. Check the store yourself to see whether a release exists.
-
If users are locked out (
maintenanceHeld: true, i.e.user_oidcis pending), your choice is between waiting with the instance dark, and accepting the state to bring it back up with SSO broken. If SSO is the only login path, "up" may be worse than "in maintenance" — that is the judgement the operator deliberately does not make for you beyond the default hold. -
If the app is gone for good, stop declaring it (
spec.apps.<app>.enabled: false, or removespec.oidcif you are abandoning SSO). The loop converges as soon as the declared set matches reality. -
Never hand-run
occ app:enableviakubectl execto "fix" it — that bypasses the per-instance occ lock and races the loop. Use aNextcloudCommandif you need a manual step; see Running occ Commands.
Patch and minor upgrades¶
Patch (31.0.8 → 31.0.9) and minor (31.0.x → 31.1.x) upgrades within the same major follow the same procedure and carry the same risk profile as any Nextcloud point release:
- Confirm a green backup exists (see Backup & rollback below) — cheap insurance, even for a patch bump.
- Edit
spec.versionto the target (exact version, or bump a prefix pin like"31"to pick up the newest 31.x automatically on the next reconcile). - Watch
status.phasegoReady → Deploying → Ready. - Confirm post-upgrade tasks ran:
status.maintenance.lastRunTriggershould showpost-upgradeandstatus.maintenance.lastVersionshould match the new resolved version.
For fleets that don't want to touch every instance by hand, see Fleet strategy below.
Major upgrades¶
Nextcloud only supports upgrading one major version at a time, and the operator enforces this — see Upgrade-path validation. Treat every major hop as its own change, not a multi-version jump.
Checklist per major hop (e.g. going from 30 → 32 means running this twice: 30 → 31, then 31 → 32):
- Back up first. Trigger (or confirm the schedule already ran) an S3 data backup via
spec.backups.data, and apgBackRestbackup if the database is managed (spec.database.managed: true). IfrestoreTestis configured, confirm the most recent test run is green before proceeding — do not upgrade on the strength of an unverified backup. - Validate on staging first if you have a staging instance for this tenant/profile (see Staging / dry-run workaround) — restore the backup there and run the hop before touching production.
- Set
spec.versionto the next major (e.g."31"or an exact"31.0.9"), not the final target. The operator's validation will reject a skip (e.g.30 → 32in one step) with a permanent error. - Wait for
status.phase: Ready. Do not queue the next major change while the current hop is stillDeploying— let the HelmRelease finish rolling and the pod come back healthy first. - Wait for post-upgrade tasks to complete for this hop (
status.maintenance.lastVersionmatches,lastRunTrigger: post-upgrade) before starting the next hop. Runningocc maintenance:repairanddb:add-missing-indicesagainst a schema that's about to be migrated again is wasted work and adds risk. - Spot-check the app — log in, check a few
NextcloudCommandocc status/occ app:listcalls (see occ during an upgrade) — before starting the next hop. - Repeat from step 3 for each remaining major, one at a time.
Do not try to shortcut this by setting spec.version directly to a version two majors ahead "to save a step" — the operator will reject it (see below), and even if it didn't, Nextcloud's own migrations aren't designed to skip majors.
Upgrade-path validation¶
Every reconcile that changes the resolved Nextcloud version is checked against status.versionResolution.resolvedVersion (the currently-running version) before the HelmRelease is touched. Downgrades are always unsafe, and so is skipping more than one major in a step (e.g. current 30.x, target 32.x) — but what happens next depends on how the target version was arrived at:
- An explicit choice fails loud. If the unsafe target comes from something typed directly — an exact
spec.version, or thespec.helm.versionescape hatch (status.versionResolution.resolvedByiscrd,configmap, orspec.helm.version) — the reconcile fails with a permanent error andstatus.phasereportsFailed. Correctspec.version(step through the intermediate major, or fix the typo) to unblock it. - Resolution drift holds instead of failing. If the unsafe target instead comes from the operator re-resolving an alias or prefix that was already set (
resolvedBy: aliasorresolvedBy: prefix— e.g.latestgets repointed, or an edit to theNextcloudVersionMapchanges what"31"resolves to), the instance does not goFailed. The operator holds the instance on its current chart version, logs a warning, and continues reconciling everything else normally — no status condition, no event, just a warning-level log line. This is the ops scenario the hold behavior exists for — see What happens on a version-map rollback below. - Patch and minor changes within the same major, and same-major re-resolutions (e.g.
latestmoving from32.0.5to32.0.6), are unaffected either way — nothing is held or blocked. - A target that's a bare prefix of the currently-running version (e.g. target
"32"whilestatus.versionResolution.resolvedVersionis32.0.6— this happens with thespec.helm.versionescape hatch) counts as the same version, never a downgrade. - Instances with no prior resolved version (first install) or a non-dotted-integer version string (custom/nightly builds) are never blocked — validation only applies once there's a real "current" version to compare against.
What happens on a version-map rollback¶
This is the scenario the hold behavior exists for: an ops team ships a bad Nextcloud release, then rolls the NextcloudVersionMap/default CRD back to remove or repoint that entry. Every instance tracking the affected alias or prefix (spec.version: latest, or a prefix pin like "31") re-resolves on its next reconcile — and the new target is now older than what's already deployed, or in the worst case points at a different major entirely.
Without the hold behavior, that one CRD rollback would flip status.phase: Failed across every instance tracking it, fleet-wide, in a single tick — a self-inflicted outage caused by the fix, not the original bad release. Instead, each affected instance simply keeps running its current chart version and logs a warning (grep operator logs for the instance name, or for "upgrade-path"). Nothing else about reconciliation is blocked — secrets, ingress, database, and every other field keep syncing normally in the meantime. Once the version map is corrected (or you explicitly edit that instance's spec.version yourself, which always takes the hard-fail path instead), the hold clears on the next reconcile.
The escape hatch¶
Annotation: k8s.bnerd.com/allow-unsafe-version-change: "true"
Set this on the NextcloudInstance to bypass both the downgrade check and the major-skip check for as long as the annotation is present — which is why removing it afterwards is part of the procedure. This exists only for documented restore/rollback procedures (see Backup & rollback) — never for routine upgrades. Using it to skip majors in one hop leaves Nextcloud's own migration path unverified; using it to "downgrade" without a genuine restore underneath will corrupt state.
kubectl annotate nci my-nextcloud \
k8s.bnerd.com/allow-unsafe-version-change=true \
--overwrite -n my-namespace
Remove it once the change lands — see the numbered rollback procedure below for the full sequence, including when to remove it.
occ during an upgrade¶
There is no operator-level lock that pauses occ activity while a HelmRelease version change rolls out. What guards exist are incidental, not a coordination mechanism:
NextcloudCommand(see Running occ Commands) only dispatches once the targetNextcloudInstanceisphase: Readywith aRunningandReadypod — so a command submitted mid-upgrade waits, it isn't cancelled or corrupted.- The operator's own maintenance tasks require
helmRelease.readybefore running. - Nextcloud's own maintenance mode (set by the image entrypoint during
occ upgrade) blocks mostoccsubcommands for the duration of the migration.
None of this is a substitute for good sequencing. In-flight commands are not cancelled when an upgrade starts, and a manual kubectl exec ... occ ... bypasses every one of these guards — it will happily race a migration.
Recommendation: run all occ invocations through NextcloudCommand, never kubectl exec, and don't submit new commands while a version change is rolling out (status.phase: Deploying). Wait for Ready first.
Fleet strategy¶
For more than a handful of instances, pick a mix of these rather than editing every spec.version by hand.
Automatic upgrades: spec.upgradePolicy¶
The primary automation surface, as of 0.20.0. A sibling of maintenance, not nested under it:
spec:
version: "32" # prefix pin — the operator resolves it, upgradePolicy decides what happens on drift
upgradePolicy:
mode: auto # manual (default) | auto
patchUpgrades: true # auto-apply same-minor patch drift, e.g. 32.0.1 -> 32.0.2
minorUpgrades: false # auto-apply same-major minor drift, e.g. 32.0.x -> 32.1.0
maintenance:
windowStart: 2 # still governs *when* auto-applied upgrades happen
mode: manual(default) — no automatic upgrades; aspec.versionedit is the only way to move version.UpdateAvailable(below) is still computed and surfaced regardless of mode — this is the fleet-visibility use case for staying in manual mode deliberately.mode: auto— inside the maintenance window, the operator re-resolvesspec.versionand applies drift per the flags independently: a resolved patch bump needspatchUpgrades: true; a resolved minor bump (same major, different minor) needsminorUpgrades: true. There is nomajorUpgradesfield, deliberately — no flag combination automates a major hop; it always requires an explicitspec.versionedit, same as today, still fully subject to upgrade-path validation.spec.maintenance.windowStartstill governs when an auto-applied upgrade happens — set it alongsideupgradePolicy.mode: auto, exactly as it did forautoUpdatebefore.- Per-profile default:
NextcloudProfile.spec.defaults.upgradePolicysets a fleet-wide default for every instance using the profile, with the same deep-merge cascade semantics as every other profile default (see Profiles → How Defaults Merge) — a profile can setmode: auto, patchUpgrades: trueand an instance that only overridesminorUpgrades: truestill inheritsmode/patchUpgradesfrom the profile; instance values win on conflict.
UpdateAvailable status condition¶
status.conditions[type=UpdateAvailable] reports whether a newer version currently resolves for this instance — independent of upgradePolicy.mode and independent of whether spec.maintenance.windowStart is even set. The window gates upgrade application, never visibility: a mode: manual fleet, or any instance with no maintenance window configured at all, still gets this signal.
# Block until an update is confirmed available
kubectl wait nci my-nextcloud --for=condition=UpdateAvailable=True --timeout=600s
# Or just read it
kubectl get nci my-nextcloud -o jsonpath='{.status.conditions[?(@.type=="UpdateAvailable")]}' | jq
status is set explicitly to "False" once up to date — not left absent — so kubectl wait works whichever way it resolves. reason is NewerVersionResolved (True) or UpToDate (False); the message names the resolved target version and chart. The condition isn't computed at all (left absent, not False) until the instance has completed its first deploy (status.versionResolution exists) — there's nothing to compare against yet.
On-demand upgrades: k8s.bnerd.com/upgrade-now¶
Applies the currently-resolved update immediately, bypassing both the maintenance window and upgradePolicy.mode entirely — the point of this annotation is that a mode: manual fleet (or an instance outside its window) can still apply a known-good, already-resolved update on request, without waiting for either.
kubectl annotate nci my-nextcloud \
k8s.bnerd.com/upgrade-now=$(date -u +"%Y-%m-%dT%H:%M:%SZ") \
--overwrite -n my-namespace
- Applies on any drift class — patch, minor, or a resolvedVersion-only/chart-only bump.
upgradePolicy.patchUpgrades/minorUpgradesare never consulted here; asking explicitly means getting whatever's pending, not a policy-narrowed subset. - Still fully protected by upgrade-path validation — a major-skip or downgrade target is held or fails exactly as it would through any other reconcile trigger;
k8s.bnerd.com/allow-unsafe-version-changeis still required to bypass that. This annotation is an on-demand trigger, not a validation bypass. - No-ops with an event (
reason: NoUpdateAvailable) when nothing is pending — checkUpdateAvailablefirst if you're not sure there's anything to apply. - Refuses an unparseable resolved target (e.g. a corrupted
NextcloudVersionMapentry) rather than guessing at its class — the same guard themode: autoapply path uses. - Firing this at the same moment a
mode: autoinstance's maintenance-window tick would also apply an update collapses to a single upgrade, not a double-apply — safe to use even mid-window on anauto-mode instance.
The maintenance.autoUpdate alias (deprecated)¶
spec.maintenance.autoUpdate: true still works. It's a deprecated alias for upgradePolicy: {mode: auto, patchUpgrades: true, minorUpgrades: false}, applied automatically — no spec edit required for existing instances. If spec.upgradePolicy is also set on the same instance, upgradePolicy wins outright and autoUpdate is ignored entirely; a reason: UpgradePolicyAliasIgnored warning event fires (once per reconcile that has pending drift, not on every tick) so the conflict doesn't go unnoticed.
The alias is not removed in 0.20.0 and won't be removed before 0.21.0 — existing autoUpdate: true specs keep working unchanged. But its exact behavior changed underneath it in 0.20.0; read both notes below before assuming it behaves identically to before.
Upgrade notes: two separate, opposite-direction changes in 0.20.0¶
Don't conflate these — they affect different things and point in opposite directions.
1. Narrowing — the alias now only auto-applies patches, not minors. Before 0.20.0, autoUpdate: true applied any same-major drift, patch or minor, without distinguishing between them. As of 0.20.0 the alias maps to patchUpgrades: true, minorUpgrades: false specifically — a resolved minor bump (e.g. 32.0.x → 32.1.0) no longer auto-applies under the alias alone. If you were relying on minor drift auto-applying, add spec.upgradePolicy: {mode: auto, patchUpgrades: true, minorUpgrades: true} explicitly (this also moves you off the alias entirely, onto upgradePolicy, so the ignored-warning never fires for you).
2. Widening — drift detection itself now also catches resolvedVersion-only bumps. Before 0.20.0, drift detection compared only the resolved chart version. Several Nextcloud patch releases can map to the same chart version in NextcloudVersionMap (e.g. Nextcloud 32.0.6/32.0.8/32.0.9 all resolving to chart 8.9.1) — that class of drift was previously invisible to autoUpdate entirely; it silently never applied, with no signal that anything was pending. As of 0.20.0, drift detection compares both the chart version and the resolved Nextcloud version (the same basis UpdateAvailable uses). This class of drift is now detected — and, with patchUpgrades: true (which the alias sets), auto-applied — on the very next maintenance-window tick after upgrading to 0.20.0, for any instance already inside its window with autoUpdate/upgradePolicy set. This is a real, immediate behavior change for existing fleets on their next window tick, not just a change for newly-created instances.
Other fleet patterns¶
- Exact pin +
mode: manual.spec.version: "31.0.9"with noupgradePolicy(ormode: manual, or the equivalent unset/falseautoUpdate) is a true no-op fleet-wide — nothing moves until you edit the spec. Use for instances under strict change control.UpdateAvailablestill tells you when something's pending. - Profile-level version pinning.
NextcloudProfile.spec.defaults.versionpins every instance using that profile at once — useful for rolling a cohort of tenants to a known-good version without touching eachNextcloudInstance. An instance's ownspec.versionalways overrides the profile default. See Version Management → Profile-Level Version Pinning. - Mixed versions across the fleet are normal and supported. Different tenants/profiles can sit on different majors simultaneously — there is no fleet-wide version requirement. Stagger major hops by profile or by tenant tier (e.g. staging profile a major ahead of production) to get a soak period for free.
Recommended production default: prefix pin ("31") + upgradePolicy: {mode: auto, patchUpgrades: true} + a windowStart for automatic patch upgrades, with explicit spec.version edits reserved for major hops, following the Major upgrades checklist. Use k8s.bnerd.com/upgrade-now when you need a specific instance upgraded before its next window tick.
Backup & rollback¶
Nextcloud has no in-place downgrade. "Rollback" always means restoring pre-upgrade backups, not reverting spec.version on a live, already-migrated instance.
Before any upgrade:
- Confirm
spec.backups.datais configured (S3 data backup via the bnerd backup operator) and, ifrestoreTestis set, that the last automated restore test succeeded. - If the database is managed (
spec.database.managed: true), confirm scheduledpgBackRestbackups are enabled and current (spec.database.postgres.backup) — see Managed PostgreSQL. - Note the pre-upgrade
status.versionResolution(requested/resolved/chart versions) so you know exactly what "rollback" means for this instance.
If an upgrade needs to be rolled back:
- Stop traffic to the instance if it's serving users (scale down, maintenance page at the ingress level, or similar — outside the operator's scope).
- Restore the pre-upgrade S3 data backup and the pre-upgrade database backup (
pgBackRestrestore for managed Postgres, or your external DB's own restore procedure) to the point captured before the failed upgrade. - Set the annotation
k8s.bnerd.com/allow-unsafe-version-change: "true"on theNextcloudInstance— required because settingspec.versionback to the pre-upgrade value is, from the operator's point of view, a downgrade. - Set
spec.versionback to the pre-upgrade value (theresolvedVersionyou noted above, not a prefix or alias — be exact so the resolved version matches what you just restored). - Wait for
status.phase: Readyand spot-check the app. - Remove the
k8s.bnerd.com/allow-unsafe-version-changeannotation once the instance is stable. Leaving it set disables downgrade/major-skip protection for every future reconcile, not just this one.
There is no automated rollback — this is a manual procedure end to end. Practice it against a staging instance before you need it in production.
Staging / dry-run workaround¶
There is no built-in dry-run or clone-for-upgrade-testing feature. Until one exists, approximate it manually:
- Create a staging
NextcloudInstanceusing the same profile as the production instance you're planning to upgrade. - Restore a recent production backup into the staging instance (same restore procedure as Backup & rollback, steps 2–3, minus the annotation since staging starts from the pre-upgrade version already).
- Run the upgrade (or the next major hop) against staging first, following the Major upgrades checklist.
- Validate via
NextcloudCommand—occ status,occ app:list --output=json, and any app-specific sanity checks — before repeating the change against production. - Tear down or reuse the staging instance for the next hop.
App compatibility¶
spec.apps installs run as occ app:install <app> || true in the chart's before-starting hook — best-effort only. There is no compatibility check against the target Nextcloud version, and a failed install does not surface in status; it silently no-ops. Before a major upgrade, check each installed app's compatibility with the target Nextcloud version against the Nextcloud app store yourself, and validate via a staging run (above) rather than trusting spec.apps to fail loudly if an app breaks.
See also¶
- Standard Upgrade Procedure — the single consolidated checklist built from this guide.
- Version Management — how
spec.versionresolves to a chart version, theNextcloudVersionMapCRD, image resolution, profile-level pinning. - Operations & Annotations — the full annotation/label reference, including
k8s.bnerd.com/reconcileandk8s.bnerd.com/run-maintenance. - Running occ Commands — the
NextcloudCommandCRD used for validation and spot-checks above. - Managed PostgreSQL —
pgBackRestbackup configuration fordatabase.managed: true.