Skip to content

Upgrades

Upgrading Nextcloud is a version bump on the CR, but doing it safely across a fleet is a platform-team procedure: back up first, step through majors one at a time, and validate before moving on. This guide is the production runbook — the mechanism behind every step. For a single top-to-bottom checklist that assembles this guide with Operations, occ Commands, and Version Management into one procedure, see Standard Upgrade Procedure. For how spec.version resolves to a chart version in the first place, see Version Management.

What the operator automates vs. what you do

Operator You
Resolve spec.version to a chart + image ✅ via NextcloudVersionMap —
Roll the Helm release ✅ Flux applies the new HelmRelease —
Hold maintenance mode across the WHOLE window ✅ 0.23.0 — opened before the roll, lifted only after apps reconverge —
Run occ upgrade / DB migration ✅ official image entrypoint, on pod start, inside the operator's window —
Re-enable apps the upgrade disabled ✅ 0.23.0 — app:update --all + re-assert, retried until it converges —
Run post-upgrade occ repair tasks ✅ in-window (fast repair) + maintenance timer (expensive repair) —
Block unsafe major-skips and downgrades ✅ upgrade-path validation —
Decide when to change spec.version — ✅ you
Pre-upgrade backup + restore test — ✅ you
Step through intermediate majors — ✅ you
App compatibility check ⚠️ 0.23.0 — advisory pre-flight warning only, never a gate ✅ you (still your call, see below)
Rollback — ✅ you (restore from backup)

The operator still does not exec occ upgrade itself — it never touches the pod for the migration. It changes the HelmRelease; the Nextcloud container image's own entrypoint detects the version bump and runs the DB migration before serving traffic again. What changed in 0.23.0 is everything around that step: the operator now opens Nextcloud's maintenance mode before the roll and holds it until the instance's declared apps are verified enabled again — see Post-upgrade app convergence. The image's own maintenance toggle only ever spanned the schema migration, and Nextcloud's updater leaves a maintenance flag it did not set alone, so the operator's outer window survives occ upgrade intact.

A failed upgrade holds — it does not roll back (0.21.0+)

As of 0.21.0, a Helm upgrade that fails partway (spec.timeout elapses, or the release comes back unhealthy) is held on the failed revision, not automatically rolled back. This is a deliberate change from earlier versions, and it matters because Nextcloud's own DB migration is not reversible: if the failed release already ran occ upgrade far enough to migrate the schema forward before crashing, rolling the Helm release back to the previous chart/image is not a rollback at all — it re-deploys the old Nextcloud code against the new, already-migrated schema, which crash-loops permanently. There is no way for Flux or the operator to know from the outside whether a given failure happened before or after the migration point, so as of 0.21.0 the operator never lets Flux guess: it disables Flux's automatic upgrade remediation entirely (spec.upgrade.remediation.retries: 0 on the generated HelmRelease) and instead surfaces the stuck state for a human to diagnose. Install remediation is unchanged (retries: 3) — a failed first install has no migrated data at risk, so self-healing there is still safe.

Two related changes ship alongside this:

  • Default Helm timeout raised to 15 minutes. Flux's own default is 5 minutes, which large-migration upgrades (a big Nextcloud instance moving a major version) routinely exceed — a slow-but-otherwise-healthy migration hitting the old 5m ceiling looked identical to a genuinely broken upgrade. Override per instance with spec.helm.timeout (a Go duration string, e.g. "30m", "1h") if even 15 minutes isn't enough for a particular instance's data size.
spec:
  helm:
    timeout: "30m"
  • UpgradeFailed status condition. The operator's status-sync timer compares the HelmRelease's lastAttemptedRevision against lastAppliedRevision: whenever the release is not ready and the two differ, it sets status.conditions[type=UpgradeFailed] to True with reason: HelmUpgradeFailed, and the message carries both revisions plus Flux's own failure message. It also emits a HelmUpgradeStuck Warning event — once per stuck revision, not once per 30-second reconcile tick, so a human investigating isn't flooded with duplicate events for the same still-unresolved failure. Once the revisions converge again (the upgrade is fixed and completes, or a corrected upgrade-now re-applies successfully), the condition flips to False with reason: Recovered. An instance that was never stuck gets no UpgradeFailed write at all — kubectl describe stays clean for healthy instances, it doesn't carry a permanent False condition nobody asked for.
kubectl get nci my-nextcloud -o jsonpath='{.status.conditions[?(@.type=="UpgradeFailed")]}' | jq

Where the applied revision comes from depends on your Flux version, and the operator handles both. Flux ≥ 2.3 serves helm.toolkit.fluxcd.io/v2, which dropped status.lastAppliedRevision in favour of status.history — so the operator reads the chart version of the newest history snapshot Helm still reports as deployed (a failed upgrade leaves the previous release deployed and the new one failed, which is exactly the comparison this condition wants). On Flux 2.0–2.2, where a beta API was the storage version, it reads status.lastAppliedRevision directly. An instance whose first install never succeeded has no deployed snapshot at all, so it is never reported as a stuck upgrade.

kubectl get helmrelease my-nextcloud-nextcloud \
  -o jsonpath='{range .status.history[*]}{.chartVersion}{"\t"}{.status}{"\n"}{end}'

Existing HelmReleases converge automatically — the disabled remediation and raised timeout apply the next time the operator writes to an instance's HelmRelease (any on_create/on_update/force-reconcile patch), not via a one-time migration step. A fleet upgrading the operator to 0.21.0 doesn't need to touch every NextcloudInstance by hand; the next normal reconcile (a spec edit, a version bump, or a k8s.bnerd.com/reconcile annotation) picks up the new remediation/timeout settings.

Runbook: a stuck upgrade

  1. Diagnose. Start with the condition itself — it already names both revisions and carries Flux's failure message:
kubectl describe nci my-nextcloud
# look for: Type: UpgradeFailed, Status: True, Reason: HelmUpgradeFailed

Then pull the full Helm history and the Flux controller's own logs for more detail than fits in the condition message:

helm history my-nextcloud-nextcloud -n my-namespace
kubectl logs -n flux-system deploy/helm-controller --since=1h | grep my-nextcloud-nextcloud
  1. Fix the underlying cause. This is almost always outside the operator's control — a chart value that doesn't validate, a resource limit too low for the migration workload, an image pull failure, or the migration itself failing inside the pod (check kubectl logs on the Nextcloud pod for the occ upgrade output). Fix whatever helm history/the pod logs point at.

  2. Let it retry, or re-trigger explicitly. Once the underlying cause is fixed, either wait for Flux's own reconcile interval to retry the same revision, or force it immediately with the on-demand upgrade annotation:

kubectl annotate nci my-nextcloud \
  k8s.bnerd.com/upgrade-now=$(date -u +"%Y-%m-%dT%H:%M:%SZ") \
  --overwrite -n my-namespace

UpgradeFailed flips to False/Recovered once lastAttemptedRevision and lastAppliedRevision converge again.

Never manually roll back the image once DB migrations may have run. Editing spec.version back to the previous value, or otherwise forcing the chart/image backward, is exactly the crash-loop scenario this feature exists to prevent — see the explanation above. If you genuinely need to undo a bad upgrade, that's a data restore, not a version edit: follow Backup & rollback below, which restores from a pre-upgrade backup rather than pointing a migrated DB at old code.

How an upgrade happens

  1. You change spec.version (or a pinning profile's defaults.version changes and the instance re-resolves — see Version Management).
  2. The operator resolves the new value through the NextcloudVersionMap to a Nextcloud version + Helm chart version + image.
  3. Upgrade-path validation runs against the resolved current and target versions (see below). An unsafe transition fails the reconcile with a permanent error before anything is touched.
  4. The operator turns Nextcloud's maintenance mode ON and records status.upgrade (phase: MaintenanceOn, from, to) plus the UpgradeInProgress condition. An advisory app-store pre-flight runs here too (below). (0.23.0)
  5. The operator updates the HelmRelease. status.phase moves Ready → Deploying, status.upgrade.phase moves to Rolling.
  6. Flux reconciles the HelmRelease; the new pod starts and the Nextcloud entrypoint runs occ upgrade. Because maintenance mode was already on, Nextcloud's updater leaves it on when it finishes (it only clears a flag it set itself) — so the instance stays closed to users.
  7. The operator waits for occ status to report the target version, then converges the apps: app:update --all, re-assert every declared app, the in-window post-upgrade tasks, then verify with app:list. (0.23.0 — status.upgrade.phase: Converging)
  8. Maintenance mode is lifted and status.upgrade.phase becomes Completed, with an UpgradeCompleted event — or AppsPending if declared apps are still not enabled (see below).
  9. status.versionResolution records requestedVersion, resolvedVersion, and chartVersion for the completed change.
  10. The expensive post-upgrade task set (maintenance:repair --include-expensive, db:convert-filecache-bigint, …) runs too, unchanged from earlier releases, as a separate pass from step 7's fast in-window repair. Which trigger you used decides when: a spec.version edit runs it from the update handler's own tail, right after step 5 and therefore inside the maintenance window against the still-running old pod; an auto-update, versionmap drift, upgrade-now or a reconcile annotation has no such tail. Either way the maintenance timer picks it up on its next tick once status.maintenance.lastVersion differs — for every instance, with or without a spec.maintenance.windowStart, and it is skipped while an upgrade is still in flight. See the note below on seeing maintenance:repair twice.
kubectl get nci my-nextcloud -o jsonpath='{.status.phase}'
kubectl get nci my-nextcloud -o jsonpath='{.status.upgrade}' | jq
kubectl get nci my-nextcloud -o jsonpath='{.status.versionResolution}' | jq
kubectl get nci my-nextcloud -o jsonpath='{.status.maintenance}' | jq

Post-upgrade app convergence (0.23.0)

The failure this closes

Upgrading Nextcloud across a major disables every app that has no release compatible with the new version. That is the image entrypoint doing the right thing — an incompatible app would break the instance. It then tries to update those apps from the app store and re-enable them, which usually works, because the store usually already serves a compatible release by the time you upgrade.

When it does not, nothing ever tries again:

  • the lifecycle hooks that install and enable apps run only on pod start;
  • occ app:install on an already-installed app reports "already installed" and does not update it;
  • occ app:enable refuses the installed, incompatible version;
  • every one of those hook commands ends in || true, so all of it is silent.

The app stays disabled until a human notices. For user_oidc — an instance whose only login path is SSO — "until a human notices" is an authentication outage.

Two states reach it: the store has not published a compatible release yet (normal in the days after a major Nextcloud release), or the store is unreachable from the cluster at all (egress policy, DNS, an outage).

What the operator does instead

The recovery is a convergence loop owned by the operator's timer, not a hook firing once at the wrong moment. Every version change goes through it — a manual spec.version edit, a version-map alias moving, an upgradePolicy auto-update, or k8s.bnerd.com/upgrade-now.

Phase (status.upgrade.phase) What is happening
MaintenanceOn Maintenance mode enabled; the HelmRelease has not been patched yet.
Rolling HelmRelease patched. Waiting for the pod to report the target version via occ status — Flux reporting Ready is not enough, because converging against the old pod would prove nothing.
Converging occ app:update --all → re-assert each declared app (app:install, then app:enable) → db:add-missing-indices + a fast maintenance:repair → verify with occ app:list.
AppsPending Declared apps are still not enabled. The loop retries; see below.
Completed Every declared app verified enabled, maintenance lifted.

app:update --all runs first on every pass, and that ordering is the entire fix: unlike app:install, it updates apps that are currently disabled — which is exactly the state the upgrade left them in. Re-asserting before updating would just fail against the same incompatible version again.

The declared set is spec.apps (every entry not explicitly enabled: false) plus user_oidc whenever spec.oidc.enabled is set. user_oidc is reserved in spec.apps because it is managed through spec.oidc, which is precisely why the older AppsHealthy check never saw it — and it is the app whose disablement caused the outage this feature exists for.

Each step's result is recorded in status.upgrade.steps[] with its exit code. Nothing is swallowed.

maintenance:repair runs twice per upgrade — and only one of them is in the window

The in-window pass above runs the fast maintenance:repair, so an instance is not held closed for a repair that can take hours. The pre-existing maintenance:repair --include-expensive set still runs as well, but since 0.23.1 it never runs while an upgrade is in flight: the maintenance timer picks it up once the upgrade completes, at a per-instance jittered time.

Earlier releases ran it inline from the update handler the moment spec.version changed — inside the maintenance window, against the pod Flux was in the middle of replacing. That is not merely wasteful: the exec is killed when the pod goes away (exit 143), and if the upgrade then stalls, nothing retries it. 0.23.0's docs called the double-run "expected, not a bug", which was true of the double-run and wrong about the kill. Fixed in 0.23.1.

Telling the two apart: the fast in-window pass is recorded in status.upgrade.steps[] as maintenanceRepair; the expensive pass is recorded in status.maintenance.tasks as repair with lastRunTrigger: post-upgrade. Both honour spec.maintenance.tasks.repairAfterUpgrade: false.

kubectl get nci my-nextcloud -o jsonpath='{.status.upgrade}' | jq
kubectl get nci my-nextcloud -o jsonpath='{.status.upgrade.steps}' | jq '.[] | select(.success == false)'

The maintenance window, and when it is not lifted

The window opens before the roll and closes only when the apps are verified. If some declared app is still disabled at the end of a pass, what happens next depends on which app:

  • A non-critical app (calendar, contacts, groupfolders, …): maintenance is lifted. Availability beats completeness — a missing calendar app must not keep an entire instance dark. The instance serves; the affected apps are unavailable; AppsHealthy reports PostUpgradeAppsDisabled.
  • A critical app: maintenance stays on. The critical set is {user_oidc} when spec.oidc.enabled is set, and empty otherwise. An instance whose only login path is dead is better presented as "in maintenance", with a clear page, than as "up" and unusable to everyone who tries to log in.

An instance with OIDC disabled — including every pool spare — therefore has an empty critical set and can never be held dark by this rule.

If the operator cannot reach the pod to lift the window, it does not report Completed: it holds in Converging and retries. "Completed" is terminal, and claiming it on a locked-out instance would be a lie the loop could no longer fix.

The retry loop

While AppsPending, every instance-status tick re-runs the update / re-assert / verify steps (not the migrations — those already ran):

  • no attempt cap. Store lag resolves in days; giving up would restore exactly the terminal state this feature exists to end.
  • at least 15 minutes apart, plus a deterministic per-instance jitter, so a fleet upgraded in lockstep does not poll the store in lockstep. Tune with UPGRADE_RETRY_INTERVAL / UPGRADE_RETRY_JITTER on the operator deployment.
  • it heals by itself. The moment the store serves a compatible release, the next pass's app:update --all pulls it, app:enable succeeds, and the instance goes Completed — no human action, no re-upgrade.
  • a lifted window is never re-closed by a retry. Once users have the instance back, a routine retry does not take it away again.

While an upgrade is in flight, the upgrade flow owns the AppsHealthy condition: the 15-minute maintenance timer's own apps-health check keeps refreshing status.apps (pure observation) but stops writing the condition until the upgrade reaches a terminal phase. Without that, the timer would flip AppsHealthy back to True every 15 minutes on exactly the instance this feature exists for — it reads spec.apps, which excludes user_oidc, so "SSO dead, everything in spec.apps fine" looks healthy to it.

The alertable signals are the AppsHealthy=False/PostUpgradeAppsDisabled condition and the nextcloud_operator_upgrade_apps_pending{namespace,instance} gauge (see Monitoring). The Warning event fires when the pending set is first seen or changes — not once per retry for the days a store lag can last.

Pin exact versions, or know what a floating one costs

spec.version accepts an exact version ("34.0.2"), a bare major or prefix ("34"), or an alias (latest, stable). All three are supported. What changed in 0.23.1 is that the rendered container tag is always the resolved version — so "34" produces image.tag: 34.0.2, not a floating 34.

That matters even when you are not upgrading. A floating tag means the running version is whatever the registry served when each pod started: two pods of the same instance, started either side of an upstream release, run different Nextcloud versions, and a node reschedule is enough to cause it. Pinning the tag to the resolved version makes an instance reproducible.

Recommendation: set an exact version for anything you care about reproducing, and use a prefix or alias deliberately, understanding that the resolved version — and therefore the instance — moves when the NextcloudVersionMap moves. That movement is the upgradePolicy drift path, not an accident.

Migration note (0.23.1)

An instance whose spec.version is a bare major or alias currently runs a floating tag. On its first reconcile under 0.23.1 the tag re-renders to the pinned, resolved version — a normal rolling restart of that instance's pods, even though no version changed. One restart per affected instance, once.

An instance with an exact spec.version, or with none at all, renders exactly what it rendered before and does not restart.

If your spec.image.repository points at a mirror, check that the pinned tag exists there before rolling the fleet. A floating 34 may have been the only tag your mirror published; 34.0.2 has to be there too. spec.image.tag remains the escape hatch and still overrides everything.

When the running version is not the version you asked for

The upgrade flow waits for the pod to report a version at or beyond the target, with migrations finished. If what actually comes up is newer than the target — a mirror that resolved differently, a hand-pinned spec.image.tag, an upstream release published between the version-map snapshot and the roll — the operator adopts the running version as the target, emits a Normal UpgradeTargetDrifted event naming both, and carries on:

Expected 34.0.2 but the instance is running 34.0.3; continuing with 34.0.3 as the target.

It does not accept a pod more than one major beyond the target — that is not drift, it is a mistake, and converging apps on top of it would bury it. Nor does it accept a version it cannot parse (a nightly or custom build), because there is no ordering to compare.

A pod older than the target is simply still rolling, and the flow keeps waiting. Since 0.23.1 that wait is announced at INFO when it starts and whenever the observation changes, so a stalled roll is visible in the logs immediately rather than at the two-hour UpgradeStuck bound.

Resync: when status and reality disagree

kubectl annotate nci my-nextcloud \
  k8s.bnerd.com/upgrade-resync=true --overwrite -n my-namespace

Re-derives status.upgrade from what the instance actually reports (its version, whether migrations are pending, and whether maintenance mode is on). Use it when the operator's bookkeeping has come adrift from the instance — the case it was built for is an upgrade that really completed while status.upgrade still describes a roll in progress.

What it finds What it does
The instance is at or beyond the window's target, migrations done Adopts the running version as the target and hands the window back to app convergence (UpgradeResynced). It does not declare the upgrade complete — the apps have not been verified, and that check is what finishing an upgrade means.
The instance is behind the target, or still migrating Nothing. The upgrade genuinely has not landed (UpgradeResyncInconclusive).
An active window describing nothing Clears the window (UpgradeResynced). A completed window is left alone — it is the record of an upgrade that really happened.
Maintenance mode on, and one of the operator's OWN windows is recorded as holding it Lifts it (UpgradeResynced).
Maintenance mode on, with no such record Leaves it alone and reports MaintenanceModeNotOurs. Clear it yourself with occ maintenance:mode --off if it is stale.
Nothing wrong Says so (UpgradeResyncNoop) rather than reporting a fix that did not happen.

If the instance cannot be read at all — no pod, occ unreachable — the resync changes nothing and reports UpgradeResyncFailed. Cleaning up real state on the strength of not knowing would be worse than the state being wrong.

Adopted state converges on the next timer tick (within ~30s), not inside the annotation.

The operator spends the annotation itself once it has acted — it is a request, not a mode, and leaving it set would otherwise re-run this every 30 seconds. You do not need to remove it. (It is deliberately not cleared when the instance could not be read at all: that outcome should retry.)

MaintenanceModeOrphaned: closed, with nothing to explain it

The maintenance timer checks whether an instance is in maintenance mode while no upgrade is in flight, and reports it:

kubectl get nci my-nextcloud -o jsonpath='{.status.conditions[?(@.type=="MaintenanceModeOrphaned")]}' | jq

It never lifts the flag by itself, and that asymmetry is deliberate: you may have run occ maintenance:mode --on yourself to work on the instance, and an operator that undoes that mid-task is worse than one that leaves a stale flag.

Nor will the resync annotation lift this particular flag. The operator only lifts maintenance mode when one of its own upgrade windows is recorded as holding it (status.upgrade.maintenanceHeld), and by the time this condition can fire, no such window exists — an upgrade that completed had already lifted its own flag, and one that could not lift it is still recorded as in progress. So a flag reported here is, by construction, not one the operator set.

The recovery is manual, and that is the honest answer rather than a limitation to work around:

apiVersion: k8s.bnerd.com/v1alpha1
kind: NextcloudCommand
metadata:
  name: clear-stale-maintenance
  namespace: my-namespace
spec:
  targetRef: {kind: NextcloudInstance, name: my-nextcloud}
  commands: [["maintenance:mode", "--off"]]

Use a NextcloudCommand rather than kubectl exec so it queues through the same per-instance lock as everything else the operator runs.

The condition is not raised while an upgrade is in flight, because the flag is correct for most of every upgrade.

If the flow stalls: UpgradeStuck

A phase that has not progressed for two hours (UPGRADE_STUCK_TIMEOUT) raises status.conditions[type=UpgradeStuck] and one Warning event per onset. This covers MaintenanceOn, Rolling, Converging — and AppsPending when the window is still held, because that is the one case where users stay locked out indefinitely. A long-lived AppsPending with maintenance lifted is degraded, not stuck, and stays quiet.

Instances with the app store disabled

An instance with appstoreenabled=false (or no egress to the store) cannot update or install store apps at all. The convergence loop behaves exactly as it does for store lag: the steps fail, the failures are recorded with their exit codes, and the instance lands in AppsPending — held or lifted by the critical-set rule. There is no special-casing. Either re-enable the store, install the app by another route, or accept the state with the annotation below.

Escape hatch: accept the current state

kubectl annotate nci my-nextcloud \
  k8s.bnerd.com/upgrade-apps-accept=true --overwrite -n my-namespace

Lifts the maintenance window, stops the convergence retry loop, and records status.upgrade.accepted: true.

AppsHealthy deliberately stays False. Accepting means "stop holding this instance", never "this is healthy" — the degraded app state stays visible until it actually changes.

Use it when you have decided to run without the affected app for now. The alternative is to stop declaring it: remove it from spec.apps (or set enabled: false), which is the honest fix if you are not coming back to it.

It works in every phase — including before anything was checked

The annotation is honoured in every phase of an in-flight upgrade, not only AppsPending: a convergence pass that fails repeatedly holds the window open too, and that state needs an exit as much as a store-lagged one does.

That means "accepted" covers three different situations, and the UpgradeInProgress reason tells you which one you are looking at:

Reason What it means
AcceptedWithPendingApps A pass verified the apps; these ones are still disabled and you accepted that. The message names them.
AcceptedAfterVerification A pass verified every declared app as enabled; the annotation only lifted the window (typically because the automatic lift kept failing). Nothing is known to be wrong.
AcceptedBeforeConvergence You accepted while the upgrade was still rolling, so the declared apps were never checked for this upgrade and the operator will not re-assert them. status.upgrade.pendingApps is empty here because nothing looked — not because nothing is wrong. Check the instance's apps yourself.

status.upgrade.verifiedAt is the underlying signal: absent means no pass ever read the live app state for this upgrade.

Remove it when you are done with it

The annotation applies to the upgrade that was in flight when you applied it. An annotation that is already on the CR when a later upgrade opens its window is treated as left over and ignored, with a StaleUpgradeAcceptIgnored warning event — otherwise a forgotten annotation would abort the next upgrade's orchestration before it converged, lifting the maintenance window mid-roll and leaving the apps unverified.

So a left-behind annotation is no longer dangerous, but it is not harmless either: while it stays on the CR you cannot accept a later upgrade with it (the operator has no way to tell a fresh decision from the old one). Remove it once the situation is resolved; re-apply it if and when you want to accept a different upgrade.

Turning maintenance mode off by hand does not stop the operator

While an upgrade is in flight and the window is held, every convergence pass re-asserts occ maintenance:mode --on. So an occ maintenance:mode --off you run yourself is undone within one retry interval (~15 minutes) — the operator is not fighting you deliberately, it simply has no way to distinguish your deliberate override from the flag being lost to a pod restart. Use the annotation above instead: it is the supported way to say "stop holding this instance", and it records that decision in status where the next person can see it.

Who can do this — and the condition under which that changes

The annotation is an ordinary CR metadata write, so the authorisation boundary is Kubernetes RBAC: it needs patch on nextcloudinstances in the instance's namespace, and nothing weaker. That makes it a platform-operator action in any deployment where tenants do not hold credentials against the cluster API.

Note what the annotation actually overrides, though: the critical-app hold is a protection applied to the tenant — it keeps an instance closed when its only login path is dead. So if a control-plane surface (a self-service portal, a support tool, any automation that writes CR annotations on a tenant's behalf) ever proxies this write, that protection becomes defeasible by whoever can reach that surface, and the RBAC boundary above stops being the real boundary. If you are building such a surface, treat proxying this particular annotation as a privilege decision, not a convenience feature — the failure mode it unlocks is "instance reachable, nobody can log in", presented as if it were healthy.

Advisory pre-flight (AppCompatWarning)

Before the roll, the operator asks the app store whether each declared app has a release for the target major, and emits a Warning event listing any that do not. It never blocks the upgrade: the store is a third party that can be slow, down, or firewalled off, and none of that is a reason to refuse a version bump you asked for.

  • AppCompatWarning — checked, and these apps have no compatible release published.
  • AppCompatUnknown — the store could not be consulted at all. Deliberately a different event: "could not check" must never read as a clean bill of health.

Note the imprecision: the store's platform listing includes apps that have a release for that platform, so an app that ships bundled with the server (not published in the store) is indistinguishable from one with no compatible release, and will appear in the warning. That is acceptable for an advisory signal; it is why this is not a gate.

Disable it entirely with UPGRADE_APP_COMPAT_PREFLIGHT=false (air-gapped landscapes), or point it at a mirror with APP_STORE_URL.

Runbook: apps still disabled after an upgrade

  1. See what is pending and why.
kubectl get nci $NAME -n $NS -o jsonpath='{.status.upgrade.pendingApps}' | jq
kubectl get nci $NAME -n $NS -o jsonpath='{.status.upgrade.steps}' | jq '.[] | select(.success == false)'
kubectl get nci $NAME -n $NS -o jsonpath='{.status.upgrade.maintenanceHeld}'

The failed step's message carries occ's own stderr — an incompatible version reads differently from a store that could not be reached.

  1. Decide whether to wait. If the app simply has no compatible release yet, waiting is the correct action: the loop is already retrying and will heal the instance without you. Check the store yourself to see whether a release exists.

  2. If users are locked out (maintenanceHeld: true, i.e. user_oidc is pending), your choice is between waiting with the instance dark, and accepting the state to bring it back up with SSO broken. If SSO is the only login path, "up" may be worse than "in maintenance" — that is the judgement the operator deliberately does not make for you beyond the default hold.

  3. If the app is gone for good, stop declaring it (spec.apps.<app>.enabled: false, or remove spec.oidc if you are abandoning SSO). The loop converges as soon as the declared set matches reality.

  4. Never hand-run occ app:enable via kubectl exec to "fix" it — that bypasses the per-instance occ lock and races the loop. Use a NextcloudCommand if you need a manual step; see Running occ Commands.

Patch and minor upgrades

Patch (31.0.8 → 31.0.9) and minor (31.0.x → 31.1.x) upgrades within the same major follow the same procedure and carry the same risk profile as any Nextcloud point release:

  1. Confirm a green backup exists (see Backup & rollback below) — cheap insurance, even for a patch bump.
  2. Edit spec.version to the target (exact version, or bump a prefix pin like "31" to pick up the newest 31.x automatically on the next reconcile).
  3. Watch status.phase go Ready → Deploying → Ready.
  4. Confirm post-upgrade tasks ran: status.maintenance.lastRunTrigger should show post-upgrade and status.maintenance.lastVersion should match the new resolved version.

For fleets that don't want to touch every instance by hand, see Fleet strategy below.

Major upgrades

Nextcloud only supports upgrading one major version at a time, and the operator enforces this — see Upgrade-path validation. Treat every major hop as its own change, not a multi-version jump.

Checklist per major hop (e.g. going from 30 → 32 means running this twice: 30 → 31, then 31 → 32):

  1. Back up first. Trigger (or confirm the schedule already ran) an S3 data backup via spec.backups.data, and a pgBackRest backup if the database is managed (spec.database.managed: true). If restoreTest is configured, confirm the most recent test run is green before proceeding — do not upgrade on the strength of an unverified backup.
  2. Validate on staging first if you have a staging instance for this tenant/profile (see Staging / dry-run workaround) — restore the backup there and run the hop before touching production.
  3. Set spec.version to the next major (e.g. "31" or an exact "31.0.9"), not the final target. The operator's validation will reject a skip (e.g. 30 → 32 in one step) with a permanent error.
  4. Wait for status.phase: Ready. Do not queue the next major change while the current hop is still Deploying — let the HelmRelease finish rolling and the pod come back healthy first.
  5. Wait for post-upgrade tasks to complete for this hop (status.maintenance.lastVersion matches, lastRunTrigger: post-upgrade) before starting the next hop. Running occ maintenance:repair and db:add-missing-indices against a schema that's about to be migrated again is wasted work and adds risk.
  6. Spot-check the app — log in, check a few NextcloudCommand occ status/occ app:list calls (see occ during an upgrade) — before starting the next hop.
  7. Repeat from step 3 for each remaining major, one at a time.

Do not try to shortcut this by setting spec.version directly to a version two majors ahead "to save a step" — the operator will reject it (see below), and even if it didn't, Nextcloud's own migrations aren't designed to skip majors.

Upgrade-path validation

Every reconcile that changes the resolved Nextcloud version is checked against status.versionResolution.resolvedVersion (the currently-running version) before the HelmRelease is touched. Downgrades are always unsafe, and so is skipping more than one major in a step (e.g. current 30.x, target 32.x) — but what happens next depends on how the target version was arrived at:

  • An explicit choice fails loud. If the unsafe target comes from something typed directly — an exact spec.version, or the spec.helm.version escape hatch (status.versionResolution.resolvedBy is crd, configmap, or spec.helm.version) — the reconcile fails with a permanent error and status.phase reports Failed. Correct spec.version (step through the intermediate major, or fix the typo) to unblock it.
  • Resolution drift holds instead of failing. If the unsafe target instead comes from the operator re-resolving an alias or prefix that was already set (resolvedBy: alias or resolvedBy: prefix — e.g. latest gets repointed, or an edit to the NextcloudVersionMap changes what "31" resolves to), the instance does not go Failed. The operator holds the instance on its current chart version, logs a warning, and continues reconciling everything else normally — no status condition, no event, just a warning-level log line. This is the ops scenario the hold behavior exists for — see What happens on a version-map rollback below.
  • Patch and minor changes within the same major, and same-major re-resolutions (e.g. latest moving from 32.0.5 to 32.0.6), are unaffected either way — nothing is held or blocked.
  • A target that's a bare prefix of the currently-running version (e.g. target "32" while status.versionResolution.resolvedVersion is 32.0.6 — this happens with the spec.helm.version escape hatch) counts as the same version, never a downgrade.
  • Instances with no prior resolved version (first install) or a non-dotted-integer version string (custom/nightly builds) are never blocked — validation only applies once there's a real "current" version to compare against.

What happens on a version-map rollback

This is the scenario the hold behavior exists for: an ops team ships a bad Nextcloud release, then rolls the NextcloudVersionMap/default CRD back to remove or repoint that entry. Every instance tracking the affected alias or prefix (spec.version: latest, or a prefix pin like "31") re-resolves on its next reconcile — and the new target is now older than what's already deployed, or in the worst case points at a different major entirely.

Without the hold behavior, that one CRD rollback would flip status.phase: Failed across every instance tracking it, fleet-wide, in a single tick — a self-inflicted outage caused by the fix, not the original bad release. Instead, each affected instance simply keeps running its current chart version and logs a warning (grep operator logs for the instance name, or for "upgrade-path"). Nothing else about reconciliation is blocked — secrets, ingress, database, and every other field keep syncing normally in the meantime. Once the version map is corrected (or you explicitly edit that instance's spec.version yourself, which always takes the hard-fail path instead), the hold clears on the next reconcile.

The escape hatch

Annotation: k8s.bnerd.com/allow-unsafe-version-change: "true"

Set this on the NextcloudInstance to bypass both the downgrade check and the major-skip check for as long as the annotation is present — which is why removing it afterwards is part of the procedure. This exists only for documented restore/rollback procedures (see Backup & rollback) — never for routine upgrades. Using it to skip majors in one hop leaves Nextcloud's own migration path unverified; using it to "downgrade" without a genuine restore underneath will corrupt state.

kubectl annotate nci my-nextcloud \
  k8s.bnerd.com/allow-unsafe-version-change=true \
  --overwrite -n my-namespace

Remove it once the change lands — see the numbered rollback procedure below for the full sequence, including when to remove it.

occ during an upgrade

There is no operator-level lock that pauses occ activity while a HelmRelease version change rolls out. What guards exist are incidental, not a coordination mechanism:

  • NextcloudCommand (see Running occ Commands) only dispatches once the target NextcloudInstance is phase: Ready with a Running and Ready pod — so a command submitted mid-upgrade waits, it isn't cancelled or corrupted.
  • The operator's own maintenance tasks require helmRelease.ready before running.
  • Nextcloud's own maintenance mode (set by the image entrypoint during occ upgrade) blocks most occ subcommands for the duration of the migration.

None of this is a substitute for good sequencing. In-flight commands are not cancelled when an upgrade starts, and a manual kubectl exec ... occ ... bypasses every one of these guards — it will happily race a migration.

Recommendation: run all occ invocations through NextcloudCommand, never kubectl exec, and don't submit new commands while a version change is rolling out (status.phase: Deploying). Wait for Ready first.

Fleet strategy

For more than a handful of instances, pick a mix of these rather than editing every spec.version by hand.

Automatic upgrades: spec.upgradePolicy

The primary automation surface, as of 0.20.0. A sibling of maintenance, not nested under it:

spec:
  version: "32"              # prefix pin — the operator resolves it, upgradePolicy decides what happens on drift
  upgradePolicy:
    mode: auto                # manual (default) | auto
    patchUpgrades: true        # auto-apply same-minor patch drift, e.g. 32.0.1 -> 32.0.2
    minorUpgrades: false       # auto-apply same-major minor drift, e.g. 32.0.x -> 32.1.0
  maintenance:
    windowStart: 2             # still governs *when* auto-applied upgrades happen
  • mode: manual (default) — no automatic upgrades; a spec.version edit is the only way to move version. UpdateAvailable (below) is still computed and surfaced regardless of mode — this is the fleet-visibility use case for staying in manual mode deliberately.
  • mode: auto — inside the maintenance window, the operator re-resolves spec.version and applies drift per the flags independently: a resolved patch bump needs patchUpgrades: true; a resolved minor bump (same major, different minor) needs minorUpgrades: true. There is no majorUpgrades field, deliberately — no flag combination automates a major hop; it always requires an explicit spec.version edit, same as today, still fully subject to upgrade-path validation.
  • spec.maintenance.windowStart still governs when an auto-applied upgrade happens — set it alongside upgradePolicy.mode: auto, exactly as it did for autoUpdate before.
  • Per-profile default: NextcloudProfile.spec.defaults.upgradePolicy sets a fleet-wide default for every instance using the profile, with the same deep-merge cascade semantics as every other profile default (see Profiles → How Defaults Merge) — a profile can set mode: auto, patchUpgrades: true and an instance that only overrides minorUpgrades: true still inherits mode/patchUpgrades from the profile; instance values win on conflict.

UpdateAvailable status condition

status.conditions[type=UpdateAvailable] reports whether a newer version currently resolves for this instance — independent of upgradePolicy.mode and independent of whether spec.maintenance.windowStart is even set. The window gates upgrade application, never visibility: a mode: manual fleet, or any instance with no maintenance window configured at all, still gets this signal.

# Block until an update is confirmed available
kubectl wait nci my-nextcloud --for=condition=UpdateAvailable=True --timeout=600s

# Or just read it
kubectl get nci my-nextcloud -o jsonpath='{.status.conditions[?(@.type=="UpdateAvailable")]}' | jq

status is set explicitly to "False" once up to date — not left absent — so kubectl wait works whichever way it resolves. reason is NewerVersionResolved (True) or UpToDate (False); the message names the resolved target version and chart. The condition isn't computed at all (left absent, not False) until the instance has completed its first deploy (status.versionResolution exists) — there's nothing to compare against yet.

On-demand upgrades: k8s.bnerd.com/upgrade-now

Applies the currently-resolved update immediately, bypassing both the maintenance window and upgradePolicy.mode entirely — the point of this annotation is that a mode: manual fleet (or an instance outside its window) can still apply a known-good, already-resolved update on request, without waiting for either.

kubectl annotate nci my-nextcloud \
  k8s.bnerd.com/upgrade-now=$(date -u +"%Y-%m-%dT%H:%M:%SZ") \
  --overwrite -n my-namespace
  • Applies on any drift class — patch, minor, or a resolvedVersion-only/chart-only bump. upgradePolicy.patchUpgrades/minorUpgrades are never consulted here; asking explicitly means getting whatever's pending, not a policy-narrowed subset.
  • Still fully protected by upgrade-path validation — a major-skip or downgrade target is held or fails exactly as it would through any other reconcile trigger; k8s.bnerd.com/allow-unsafe-version-change is still required to bypass that. This annotation is an on-demand trigger, not a validation bypass.
  • No-ops with an event (reason: NoUpdateAvailable) when nothing is pending — check UpdateAvailable first if you're not sure there's anything to apply.
  • Refuses an unparseable resolved target (e.g. a corrupted NextcloudVersionMap entry) rather than guessing at its class — the same guard the mode: auto apply path uses.
  • Firing this at the same moment a mode: auto instance's maintenance-window tick would also apply an update collapses to a single upgrade, not a double-apply — safe to use even mid-window on an auto-mode instance.

The maintenance.autoUpdate alias (deprecated)

spec.maintenance.autoUpdate: true still works. It's a deprecated alias for upgradePolicy: {mode: auto, patchUpgrades: true, minorUpgrades: false}, applied automatically — no spec edit required for existing instances. If spec.upgradePolicy is also set on the same instance, upgradePolicy wins outright and autoUpdate is ignored entirely; a reason: UpgradePolicyAliasIgnored warning event fires (once per reconcile that has pending drift, not on every tick) so the conflict doesn't go unnoticed.

The alias is not removed in 0.20.0 and won't be removed before 0.21.0 — existing autoUpdate: true specs keep working unchanged. But its exact behavior changed underneath it in 0.20.0; read both notes below before assuming it behaves identically to before.

Upgrade notes: two separate, opposite-direction changes in 0.20.0

Don't conflate these — they affect different things and point in opposite directions.

1. Narrowing — the alias now only auto-applies patches, not minors. Before 0.20.0, autoUpdate: true applied any same-major drift, patch or minor, without distinguishing between them. As of 0.20.0 the alias maps to patchUpgrades: true, minorUpgrades: false specifically — a resolved minor bump (e.g. 32.0.x → 32.1.0) no longer auto-applies under the alias alone. If you were relying on minor drift auto-applying, add spec.upgradePolicy: {mode: auto, patchUpgrades: true, minorUpgrades: true} explicitly (this also moves you off the alias entirely, onto upgradePolicy, so the ignored-warning never fires for you).

2. Widening — drift detection itself now also catches resolvedVersion-only bumps. Before 0.20.0, drift detection compared only the resolved chart version. Several Nextcloud patch releases can map to the same chart version in NextcloudVersionMap (e.g. Nextcloud 32.0.6/32.0.8/32.0.9 all resolving to chart 8.9.1) — that class of drift was previously invisible to autoUpdate entirely; it silently never applied, with no signal that anything was pending. As of 0.20.0, drift detection compares both the chart version and the resolved Nextcloud version (the same basis UpdateAvailable uses). This class of drift is now detected — and, with patchUpgrades: true (which the alias sets), auto-applied — on the very next maintenance-window tick after upgrading to 0.20.0, for any instance already inside its window with autoUpdate/upgradePolicy set. This is a real, immediate behavior change for existing fleets on their next window tick, not just a change for newly-created instances.

Other fleet patterns

  • Exact pin + mode: manual. spec.version: "31.0.9" with no upgradePolicy (or mode: manual, or the equivalent unset/false autoUpdate) is a true no-op fleet-wide — nothing moves until you edit the spec. Use for instances under strict change control. UpdateAvailable still tells you when something's pending.
  • Profile-level version pinning. NextcloudProfile.spec.defaults.version pins every instance using that profile at once — useful for rolling a cohort of tenants to a known-good version without touching each NextcloudInstance. An instance's own spec.version always overrides the profile default. See Version Management → Profile-Level Version Pinning.
  • Mixed versions across the fleet are normal and supported. Different tenants/profiles can sit on different majors simultaneously — there is no fleet-wide version requirement. Stagger major hops by profile or by tenant tier (e.g. staging profile a major ahead of production) to get a soak period for free.

Recommended production default: prefix pin ("31") + upgradePolicy: {mode: auto, patchUpgrades: true} + a windowStart for automatic patch upgrades, with explicit spec.version edits reserved for major hops, following the Major upgrades checklist. Use k8s.bnerd.com/upgrade-now when you need a specific instance upgraded before its next window tick.

Backup & rollback

Nextcloud has no in-place downgrade. "Rollback" always means restoring pre-upgrade backups, not reverting spec.version on a live, already-migrated instance.

Before any upgrade:

  1. Confirm spec.backups.data is configured (S3 data backup via the bnerd backup operator) and, if restoreTest is set, that the last automated restore test succeeded.
  2. If the database is managed (spec.database.managed: true), confirm scheduled pgBackRest backups are enabled and current (spec.database.postgres.backup) — see Managed PostgreSQL.
  3. Note the pre-upgrade status.versionResolution (requested/resolved/chart versions) so you know exactly what "rollback" means for this instance.

If an upgrade needs to be rolled back:

  1. Stop traffic to the instance if it's serving users (scale down, maintenance page at the ingress level, or similar — outside the operator's scope).
  2. Restore the pre-upgrade S3 data backup and the pre-upgrade database backup (pgBackRest restore for managed Postgres, or your external DB's own restore procedure) to the point captured before the failed upgrade.
  3. Set the annotation k8s.bnerd.com/allow-unsafe-version-change: "true" on the NextcloudInstance — required because setting spec.version back to the pre-upgrade value is, from the operator's point of view, a downgrade.
  4. Set spec.version back to the pre-upgrade value (the resolvedVersion you noted above, not a prefix or alias — be exact so the resolved version matches what you just restored).
  5. Wait for status.phase: Ready and spot-check the app.
  6. Remove the k8s.bnerd.com/allow-unsafe-version-change annotation once the instance is stable. Leaving it set disables downgrade/major-skip protection for every future reconcile, not just this one.

There is no automated rollback — this is a manual procedure end to end. Practice it against a staging instance before you need it in production.

Staging / dry-run workaround

There is no built-in dry-run or clone-for-upgrade-testing feature. Until one exists, approximate it manually:

  1. Create a staging NextcloudInstance using the same profile as the production instance you're planning to upgrade.
  2. Restore a recent production backup into the staging instance (same restore procedure as Backup & rollback, steps 2–3, minus the annotation since staging starts from the pre-upgrade version already).
  3. Run the upgrade (or the next major hop) against staging first, following the Major upgrades checklist.
  4. Validate via NextcloudCommand — occ status, occ app:list --output=json, and any app-specific sanity checks — before repeating the change against production.
  5. Tear down or reuse the staging instance for the next hop.

App compatibility

spec.apps installs run as occ app:install <app> || true in the chart's before-starting hook — best-effort only. There is no compatibility check against the target Nextcloud version, and a failed install does not surface in status; it silently no-ops. Before a major upgrade, check each installed app's compatibility with the target Nextcloud version against the Nextcloud app store yourself, and validate via a staging run (above) rather than trusting spec.apps to fail loudly if an app breaks.

See also

  • Standard Upgrade Procedure — the single consolidated checklist built from this guide.
  • Version Management — how spec.version resolves to a chart version, the NextcloudVersionMap CRD, image resolution, profile-level pinning.
  • Operations & Annotations — the full annotation/label reference, including k8s.bnerd.com/reconcile and k8s.bnerd.com/run-maintenance.
  • Running occ Commands — the NextcloudCommand CRD used for validation and spot-checks above.
  • Managed PostgreSQL — pgBackRest backup configuration for database.managed: true.