Operations & Annotations¶
This guide lists the annotations and labels you can set on managed resources to drive day-2 operations — forcing a reconciliation, triggering maintenance on demand, or overriding deletion protection. It also documents the labels the operator sets itself so you know what is safe to rely on in kubectl -l selectors.
All keys use the k8s.bnerd.com/ prefix.
Summary¶
| Key | Kind | Resource | Purpose |
|---|---|---|---|
k8s.bnerd.com/reconcile |
annotation | Nextcloud, NextcloudInstance |
Force an immediate reconciliation |
k8s.bnerd.com/run-maintenance |
annotation | NextcloudInstance |
Run OCC maintenance tasks immediately |
k8s.bnerd.com/force-delete |
label | NextcloudInstance |
Allow deletion of an assigned pool instance |
k8s.bnerd.com/allow-unsafe-version-change |
annotation | NextcloudInstance |
Bypass upgrade-path validation (downgrade / major-skip) for a restore or rollback |
k8s.bnerd.com/upgrade-now |
annotation | NextcloudInstance |
Apply a pending version update immediately, bypassing the maintenance window and upgradePolicy.mode |
k8s.bnerd.com/upgrade-apps-accept |
annotation | NextcloudInstance |
Accept an upgrade whose declared apps are still disabled: lift the maintenance hold and stop the convergence retry loop |
k8s.bnerd.com/upgrade-resync |
annotation | NextcloudInstance |
Re-derive status.upgrade from the instance's live state; clears bookkeeping that no longer describes it. Lifts maintenance mode only if one of the operator's own upgrade windows is recorded as holding it — anything else is reported (MaintenanceModeNotOurs) and left alone. Spent automatically once it has acted |
Force Reconcile¶
Annotation: k8s.bnerd.com/reconcile
Applies to: Nextcloud, NextcloudInstance
Value: any string; convention is an ISO 8601 timestamp or date +%s. Only changes to the value trigger a reconcile — setting the same value twice is a no-op.
Normal reconciliation runs whenever a relevant spec field changes or on the periodic timer (30 s for NextcloudInstance, 60 s for Nextcloud). Use this annotation to trigger a reconcile immediately in cases where the operator would not otherwise notice a change.
When to use¶
- After editing a
NextcloudProfile— existing instances don't pick up profile changes automatically - To retry provisioning after a transient error (managed database creation, S3 bucket auto-creation, HelmRelease failure)
- After rotating a referenced secret (
credentialsSecret) so the new values are read and re-applied - To propagate a
Nextcloudspec change to its assignedNextcloudInstanceif the sync appears stuck
How to trigger¶
# Force reconcile a NextcloudInstance
kubectl annotate nci my-instance \
k8s.bnerd.com/reconcile=$(date -u +"%Y-%m-%dT%H:%M:%SZ") \
--overwrite -n my-namespace
# Force reconcile a Nextcloud (logical resource)
kubectl annotate nc my-tenant \
k8s.bnerd.com/reconcile=$(date +%s) \
--overwrite -n my-namespace
A ready-made helper script is shipped with the repository:
Watch the operator logs to confirm the trigger fired — you should see Manual reconciliation triggered for ....
On-Demand Maintenance¶
Annotation: k8s.bnerd.com/run-maintenance
Applies to: NextcloudInstance
Value: any string; use a fresh timestamp on each run.
By default, periodic OCC maintenance (file cleanup, missing-index checks, etc.) runs once per day during the window configured in spec.maintenance.maintenanceWindow. Set this annotation to run the same tasks immediately, outside the window.
This does not re-run post-upgrade tasks (those are tied to a version change). It runs whichever periodic tasks are enabled in spec.maintenance.tasks.
When to use¶
- You want to reclaim disk space or tidy up orphaned files right now without waiting for the window
- You just enabled a new task in
spec.maintenance.tasksand want to run it once straight away - You are investigating an issue and want fresh output from
db:add-missing-indicesor similar
How to trigger¶
kubectl annotate nci my-instance \
k8s.bnerd.com/run-maintenance=$(date +%s) \
--overwrite -n my-namespace
After the run completes, status.maintenance.lastRunTrigger is set to annotation and status.maintenance.lastRunAt is updated:
See the API reference for the full list of tasks and timeout knobs under spec.maintenance.
For arbitrary occ commands (not just the canned maintenance tasks), use the NextcloudCommand CRD. Where run-maintenance triggers a fixed task set, NextcloudCommand lets you declaratively run any occ invocation with per-command result reporting.
All occ execs against a given instance — triggered by this annotation, NextcloudCommand, admin-credential apply, or the upgrade/apps-health checks — serialize automatically per instance (0.20.0+); see Running occ Commands → Limitations for the exact guarantee and the lock-timeout fallback.
Force Delete¶
Label: k8s.bnerd.com/force-delete=true
Applies to: NextcloudInstance
Note: This is a label, not an annotation. Use
kubectl labelrather thankubectl annotate.
Pool-provisioned NextcloudInstance resources carry a finalizer (k8s.bnerd.com/assigned-instance-protection) and labels linking them to their assigned Nextcloud. Deleting an assigned instance directly is normally blocked with a TemporaryError so that the Nextcloud doesn't end up pointing at a gone backend.
Set k8s.bnerd.com/force-delete=true on the instance to bypass that protection. Use as a last-resort escape hatch: this will leave the assigned Nextcloud in a broken state until you re-point it at another instance or delete it.
When to use¶
- The instance is hard-stuck in a failed state and needs to be removed before the
Nextcloudcan be re-assigned - You are manually draining a pool and intend to delete both the
Nextcloudand the instance
How to trigger¶
# Check what the instance is assigned to first
kubectl get nci my-instance -o jsonpath='{.metadata.labels}' | jq
# Override protection and delete
kubectl label nci my-instance k8s.bnerd.com/force-delete=true -n my-namespace
kubectl delete nci my-instance -n my-namespace
The preferred path is to delete the owning Nextcloud first — that clears the assignment and the instance deletes cleanly without the force label.
Upgrade-Path Override¶
Annotation: k8s.bnerd.com/allow-unsafe-version-change
Applies to: NextcloudInstance
Value: "true" to bypass; remove (or unset) to restore normal protection.
The operator refuses to resolve spec.version to a downgrade or a major-version skip (e.g. 30 → 32 in one hop) — see Upgrades → Upgrade-path validation. This annotation is the escape hatch, intended only for documented restore/rollback procedures where you're intentionally setting spec.version back to a pre-upgrade value after restoring backups. Remove it once the instance is stable; leaving it set disables the protection for every subsequent reconcile, not just the one you needed it for. Full procedure: Upgrades → Backup & rollback.
On-Demand Upgrade¶
Annotation: k8s.bnerd.com/upgrade-now
Applies to: NextcloudInstance
Value: any string; use a fresh timestamp on each trigger.
Applies the currently-resolved version update immediately, bypassing both spec.maintenance.windowStart and spec.upgradePolicy.mode entirely. See Upgrades → On-demand upgrades for the full behavior (drift-class bypass, upgrade-path validation still applies, the NoUpdateAvailable no-op, the annotation-vs-timer race handling).
When to use¶
- A
mode: manualinstance hasUpdateAvailable: Trueand you want it applied right now, without switching the instance tomode: autoand waiting for a window. - An
auto-mode instance's maintenance window hasn't arrived yet but you need the update applied sooner (an out-of-band security patch, a coordinated maintenance slot). - You want to apply a resolved patch bump to one specific instance without waiting for its next scheduled window tick.
What to expect¶
- If nothing is pending (
UpdateAvailable: Falseor absent), this is a no-op — areason: NoUpdateAvailableevent fires and nothing else happens. - If something is pending, the operator bumps
k8s.bnerd.com/reconcileinternally to trigger an immediate reconcile with the resolved target — the sameHelmReleaseupdate / Flux rollout / post-upgrade task sequence as any other upgrade (see Upgrades → How an upgrade happens). - A resolved target that's a major skip or downgrade is still held or blocked by upgrade-path validation — this annotation does not bypass that; you still need
k8s.bnerd.com/allow-unsafe-version-changeon top for a documented rollback.
How to trigger¶
kubectl annotate nci my-instance \
k8s.bnerd.com/upgrade-now=$(date -u +"%Y-%m-%dT%H:%M:%SZ") \
--overwrite -n my-namespace
How to verify¶
# Confirm there was something to apply, before triggering
kubectl get nci my-instance -n my-namespace \
-o jsonpath='{.status.conditions[?(@.type=="UpdateAvailable")]}' | jq
# After triggering, watch the phase and the resolved version update
kubectl get nci my-instance -n my-namespace -o jsonpath='{.status.phase}'
kubectl get nci my-instance -n my-namespace -o jsonpath='{.status.versionResolution}' | jq
# If it no-op'd, confirm why
kubectl get events -n my-namespace --field-selector involvedObject.name=my-instance | grep -i upgrade
Accepting a stalled post-upgrade app convergence¶
Annotation: k8s.bnerd.com/upgrade-apps-accept
Applies to: NextcloudInstance
Value: "true".
Since 0.23.0 the operator holds Nextcloud's maintenance mode across a version upgrade until it has verified that every declared app is enabled again, and keeps holding it when a critical app (user_oidc, while spec.oidc.enabled) cannot be enabled. It retries indefinitely, because the usual cause — an app store that has not yet published a release for the new major — resolves on its own within days. See Upgrades → Post-upgrade app convergence.
This annotation is the human override for "I know, stop holding this instance."
When to use¶
- SSO is broken by a pending
user_oidc, and you would rather have the instance up with local logins than down with a maintenance page. - The affected app is not coming back and you are not ready to edit
spec.appsyet.
What to expect¶
- Maintenance mode is lifted,
status.upgrade.phasebecomesCompletedwithaccepted: true, and thenextcloud_operator_upgrade_apps_pendingseries is dropped. AppsHealthystaysFalse. Accepting is not a claim of health — the degraded app state stays visible until it actually changes.- The retry loop stops. If you want the operator to try again afterwards, trigger a reconcile (
k8s.bnerd.com/reconcile) after fixing the underlying cause, or let the next version change open a fresh window. - If the operator cannot reach the pod to lift the window, the acceptance is recorded but the phase stays
AppsPendingand the lift is retried — you will not get aCompletedon an instance that is still dark.
How to trigger¶
kubectl annotate nci my-instance \
k8s.bnerd.com/upgrade-apps-accept=true --overwrite -n my-namespace
How to verify¶
kubectl get nci my-instance -n my-namespace -o jsonpath='{.status.upgrade}' | jq '{phase, accepted, maintenanceHeld, pendingApps}'
Remove the annotation once the situation is resolved. It applies to the upgrade that was
in flight when you applied it: an annotation still on the CR when a later upgrade opens
its window is treated as left over and ignored (with a StaleUpgradeAcceptIgnored
warning event), so a forgotten one cannot abort the next upgrade — but while it stays
there you also cannot accept a later upgrade with it. Remove it, and re-apply it if and
when you mean to accept a different one.
Authorisation: this is a plain CR annotation write (patch on nextcloudinstances in
the instance's namespace) and it overrides a protection applied to the tenant — see
Upgrades → Escape hatch for why
proxying it from a self-service surface is a privilege decision rather than a
convenience.
Admin Credential Management¶
As of 0.20.0 (#7866), spec.admin.{username,password} (inline or via
credentialsSecret) is not just recorded on the NextcloudInstance — the
operator actively applies it to the live Nextcloud instance. This is the
"Rails-as-source-of-truth" pattern: a control plane (or any other CR author)
declares the credential once on the CR, and can rely on the operator to make
that credential work against the running instance, rather than separately
tracking whether it was ever installed.
Trigger points¶
The operator converges the live admin credential to spec.admin from three
places, all funneling into the same idempotent apply step:
- PoolAssignment — see Pool Provisioning → Admin Credential Handoff at Assignment.
- A
spec.adminchange on an already-assigned instance (aspec.adminfield watcher). - The periodic instance-status timer — a safety net that re-checks on every reconcile tick so a missed event still converges.
How application works¶
- The desired credential is resolved via the same
credentialsSecret/inline loader used elsewhere (admin.credentialsSecret, oradmin.username/admin.passworddirectly). - Precedence (0.21.5): a
credentialsSecretyou supply outranks an inlinepassword, as documented. The operator's own<instance>-nextcloud-adminSecret does not: whenadmin.credentialsSecretnames that Secret andadmin.passwordis non-empty, the inline password wins, because that Secret is a materialization of the desired credential rather than a source of truth. See Credentials guide → Changing a declared admin password for the full table and the rewrite-then-blank sequence. - The literal
changemeis never treated as a declared password (0.21.5). TheNextcloudInstanceCRD defaultedspec.admin.passwordto it until this release, and the API server injects a schema default into every admin object that omits the key. On 0.21.0–0.21.4 that made the literal an instance's real admin password; 0.21.5 refuses it in the resolver, in both auto-generate gates and in the shadow rule, so an old separately vendored CRD cannot reintroduce it from the spec. It does not rotate a password already provisioned that way — that value lives in the admin Secret, which is the resolved credential, so it keeps being applied until someone rotates it. Audit the<instance>-nextcloud-adminSecrets. - The blank-back-out is conflict-safe: the operator re-reads the live spec and
clears only the exact password it just applied, via a JSON-Patch
testop. A password changed again mid-apply is never swallowed — it wins the next reconcile. - Existence is probed with
occ user:info <username>: - Exists →
occ user:resetpasswordresets the password. - Missing →
occ user:add --group admincreates the user (create-if-missing). - The password is transported to the pod exec's stdin only — never in
command-line arguments, never written to any CR, and never logged (asserted
by tests). It is not recoverable from
kubectl describe, operator logs, or Kubernetes events. - The previous admin account is never deleted, renamed, or disabled. If
spec.admin.usernamechanges, the operator creates the new account and leaves the old one exactly as it was — cleanup, if wanted, is a manual step.
Idempotency¶
status.appliedAdminCredential records {username, passwordHash, appliedAt}
for the credential last applied to the live instance. passwordHash is
PBKDF2-HMAC-SHA256 (600,000 iterations) of the password, salted with the
instance UID — self-describing format pbkdf2-sha256$600000$<hex>. It is
never the plaintext password, only a drift-detection fingerprint. The apply step
no-ops (no exec, no Secret rewrite) once username and passwordHash both
already match what's resolved from spec.admin.
Failure surface¶
A condition AdminCredentialApplied is set on the instance:
reason |
status |
Meaning |
|---|---|---|
Applied |
True |
Password reset succeeded against an existing admin account. |
UserCreated |
True |
The declared username didn't exist on the live instance; occ user:add --group admin created it. |
ExecFailed |
False |
The occ exec failed (pod not reachable, command error, etc.). A Kubernetes warning event is also emitted. The operator retries (kopf.TemporaryError, never a permanent failure). |
CredentialsUnavailable |
False |
The declared credential couldn't be resolved — most often a spec.admin.credentialsSecret that doesn't exist yet. A Kubernetes warning event is also emitted, and the operator retries indefinitely (never a permanent failure), so a late-arriving Secret converges on its own. |
The condition is re-asserted on every attempt, so it always reflects the current state.
The event is not: the operator emits one Warning per distinct failure, not one per
reconcile tick, so a persistently broken instance doesn't bury its own timeline. A rotated
credential that also fails counts as a new failure and gets its own event. The marker lives
in status.adminCredentialLastNotifiedFailure and is cleared once the credential applies
(or already matches), so nothing is ever silently suppressed.
Failure messages now include the underlying error (0.21.1). Both ExecFailed
and CredentialsUnavailable condition messages fold in a sanitized snippet of
the actual exception — never raw secret material, redaction is tested —
instead of a generic "failed to apply" message, so the real cause (a specific
occ exec error, a missing Secret key, a broken mail credentialsSecret) is
visible from kubectl describe nci without a manual pod exec. A broken
spec.mail.credentialsSecret also surfaces through this same path: it's
read here to build the admin Secret's mail_config, and before 0.21.1 a
failure there escaped as a raw, unsurfaced exception with no condition or
event at all.
Safety-net logging raised to WARNING (0.21.1). The instance-status
timer's catch-all around this apply step now logs at WARNING once per
failure onset and DEBUG for repeats of the same failure — bounded by the
same adminCredentialLastNotifiedFailure dedupe marker above, not a second
mechanism. Before 0.21.1 every failure logged at DEBUG, invisible at the
deployed INFO level.
Hash-match no-op path now self-heals a missing condition (0.21.1). A rare
race (a same-tick sibling patch seeding conditions from a stale snapshot)
could drop the condition on the credential's first successful apply while
status.appliedAdminCredential was still correctly recorded — a healthy
instance with no visible AdminCredentialApplied condition at all, and no
later tick ever rewrote it. The no-op path now re-asserts
AdminCredentialApplied=True/Applied whenever the ambient condition is
absent or not True, closing the gap within one reconcile tick; when the
condition already reads True, the tick remains a genuine zero-write no-op.
An instance that isn't Ready yet has no pod to exec into, so there is nothing
to apply to and no condition is written. How that's handled depends on which
trigger point noticed:
- The
spec.adminfield watcher returns and lets the periodic instance-status timer pick the work up on the next Ready tick. It does not retry:spec.adminexists on effectively every instance, so this watcher runs as part of kopf's create cycle for the object, and a handler that retries forever holds that cycle open — which in turn stops kopf from dispatching watch-driven handlers (aspec.versionedit going unreconciled) and from removing its own finalizer during deletion (an instance stuckTerminatinglong after the operator's own cleanup finished). Fixed in 0.21.0. - The instance-status timer only calls the apply step for an instance it has
just confirmed
Ready, so the credential converges within one tick of the instance becoming usable.
An instance carrying a metadata.deletionTimestamp is skipped entirely, from
every trigger point — nothing is applied to an object being torn down.
Generated admin credentials and the 0.21.0 migration¶
This section is about credential hygiene (what's stored where), which is a different concern from the live-apply behavior above (what's active on the running instance). See the Credentials guide → Generated Admin Credentials for the full mechanism; summary:
- Since 0.21.0, a generated admin password is stored only in the
<instance>-nextcloud-adminSecret. The CR getsspec.admin.credentialsSecretwritten back, never the plaintext, and is marked with thek8s.bnerd.com/admin-credential: generatedannotation. - On the first reconcile after upgrading to 0.21.0, pool-origin instances
whose spec still carries a plaintext generated password (from before this
change) are migrated in place to the ref form — one
MigratedAdminCredentialevent, no live credential change (the hash-gate no-ops, so nooccexec runs). - Non-pool instances with an inline password are left untouched, with a
one-time
InlineAdminPasswordDeprecatedadvisory event (see the annotations table above for the gating annotation).
Runbook: rotate the admin password of a pool-provisioned instance¶
A pool-provisioned instance's spec carries the operator's own admin Secret ref and an empty password. To set a specific password:
# 1. Declare it. Usually on the parent Nextcloud CR (it passes spec.admin
# straight through to the instance); a direct edit of the NextcloudInstance
# works the same way. Declaring it up front, on a Nextcloud CR that is about
# to be pool-assigned, goes through the same path.
kubectl patch nc <name> -n <ns> --type merge \
-p '{"spec":{"admin":{"password":"<new-password>"}}}'
# 2. Watch status.appliedAdminCredential — this is the field that moves.
# (Seconds, not minutes: the instance's spec.admin watcher fires on the
# patch itself; the 30s instance-status timer is only the safety net.)
kubectl get nci <instance> -n <ns> -o jsonpath='{.status.appliedAdminCredential}'
# -> appliedAt advances and passwordHash changes once the new credential is live
# 3. OPTIONAL, and not a pass/fail gate: the operator blanks the plaintext back
# out of the instance spec after applying it.
kubectl get nci <instance> -n <ns> -o jsonpath='{.spec.admin}'
# -> {"credentialsSecret":"<instance>-nextcloud-admin","password":""}
# ...but if the PARENT Nextcloud CR still declares the password, the next spec
# passthrough (within ~30s) puts it back. Seeing the plaintext here again is
# expected in that case and does NOT mean the apply failed — step 2 is the gate.
# It is also skipped entirely for a GitOps-managed CR (Flux owns the spec).
# 4. The Secret now holds the new password.
kubectl get secret <instance>-nextcloud-admin -n <ns> \
-o jsonpath='{.data.nextcloud-password}' | base64 -d
Don't verify this with AdminCredentialApplied's lastTransitionTime
On an instance that already had a credential applied — every pool-provisioned
instance, which got its generated password applied at provisioning time — the
condition goes True → True. Per the Kubernetes convention the operator
follows, lastTransitionTime is preserved across a same-status update, so
it does not move even though the new password was applied. Use
status.appliedAdminCredential.appliedAt / passwordHash (step 2), the admin
Secret's resourceVersion, or an actual login. An unchanged condition
timestamp is not evidence of failure.
If the credential does not apply, the condition flips to False — read its
reason and message (kubectl describe nci <instance>): ExecFailed,
CredentialsUnavailable and SecretWriteFailed all carry a sanitized snippet of
the underlying error. SecretWriteFailed means the operator could not even rewrite
the owned <instance>-nextcloud-admin Secret, so the live occ apply was never
attempted. The instance must be Ready for the credential to be applied at all —
there is no pod to occ into otherwise.
To stop the plaintext reappearing (step 3), remove spec.admin.password from the
parent Nextcloud CR once step 2 shows the new credential applied; resolution then
continues through the Secret.
Upgrade note: first reconcile after 0.20.0 resets drifted passwords fleet-wide¶
status.appliedAdminCredential is new in 0.20.0, so every already-Ready
instance starts with no recorded applied credential. On the very next
periodic instance-status timer tick after upgrading, every Ready
instance in the fleet gets its declared spec.admin applied — unconditionally,
not staged or throttled. For an instance whose live password already matches
spec.admin.password, this is a harmless extra occ exec. For an instance
where the two have drifted (exactly the scenario this feature exists to fix),
the live admin password will be reset to the CR's declared value at that
moment, with no separate trigger or advance notice. This is the feature
working as intended, not a bug — but plan for it before upgrading a fleet
where spec.admin may not reflect the current live password: audit for drift
first if that matters for your rollout, since the change happens automatically
on upgrade rather than on the next spec.admin edit.
Consumer contract: the admin Secret can briefly diverge from the live instance¶
The operator syncs the operator-owned <instance>-nextcloud-admin Secret to
the desired credential before attempting the live occ exec. If the
exec then fails, the Secret already reflects the new password while the live
instance's actual password is still the old one — AdminCredentialApplied
flips to False/ExecFailed and a warning event fires, so the failure is
visible, but only to something watching the CR's condition or events, not to
something that only reads the Secret.
Anything that reads this Secret as its source of truth for "what credential
currently works against the live instance" (e.g. an external control plane)
must check AdminCredentialApplied=True before trusting it — the Secret
being up to date does not by itself guarantee the live instance matches it.
Profile Propagation at Upgrade¶
Upgrading the operator to 0.20.0 does not, by itself, retroactively reconfigure any already-running instance — see Profiles → When Profiles Apply for the full mechanism (profile defaults apply once, at instance creation, by design). Concretely, for an existing profile-backed instance:
- The
specwrite-back is unchanged, still and forever create-time-only. The 0.20.0 deep-merge fix, the correcthookswholesale-replace behavior, and the widerspecwrite-back only take effect at instance creation — they apply to instances created after the 0.20.0 operator is live, not to ones that already exist.spec.database/spec.s3/etc. on an existing instance stay exactly whatever was persisted under the previous (possibly incorrectly-merged) logic. This part has not changed since 0.20.0. - Rendered Helm values are a different story as of 0.21.1 (review finding
E-1). Before 0.21.1, rendering also skipped re-resolving the profile once
an instance was materialized, so the previous bullet's "unchanged" claim
held for Helm values too. 0.21.1 removed that skip — the profile now
resolves fresh on every render, though the instance's own materialized
value (from the write-back above, or your own
spec.helm.values) still wins the merge and is never overridden. Net effect for an already-running instance: a profile edit to a key it already has a value for still never propagates; a key added to the profile since creation now renders on the next reconcile; and an instance whose write-back never ran at all (created before operator v0.10.2, or GitOps-managed) now gets its entire profile-derived Helm values rendered correctly, where before they silently never appeared. See Profiles → When Profiles Apply for the full contract. Rollout note: on the first reconcile after upgrading to 0.21.1, any instance whose profile gained keys since it was created — or whose spec write-back never ran — renders those keys for the first time, which changes that instance'sHelmReleaseand restarts its pod(s) once. - Managed-PostgreSQL is one exception. On the next reconcile after
upgrading, a managed-PostgreSQL, profile-backed instance's live
PerconaPGClusterpgBouncer configuration is corrected to match the current profile — regardless of whetherspec.databaseitself was ever backfilled, and regardless of GitOps management. See Profiles → When Profiles Apply → Exception. spec.mailgets a one-time heal in 0.21.0. Instances that materialized before 0.20.0's merge-kernel fix can be permanently missing profile-provided mail fields (most notablyfromAddress— #7870). Starting in 0.21.0, the periodic instance-status timer gap-fills any such missing keys from the instance's applied profile, once, for non-GitOps instances — see Profiles → When Profiles Apply → One-time heal. Unlike pgBouncer, this is not a continuous sync: after the one-time fill,spec.mailreverts to the normal once-only rule above. This write changes rendered Helm values, so the affected instance's Nextcloud pod(s) restart once the next time this runs after upgrading to 0.21.0.- New instances get the full fix from creation onward: correct
deep-merge, wider write-back (skipped for GitOps-managed CRs — logged
instead, not silently dropped),
hookswholesale-replace, and the continuously profile-aware pgBouncer reconciler.
Don't confuse this with admin credential application
above, which behaves the opposite way: spec.admin is actively
re-applied to every already-Ready instance's live Nextcloud on the first
reconcile after upgrading. Profile-derived configuration and admin
credentials follow different rules in 0.20.0 — only admin credentials and
the pgBouncer proxy self-heal on upgrade; everything else stays as-is.
HelmRelease Value Ownership¶
Starting in 0.21.1, the operator owns the managed HelmRelease's spec.values
wholesale: every update rebuilds the complete values object from the
instance's spec and profile, and replaces the live spec.values with it
entirely (a JSON Patch add, not a merge). Previously, an update
merge-PATCHed the live object, so any key present on the HelmRelease but
absent from the freshly-built values was left untouched — a stale value from
an earlier spec revision, or a value someone hand-edited on the HelmRelease
directly, could silently persist forever. A manual edit to a managed
HelmRelease's spec.values is removed on the instance's next reconcile.
Manage the underlying setting through the NextcloudInstance/Nextcloud
spec (or a NextcloudProfile), not by hand-editing the rendered
HelmRelease — the same rule that already applied to every other
operator-managed field, now also true for spec.values.
Only the fields the operator itself sets are affected: spec.interval,
spec.timeout, spec.chart, spec.values, spec.install, and
spec.upgrade. Anything else on the HelmRelease survives untouched,
because a JSON Patch only touches the exact paths it lists — most notably:
spec.suspend(set viaflux suspend hr <name> -n <namespace>) is never one of the operator's ops, so pausing aHelmReleasefor an incident survives every subsequent operator reconcile until you resume it (flux resume hr).- Any other field Flux or a different controller manages on the object
(
dependsOn,postRenderers, ...) is likewise untouched.
This closed a real incident: a stale cert-manager.io/cluster-issuer
ingress annotation left behind by a spec.ingress.tls.certManager: false
flip caused cert-manager to keep issuing a certificate into the instance's
wildcard TLS Secret, clobbering it on every renewal. See the
CHANGES.md
v0.21.1 entry for the full writeup.
Backup Runbook (0.22.0)¶
Day-to-day operational procedures for the two backup halves. The full picture, including restore, lives in Backup & Restore.
Enable off-cluster backups for the whole estate¶
One change to the operator's own Helm release plus two Secrets in its namespace — no edit to any instance, profile or CRD.
NS=nextcloud-operator-system
# 1. S3 credentials for the backup bucket
kubectl -n "$NS" create secret generic db-backup-s3 \
--from-literal=accessKey='...' --from-literal=secretKey='...'
# 2. The backup-encryption passphrase. Put this in the password manager TOO --
# a backup encrypted with a key you have lost is noise.
openssl rand -base64 48 > cipher.key
kubectl -n "$NS" create secret generic db-backup-cipher --from-file=cipher-pass=cipher.key
# 3. Point the operator at the bucket
helm upgrade nextcloud-operator ./chart -n "$NS" --reuse-values \
--set backup.dbDefaultS3.bucket=nextcloud-db-backups \
--set backup.dbDefaultS3.endpoint=s3.example.com \
--set backup.dbDefaultS3.region=eu-central-1 \
--set backup.dbDefaultS3.secretRef=db-backup-s3 \
--set backup.dbDefaultS3.cipherSecretRef=db-backup-cipher
Roll it out to one instance first (set the values, confirm the checks below, then let the rest of the estate follow on its own reconcile ticks).
The operator needs a writable working directory
Config bundles are staged on disk before upload. The operator runs with
readOnlyRootFilesystem: true, so the chart and deploy/operator.yaml both mount
an emptyDir at /tmp (workDir.sizeLimit, default 1Gi) with matching
ephemeral-storage requests and limits. If you render your own Deployment or strip
volumes, keep that mount — without it every bundle fails with a permission error
and status.configBackup.lastFailureReason will say so.
Peak usage is roughly 3x the bundle size (streamed parts, the assembled tar and the
encrypted copy exist together), so raise both the sizeLimit and the
ephemeral-storage limit together if you raise the bundle guard.
Verify an instance picked it up¶
NS=<instance-namespace>
INST=<instance-name>
# repo2 exists and carries a PER-INSTANCE path
kubectl -n "$NS" get perconapgcluster "$INST-pg" \
-o jsonpath='{.spec.backups.pgbackrest.repos[*].name}{"\n"}{.spec.backups.pgbackrest.global.repo2-path}{"\n"}'
# the derived credentials Secret was materialized
kubectl -n "$NS" get secret "$INST-pgbackrest-s3"
# the one-time bootstrap full was triggered
kubectl -n "$NS" get nci "$INST" -o jsonpath='{.status.backupRepo2Bootstrapped}{"\n"}'
kubectl -n "$NS" get perconapgbackup -o custom-columns=NAME:.metadata.name,REPO:.spec.repoName,STATE:.status.state
# the configuration bundle landed
kubectl -n "$NS" get nci "$INST" -o jsonpath='{.status.configBackup}' | jq
Expected: repo1 repo2, a repo2-path of /<namespace>/<instance>/backups/db, a
BackupRepo2Bootstrapped event, and a ConfigBackupCompleted event within a day.
A running bundle holds the instance's exec lock¶
The operator serializes pod execs per instance (the same lock occ commands take), and
a config-bundle run holds it for the whole run — the path probe, the tar stream
and the upload. On a large instance that can be a few minutes.
While it is held, anything else that execs into that instance queues behind it: an occ
command, a NextcloudCommand, a maintenance task. Those paths already retry, so what
you will see is a TemporaryError and a retry in the operator log, not a failure — the
work happens once the lock frees.
This is a deliberate trade, not an oversight. Two concurrent execs into one Nextcloud
pod can interleave, and occ is not safe to run concurrently with itself; a corrupted
occ run during a backup would be a far worse outcome than a few minutes of queueing.
If a bundle's timing is inconvenient for a particular instance, move its slot rather
than working around the lock — see the daily jittered hour above.
Trigger a configuration backup now¶
There is no dedicated annotation. Clear the daily marker and force a reconcile:
kubectl -n "$NS" patch nci "$INST" --subresource=status --type=merge \
-p '{"status":{"configBackup":{"lastSuccess":null}}}'
kubectl -n "$NS" annotate nci "$INST" k8s.bnerd.com/reconcile="$(date +%s)" --overwrite
The next maintenance tick uploads a bundle. This does not delete anything: the previous bundles stay in the bucket until retention prunes them.
Rotate the backup credentials or the encryption key¶
kubectl -n nextcloud-operator-system create secret generic db-backup-s3 \
--from-literal=accessKey='...' --from-literal=secretKey='...' \
--dry-run=client -o yaml | kubectl apply -f -
Every instance namespace re-syncs its derived <instance>-pgbackrest-s3 Secret on the
next reconcile tick — the operator compares a content hash, so nothing else is rewritten
and no repo-host pod is rolled unnecessarily.
Keep retired encryption keys
Rotating db-backup-cipher means new backups use the new key while existing
backups still need the old one. Retain every retired key at least as long as the
backups it can open, and record which key covers which date range.
Opt one instance out¶
spec:
database:
postgres:
backup:
enabled: false # no repo2, no derived Secret, no bootstrap
backups:
config:
enabled: false # no configuration bundle
database.postgres.backup.enabled: false wins over a centrally-configured bucket: a
fleet default never resurrects backups on an instance that deliberately opted out, and
no credentials Secret is written into its namespace.
When a configuration backup fails¶
kubectl -n "$NS" get nci "$INST" -o jsonpath='{.status.configBackup.lastFailureReason}{"\n"}'
kubectl -n "$NS" get events --field-selector reason=ConfigBackupFailed
The reason is redacted of credentials before it is stored. Common causes:
| Reason contains | Cause | Fix |
|---|---|---|
config/ not found |
No config/ in the pod — the instance is not installed yet |
Wait for install to complete |
tar exited |
The pod refused to read a path | Check pod permissions; check the paths still exist |
AccessDenied |
Bucket credentials or policy | Verify the central Secret and the bucket policy |
not found (yet) |
The central Secret has not synced | Check the landscape sync for the operator namespace |
exceeded the ... byte budget |
Only a warning, not a failure — custom_apps//themes/ were dropped |
Slim those directories, or accept the smaller bundle |
A failure never clears configBackup.lastSuccess/lastKey: the last good bundle is
still in the bucket.
Operator-Managed Labels (read-only)¶
The operator sets these labels on pool instances and related resources. They are useful for kubectl -l selectors and dashboards — do not edit them by hand.
| Label | Set on | Purpose |
|---|---|---|
k8s.bnerd.com/managed-by |
NextcloudInstance |
Always nextcloud-operator for pool-created instances |
k8s.bnerd.com/pool |
NextcloudInstance |
Name of the NextcloudPool that created it |
k8s.bnerd.com/assigned |
NextcloudInstance |
true / false — whether the instance is claimed by a Nextcloud |
k8s.bnerd.com/nextcloud |
NextcloudInstance |
Name of the assigned Nextcloud (when assigned) |
k8s.bnerd.com/nextcloud-ns |
NextcloudInstance |
Namespace of the assigned Nextcloud |
k8s.bnerd.com/profile |
NextcloudInstance |
Profile the pool template referenced |
Example queries:
# All unassigned instances in a pool
kubectl get nci -A -l k8s.bnerd.com/pool=my-pool,k8s.bnerd.com/assigned=false
# All instances assigned to a specific Nextcloud
kubectl get nci -A -l k8s.bnerd.com/nextcloud=my-tenant
Operator-Managed Annotations (read-only)¶
The operator writes these annotations to record credential provenance. Unlike the annotations in the Summary table above, you don't set these yourself — do not edit them by hand.
| Annotation | Value | Set on | Purpose |
|---|---|---|---|
k8s.bnerd.com/admin-credential |
generated |
NextcloudInstance |
The instance's admin credential (spec.admin.credentialsSecret) was operator-generated rather than user-supplied — see Admin Credential Management → Generated credentials. Written the same reconcile the credentialsSecret ref is written back. |
k8s.bnerd.com/inline-admin-advisory |
sent |
NextcloudInstance |
The one-time InlineAdminPasswordDeprecated advisory event has already fired for this instance (non-pool instance with an inline spec.admin.password). Prevents the advisory from repeating on every reconcile. |