Skip to content

Operations & Annotations

This guide lists the annotations and labels you can set on managed resources to drive day-2 operations — forcing a reconciliation, triggering maintenance on demand, or overriding deletion protection. It also documents the labels the operator sets itself so you know what is safe to rely on in kubectl -l selectors.

All keys use the k8s.bnerd.com/ prefix.

Summary

Key Kind Resource Purpose
k8s.bnerd.com/reconcile annotation Nextcloud, NextcloudInstance Force an immediate reconciliation
k8s.bnerd.com/run-maintenance annotation NextcloudInstance Run OCC maintenance tasks immediately
k8s.bnerd.com/force-delete label NextcloudInstance Allow deletion of an assigned pool instance
k8s.bnerd.com/allow-unsafe-version-change annotation NextcloudInstance Bypass upgrade-path validation (downgrade / major-skip) for a restore or rollback
k8s.bnerd.com/upgrade-now annotation NextcloudInstance Apply a pending version update immediately, bypassing the maintenance window and upgradePolicy.mode
k8s.bnerd.com/upgrade-apps-accept annotation NextcloudInstance Accept an upgrade whose declared apps are still disabled: lift the maintenance hold and stop the convergence retry loop
k8s.bnerd.com/upgrade-resync annotation NextcloudInstance Re-derive status.upgrade from the instance's live state; clears bookkeeping that no longer describes it. Lifts maintenance mode only if one of the operator's own upgrade windows is recorded as holding it — anything else is reported (MaintenanceModeNotOurs) and left alone. Spent automatically once it has acted

Force Reconcile

Annotation: k8s.bnerd.com/reconcile Applies to: Nextcloud, NextcloudInstance Value: any string; convention is an ISO 8601 timestamp or date +%s. Only changes to the value trigger a reconcile — setting the same value twice is a no-op.

Normal reconciliation runs whenever a relevant spec field changes or on the periodic timer (30 s for NextcloudInstance, 60 s for Nextcloud). Use this annotation to trigger a reconcile immediately in cases where the operator would not otherwise notice a change.

When to use

  • After editing a NextcloudProfile — existing instances don't pick up profile changes automatically
  • To retry provisioning after a transient error (managed database creation, S3 bucket auto-creation, HelmRelease failure)
  • After rotating a referenced secret (credentialsSecret) so the new values are read and re-applied
  • To propagate a Nextcloud spec change to its assigned NextcloudInstance if the sync appears stuck

How to trigger

# Force reconcile a NextcloudInstance
kubectl annotate nci my-instance \
  k8s.bnerd.com/reconcile=$(date -u +"%Y-%m-%dT%H:%M:%SZ") \
  --overwrite -n my-namespace

# Force reconcile a Nextcloud (logical resource)
kubectl annotate nc my-tenant \
  k8s.bnerd.com/reconcile=$(date +%s) \
  --overwrite -n my-namespace

A ready-made helper script is shipped with the repository:

./examples/trigger-manual-reconcile.sh nci my-instance my-namespace

Watch the operator logs to confirm the trigger fired — you should see Manual reconciliation triggered for ....

On-Demand Maintenance

Annotation: k8s.bnerd.com/run-maintenance Applies to: NextcloudInstance Value: any string; use a fresh timestamp on each run.

By default, periodic OCC maintenance (file cleanup, missing-index checks, etc.) runs once per day during the window configured in spec.maintenance.maintenanceWindow. Set this annotation to run the same tasks immediately, outside the window.

This does not re-run post-upgrade tasks (those are tied to a version change). It runs whichever periodic tasks are enabled in spec.maintenance.tasks.

When to use

  • You want to reclaim disk space or tidy up orphaned files right now without waiting for the window
  • You just enabled a new task in spec.maintenance.tasks and want to run it once straight away
  • You are investigating an issue and want fresh output from db:add-missing-indices or similar

How to trigger

kubectl annotate nci my-instance \
  k8s.bnerd.com/run-maintenance=$(date +%s) \
  --overwrite -n my-namespace

After the run completes, status.maintenance.lastRunTrigger is set to annotation and status.maintenance.lastRunAt is updated:

kubectl get nci my-instance -o jsonpath='{.status.maintenance}' | jq

See the API reference for the full list of tasks and timeout knobs under spec.maintenance.

For arbitrary occ commands (not just the canned maintenance tasks), use the NextcloudCommand CRD. Where run-maintenance triggers a fixed task set, NextcloudCommand lets you declaratively run any occ invocation with per-command result reporting.

All occ execs against a given instance — triggered by this annotation, NextcloudCommand, admin-credential apply, or the upgrade/apps-health checks — serialize automatically per instance (0.20.0+); see Running occ Commands → Limitations for the exact guarantee and the lock-timeout fallback.

Force Delete

Label: k8s.bnerd.com/force-delete=true Applies to: NextcloudInstance

Note: This is a label, not an annotation. Use kubectl label rather than kubectl annotate.

Pool-provisioned NextcloudInstance resources carry a finalizer (k8s.bnerd.com/assigned-instance-protection) and labels linking them to their assigned Nextcloud. Deleting an assigned instance directly is normally blocked with a TemporaryError so that the Nextcloud doesn't end up pointing at a gone backend.

Set k8s.bnerd.com/force-delete=true on the instance to bypass that protection. Use as a last-resort escape hatch: this will leave the assigned Nextcloud in a broken state until you re-point it at another instance or delete it.

When to use

  • The instance is hard-stuck in a failed state and needs to be removed before the Nextcloud can be re-assigned
  • You are manually draining a pool and intend to delete both the Nextcloud and the instance

How to trigger

# Check what the instance is assigned to first
kubectl get nci my-instance -o jsonpath='{.metadata.labels}' | jq

# Override protection and delete
kubectl label nci my-instance k8s.bnerd.com/force-delete=true -n my-namespace
kubectl delete nci my-instance -n my-namespace

The preferred path is to delete the owning Nextcloud first — that clears the assignment and the instance deletes cleanly without the force label.

Upgrade-Path Override

Annotation: k8s.bnerd.com/allow-unsafe-version-change Applies to: NextcloudInstance Value: "true" to bypass; remove (or unset) to restore normal protection.

The operator refuses to resolve spec.version to a downgrade or a major-version skip (e.g. 30 → 32 in one hop) — see Upgrades → Upgrade-path validation. This annotation is the escape hatch, intended only for documented restore/rollback procedures where you're intentionally setting spec.version back to a pre-upgrade value after restoring backups. Remove it once the instance is stable; leaving it set disables the protection for every subsequent reconcile, not just the one you needed it for. Full procedure: Upgrades → Backup & rollback.

On-Demand Upgrade

Annotation: k8s.bnerd.com/upgrade-now Applies to: NextcloudInstance Value: any string; use a fresh timestamp on each trigger.

Applies the currently-resolved version update immediately, bypassing both spec.maintenance.windowStart and spec.upgradePolicy.mode entirely. See Upgrades → On-demand upgrades for the full behavior (drift-class bypass, upgrade-path validation still applies, the NoUpdateAvailable no-op, the annotation-vs-timer race handling).

When to use

  • A mode: manual instance has UpdateAvailable: True and you want it applied right now, without switching the instance to mode: auto and waiting for a window.
  • An auto-mode instance's maintenance window hasn't arrived yet but you need the update applied sooner (an out-of-band security patch, a coordinated maintenance slot).
  • You want to apply a resolved patch bump to one specific instance without waiting for its next scheduled window tick.

What to expect

  • If nothing is pending (UpdateAvailable: False or absent), this is a no-op — a reason: NoUpdateAvailable event fires and nothing else happens.
  • If something is pending, the operator bumps k8s.bnerd.com/reconcile internally to trigger an immediate reconcile with the resolved target — the same HelmRelease update / Flux rollout / post-upgrade task sequence as any other upgrade (see Upgrades → How an upgrade happens).
  • A resolved target that's a major skip or downgrade is still held or blocked by upgrade-path validation — this annotation does not bypass that; you still need k8s.bnerd.com/allow-unsafe-version-change on top for a documented rollback.

How to trigger

kubectl annotate nci my-instance \
  k8s.bnerd.com/upgrade-now=$(date -u +"%Y-%m-%dT%H:%M:%SZ") \
  --overwrite -n my-namespace

How to verify

# Confirm there was something to apply, before triggering
kubectl get nci my-instance -n my-namespace \
  -o jsonpath='{.status.conditions[?(@.type=="UpdateAvailable")]}' | jq

# After triggering, watch the phase and the resolved version update
kubectl get nci my-instance -n my-namespace -o jsonpath='{.status.phase}'
kubectl get nci my-instance -n my-namespace -o jsonpath='{.status.versionResolution}' | jq

# If it no-op'd, confirm why
kubectl get events -n my-namespace --field-selector involvedObject.name=my-instance | grep -i upgrade

Accepting a stalled post-upgrade app convergence

Annotation: k8s.bnerd.com/upgrade-apps-accept Applies to: NextcloudInstance Value: "true".

Since 0.23.0 the operator holds Nextcloud's maintenance mode across a version upgrade until it has verified that every declared app is enabled again, and keeps holding it when a critical app (user_oidc, while spec.oidc.enabled) cannot be enabled. It retries indefinitely, because the usual cause — an app store that has not yet published a release for the new major — resolves on its own within days. See Upgrades → Post-upgrade app convergence.

This annotation is the human override for "I know, stop holding this instance."

When to use

  • SSO is broken by a pending user_oidc, and you would rather have the instance up with local logins than down with a maintenance page.
  • The affected app is not coming back and you are not ready to edit spec.apps yet.

What to expect

  • Maintenance mode is lifted, status.upgrade.phase becomes Completed with accepted: true, and the nextcloud_operator_upgrade_apps_pending series is dropped.
  • AppsHealthy stays False. Accepting is not a claim of health — the degraded app state stays visible until it actually changes.
  • The retry loop stops. If you want the operator to try again afterwards, trigger a reconcile (k8s.bnerd.com/reconcile) after fixing the underlying cause, or let the next version change open a fresh window.
  • If the operator cannot reach the pod to lift the window, the acceptance is recorded but the phase stays AppsPending and the lift is retried — you will not get a Completed on an instance that is still dark.

How to trigger

kubectl annotate nci my-instance \
  k8s.bnerd.com/upgrade-apps-accept=true --overwrite -n my-namespace

How to verify

kubectl get nci my-instance -n my-namespace -o jsonpath='{.status.upgrade}' | jq '{phase, accepted, maintenanceHeld, pendingApps}'

Remove the annotation once the situation is resolved. It applies to the upgrade that was in flight when you applied it: an annotation still on the CR when a later upgrade opens its window is treated as left over and ignored (with a StaleUpgradeAcceptIgnored warning event), so a forgotten one cannot abort the next upgrade — but while it stays there you also cannot accept a later upgrade with it. Remove it, and re-apply it if and when you mean to accept a different one.

Authorisation: this is a plain CR annotation write (patch on nextcloudinstances in the instance's namespace) and it overrides a protection applied to the tenant — see Upgrades → Escape hatch for why proxying it from a self-service surface is a privilege decision rather than a convenience.

Admin Credential Management

As of 0.20.0 (#7866), spec.admin.{username,password} (inline or via credentialsSecret) is not just recorded on the NextcloudInstance — the operator actively applies it to the live Nextcloud instance. This is the "Rails-as-source-of-truth" pattern: a control plane (or any other CR author) declares the credential once on the CR, and can rely on the operator to make that credential work against the running instance, rather than separately tracking whether it was ever installed.

Trigger points

The operator converges the live admin credential to spec.admin from three places, all funneling into the same idempotent apply step:

  1. PoolAssignment — see Pool Provisioning → Admin Credential Handoff at Assignment.
  2. A spec.admin change on an already-assigned instance (a spec.admin field watcher).
  3. The periodic instance-status timer — a safety net that re-checks on every reconcile tick so a missed event still converges.

How application works

  • The desired credential is resolved via the same credentialsSecret/inline loader used elsewhere (admin.credentialsSecret, or admin.username / admin.password directly).
  • Precedence (0.21.5): a credentialsSecret you supply outranks an inline password, as documented. The operator's own <instance>-nextcloud-admin Secret does not: when admin.credentialsSecret names that Secret and admin.password is non-empty, the inline password wins, because that Secret is a materialization of the desired credential rather than a source of truth. See Credentials guide → Changing a declared admin password for the full table and the rewrite-then-blank sequence.
  • The literal changeme is never treated as a declared password (0.21.5). The NextcloudInstance CRD defaulted spec.admin.password to it until this release, and the API server injects a schema default into every admin object that omits the key. On 0.21.0–0.21.4 that made the literal an instance's real admin password; 0.21.5 refuses it in the resolver, in both auto-generate gates and in the shadow rule, so an old separately vendored CRD cannot reintroduce it from the spec. It does not rotate a password already provisioned that way — that value lives in the admin Secret, which is the resolved credential, so it keeps being applied until someone rotates it. Audit the <instance>-nextcloud-admin Secrets.
  • The blank-back-out is conflict-safe: the operator re-reads the live spec and clears only the exact password it just applied, via a JSON-Patch test op. A password changed again mid-apply is never swallowed — it wins the next reconcile.
  • Existence is probed with occ user:info <username>:
  • Exists → occ user:resetpassword resets the password.
  • Missing → occ user:add --group admin creates the user (create-if-missing).
  • The password is transported to the pod exec's stdin only — never in command-line arguments, never written to any CR, and never logged (asserted by tests). It is not recoverable from kubectl describe, operator logs, or Kubernetes events.
  • The previous admin account is never deleted, renamed, or disabled. If spec.admin.username changes, the operator creates the new account and leaves the old one exactly as it was — cleanup, if wanted, is a manual step.

Idempotency

status.appliedAdminCredential records {username, passwordHash, appliedAt} for the credential last applied to the live instance. passwordHash is PBKDF2-HMAC-SHA256 (600,000 iterations) of the password, salted with the instance UID — self-describing format pbkdf2-sha256$600000$<hex>. It is never the plaintext password, only a drift-detection fingerprint. The apply step no-ops (no exec, no Secret rewrite) once username and passwordHash both already match what's resolved from spec.admin.

Failure surface

A condition AdminCredentialApplied is set on the instance:

reason status Meaning
Applied True Password reset succeeded against an existing admin account.
UserCreated True The declared username didn't exist on the live instance; occ user:add --group admin created it.
ExecFailed False The occ exec failed (pod not reachable, command error, etc.). A Kubernetes warning event is also emitted. The operator retries (kopf.TemporaryError, never a permanent failure).
CredentialsUnavailable False The declared credential couldn't be resolved — most often a spec.admin.credentialsSecret that doesn't exist yet. A Kubernetes warning event is also emitted, and the operator retries indefinitely (never a permanent failure), so a late-arriving Secret converges on its own.

The condition is re-asserted on every attempt, so it always reflects the current state. The event is not: the operator emits one Warning per distinct failure, not one per reconcile tick, so a persistently broken instance doesn't bury its own timeline. A rotated credential that also fails counts as a new failure and gets its own event. The marker lives in status.adminCredentialLastNotifiedFailure and is cleared once the credential applies (or already matches), so nothing is ever silently suppressed.

Failure messages now include the underlying error (0.21.1). Both ExecFailed and CredentialsUnavailable condition messages fold in a sanitized snippet of the actual exception — never raw secret material, redaction is tested — instead of a generic "failed to apply" message, so the real cause (a specific occ exec error, a missing Secret key, a broken mail credentialsSecret) is visible from kubectl describe nci without a manual pod exec. A broken spec.mail.credentialsSecret also surfaces through this same path: it's read here to build the admin Secret's mail_config, and before 0.21.1 a failure there escaped as a raw, unsurfaced exception with no condition or event at all.

Safety-net logging raised to WARNING (0.21.1). The instance-status timer's catch-all around this apply step now logs at WARNING once per failure onset and DEBUG for repeats of the same failure — bounded by the same adminCredentialLastNotifiedFailure dedupe marker above, not a second mechanism. Before 0.21.1 every failure logged at DEBUG, invisible at the deployed INFO level.

Hash-match no-op path now self-heals a missing condition (0.21.1). A rare race (a same-tick sibling patch seeding conditions from a stale snapshot) could drop the condition on the credential's first successful apply while status.appliedAdminCredential was still correctly recorded — a healthy instance with no visible AdminCredentialApplied condition at all, and no later tick ever rewrote it. The no-op path now re-asserts AdminCredentialApplied=True/Applied whenever the ambient condition is absent or not True, closing the gap within one reconcile tick; when the condition already reads True, the tick remains a genuine zero-write no-op.

An instance that isn't Ready yet has no pod to exec into, so there is nothing to apply to and no condition is written. How that's handled depends on which trigger point noticed:

  • The spec.admin field watcher returns and lets the periodic instance-status timer pick the work up on the next Ready tick. It does not retry: spec.admin exists on effectively every instance, so this watcher runs as part of kopf's create cycle for the object, and a handler that retries forever holds that cycle open — which in turn stops kopf from dispatching watch-driven handlers (a spec.version edit going unreconciled) and from removing its own finalizer during deletion (an instance stuck Terminating long after the operator's own cleanup finished). Fixed in 0.21.0.
  • The instance-status timer only calls the apply step for an instance it has just confirmed Ready, so the credential converges within one tick of the instance becoming usable.

An instance carrying a metadata.deletionTimestamp is skipped entirely, from every trigger point — nothing is applied to an object being torn down.

Generated admin credentials and the 0.21.0 migration

This section is about credential hygiene (what's stored where), which is a different concern from the live-apply behavior above (what's active on the running instance). See the Credentials guide → Generated Admin Credentials for the full mechanism; summary:

  • Since 0.21.0, a generated admin password is stored only in the <instance>-nextcloud-admin Secret. The CR gets spec.admin.credentialsSecret written back, never the plaintext, and is marked with the k8s.bnerd.com/admin-credential: generated annotation.
  • On the first reconcile after upgrading to 0.21.0, pool-origin instances whose spec still carries a plaintext generated password (from before this change) are migrated in place to the ref form — one MigratedAdminCredential event, no live credential change (the hash-gate no-ops, so no occ exec runs).
  • Non-pool instances with an inline password are left untouched, with a one-time InlineAdminPasswordDeprecated advisory event (see the annotations table above for the gating annotation).

Runbook: rotate the admin password of a pool-provisioned instance

A pool-provisioned instance's spec carries the operator's own admin Secret ref and an empty password. To set a specific password:

# 1. Declare it. Usually on the parent Nextcloud CR (it passes spec.admin
#    straight through to the instance); a direct edit of the NextcloudInstance
#    works the same way. Declaring it up front, on a Nextcloud CR that is about
#    to be pool-assigned, goes through the same path.
kubectl patch nc <name> -n <ns> --type merge \
  -p '{"spec":{"admin":{"password":"<new-password>"}}}'

# 2. Watch status.appliedAdminCredential — this is the field that moves.
#    (Seconds, not minutes: the instance's spec.admin watcher fires on the
#    patch itself; the 30s instance-status timer is only the safety net.)
kubectl get nci <instance> -n <ns> -o jsonpath='{.status.appliedAdminCredential}'
# -> appliedAt advances and passwordHash changes once the new credential is live

# 3. OPTIONAL, and not a pass/fail gate: the operator blanks the plaintext back
#    out of the instance spec after applying it.
kubectl get nci <instance> -n <ns> -o jsonpath='{.spec.admin}'
# -> {"credentialsSecret":"<instance>-nextcloud-admin","password":""}
#    ...but if the PARENT Nextcloud CR still declares the password, the next spec
#    passthrough (within ~30s) puts it back. Seeing the plaintext here again is
#    expected in that case and does NOT mean the apply failed — step 2 is the gate.
#    It is also skipped entirely for a GitOps-managed CR (Flux owns the spec).

# 4. The Secret now holds the new password.
kubectl get secret <instance>-nextcloud-admin -n <ns> \
  -o jsonpath='{.data.nextcloud-password}' | base64 -d

Don't verify this with AdminCredentialApplied's lastTransitionTime

On an instance that already had a credential applied — every pool-provisioned instance, which got its generated password applied at provisioning time — the condition goes True → True. Per the Kubernetes convention the operator follows, lastTransitionTime is preserved across a same-status update, so it does not move even though the new password was applied. Use status.appliedAdminCredential.appliedAt / passwordHash (step 2), the admin Secret's resourceVersion, or an actual login. An unchanged condition timestamp is not evidence of failure.

If the credential does not apply, the condition flips to False — read its reason and message (kubectl describe nci <instance>): ExecFailed, CredentialsUnavailable and SecretWriteFailed all carry a sanitized snippet of the underlying error. SecretWriteFailed means the operator could not even rewrite the owned <instance>-nextcloud-admin Secret, so the live occ apply was never attempted. The instance must be Ready for the credential to be applied at all — there is no pod to occ into otherwise.

To stop the plaintext reappearing (step 3), remove spec.admin.password from the parent Nextcloud CR once step 2 shows the new credential applied; resolution then continues through the Secret.

Upgrade note: first reconcile after 0.20.0 resets drifted passwords fleet-wide

status.appliedAdminCredential is new in 0.20.0, so every already-Ready instance starts with no recorded applied credential. On the very next periodic instance-status timer tick after upgrading, every Ready instance in the fleet gets its declared spec.admin applied — unconditionally, not staged or throttled. For an instance whose live password already matches spec.admin.password, this is a harmless extra occ exec. For an instance where the two have drifted (exactly the scenario this feature exists to fix), the live admin password will be reset to the CR's declared value at that moment, with no separate trigger or advance notice. This is the feature working as intended, not a bug — but plan for it before upgrading a fleet where spec.admin may not reflect the current live password: audit for drift first if that matters for your rollout, since the change happens automatically on upgrade rather than on the next spec.admin edit.

Consumer contract: the admin Secret can briefly diverge from the live instance

The operator syncs the operator-owned <instance>-nextcloud-admin Secret to the desired credential before attempting the live occ exec. If the exec then fails, the Secret already reflects the new password while the live instance's actual password is still the old one — AdminCredentialApplied flips to False/ExecFailed and a warning event fires, so the failure is visible, but only to something watching the CR's condition or events, not to something that only reads the Secret.

Anything that reads this Secret as its source of truth for "what credential currently works against the live instance" (e.g. an external control plane) must check AdminCredentialApplied=True before trusting it — the Secret being up to date does not by itself guarantee the live instance matches it.

Profile Propagation at Upgrade

Upgrading the operator to 0.20.0 does not, by itself, retroactively reconfigure any already-running instance — see Profiles → When Profiles Apply for the full mechanism (profile defaults apply once, at instance creation, by design). Concretely, for an existing profile-backed instance:

  • The spec write-back is unchanged, still and forever create-time-only. The 0.20.0 deep-merge fix, the correct hooks wholesale-replace behavior, and the wider spec write-back only take effect at instance creation — they apply to instances created after the 0.20.0 operator is live, not to ones that already exist. spec.database/spec.s3/etc. on an existing instance stay exactly whatever was persisted under the previous (possibly incorrectly-merged) logic. This part has not changed since 0.20.0.
  • Rendered Helm values are a different story as of 0.21.1 (review finding E-1). Before 0.21.1, rendering also skipped re-resolving the profile once an instance was materialized, so the previous bullet's "unchanged" claim held for Helm values too. 0.21.1 removed that skip — the profile now resolves fresh on every render, though the instance's own materialized value (from the write-back above, or your own spec.helm.values) still wins the merge and is never overridden. Net effect for an already-running instance: a profile edit to a key it already has a value for still never propagates; a key added to the profile since creation now renders on the next reconcile; and an instance whose write-back never ran at all (created before operator v0.10.2, or GitOps-managed) now gets its entire profile-derived Helm values rendered correctly, where before they silently never appeared. See Profiles → When Profiles Apply for the full contract. Rollout note: on the first reconcile after upgrading to 0.21.1, any instance whose profile gained keys since it was created — or whose spec write-back never ran — renders those keys for the first time, which changes that instance's HelmRelease and restarts its pod(s) once.
  • Managed-PostgreSQL is one exception. On the next reconcile after upgrading, a managed-PostgreSQL, profile-backed instance's live PerconaPGCluster pgBouncer configuration is corrected to match the current profile — regardless of whether spec.database itself was ever backfilled, and regardless of GitOps management. See Profiles → When Profiles Apply → Exception.
  • spec.mail gets a one-time heal in 0.21.0. Instances that materialized before 0.20.0's merge-kernel fix can be permanently missing profile-provided mail fields (most notably fromAddress — #7870). Starting in 0.21.0, the periodic instance-status timer gap-fills any such missing keys from the instance's applied profile, once, for non-GitOps instances — see Profiles → When Profiles Apply → One-time heal. Unlike pgBouncer, this is not a continuous sync: after the one-time fill, spec.mail reverts to the normal once-only rule above. This write changes rendered Helm values, so the affected instance's Nextcloud pod(s) restart once the next time this runs after upgrading to 0.21.0.
  • New instances get the full fix from creation onward: correct deep-merge, wider write-back (skipped for GitOps-managed CRs — logged instead, not silently dropped), hooks wholesale-replace, and the continuously profile-aware pgBouncer reconciler.

Don't confuse this with admin credential application above, which behaves the opposite way: spec.admin is actively re-applied to every already-Ready instance's live Nextcloud on the first reconcile after upgrading. Profile-derived configuration and admin credentials follow different rules in 0.20.0 — only admin credentials and the pgBouncer proxy self-heal on upgrade; everything else stays as-is.

HelmRelease Value Ownership

Starting in 0.21.1, the operator owns the managed HelmRelease's spec.values wholesale: every update rebuilds the complete values object from the instance's spec and profile, and replaces the live spec.values with it entirely (a JSON Patch add, not a merge). Previously, an update merge-PATCHed the live object, so any key present on the HelmRelease but absent from the freshly-built values was left untouched — a stale value from an earlier spec revision, or a value someone hand-edited on the HelmRelease directly, could silently persist forever. A manual edit to a managed HelmRelease's spec.values is removed on the instance's next reconcile. Manage the underlying setting through the NextcloudInstance/Nextcloud spec (or a NextcloudProfile), not by hand-editing the rendered HelmRelease — the same rule that already applied to every other operator-managed field, now also true for spec.values.

Only the fields the operator itself sets are affected: spec.interval, spec.timeout, spec.chart, spec.values, spec.install, and spec.upgrade. Anything else on the HelmRelease survives untouched, because a JSON Patch only touches the exact paths it lists — most notably:

  • spec.suspend (set via flux suspend hr <name> -n <namespace>) is never one of the operator's ops, so pausing a HelmRelease for an incident survives every subsequent operator reconcile until you resume it (flux resume hr).
  • Any other field Flux or a different controller manages on the object (dependsOn, postRenderers, ...) is likewise untouched.

This closed a real incident: a stale cert-manager.io/cluster-issuer ingress annotation left behind by a spec.ingress.tls.certManager: false flip caused cert-manager to keep issuing a certificate into the instance's wildcard TLS Secret, clobbering it on every renewal. See the CHANGES.md v0.21.1 entry for the full writeup.

Backup Runbook (0.22.0)

Day-to-day operational procedures for the two backup halves. The full picture, including restore, lives in Backup & Restore.

Enable off-cluster backups for the whole estate

One change to the operator's own Helm release plus two Secrets in its namespace — no edit to any instance, profile or CRD.

NS=nextcloud-operator-system

# 1. S3 credentials for the backup bucket
kubectl -n "$NS" create secret generic db-backup-s3 \
  --from-literal=accessKey='...' --from-literal=secretKey='...'

# 2. The backup-encryption passphrase. Put this in the password manager TOO --
#    a backup encrypted with a key you have lost is noise.
openssl rand -base64 48 > cipher.key
kubectl -n "$NS" create secret generic db-backup-cipher --from-file=cipher-pass=cipher.key

# 3. Point the operator at the bucket
helm upgrade nextcloud-operator ./chart -n "$NS" --reuse-values \
  --set backup.dbDefaultS3.bucket=nextcloud-db-backups \
  --set backup.dbDefaultS3.endpoint=s3.example.com \
  --set backup.dbDefaultS3.region=eu-central-1 \
  --set backup.dbDefaultS3.secretRef=db-backup-s3 \
  --set backup.dbDefaultS3.cipherSecretRef=db-backup-cipher

Roll it out to one instance first (set the values, confirm the checks below, then let the rest of the estate follow on its own reconcile ticks).

The operator needs a writable working directory

Config bundles are staged on disk before upload. The operator runs with readOnlyRootFilesystem: true, so the chart and deploy/operator.yaml both mount an emptyDir at /tmp (workDir.sizeLimit, default 1Gi) with matching ephemeral-storage requests and limits. If you render your own Deployment or strip volumes, keep that mount — without it every bundle fails with a permission error and status.configBackup.lastFailureReason will say so.

Peak usage is roughly 3x the bundle size (streamed parts, the assembled tar and the encrypted copy exist together), so raise both the sizeLimit and the ephemeral-storage limit together if you raise the bundle guard.

Verify an instance picked it up

NS=<instance-namespace>
INST=<instance-name>

# repo2 exists and carries a PER-INSTANCE path
kubectl -n "$NS" get perconapgcluster "$INST-pg" \
  -o jsonpath='{.spec.backups.pgbackrest.repos[*].name}{"\n"}{.spec.backups.pgbackrest.global.repo2-path}{"\n"}'

# the derived credentials Secret was materialized
kubectl -n "$NS" get secret "$INST-pgbackrest-s3"

# the one-time bootstrap full was triggered
kubectl -n "$NS" get nci "$INST" -o jsonpath='{.status.backupRepo2Bootstrapped}{"\n"}'
kubectl -n "$NS" get perconapgbackup -o custom-columns=NAME:.metadata.name,REPO:.spec.repoName,STATE:.status.state

# the configuration bundle landed
kubectl -n "$NS" get nci "$INST" -o jsonpath='{.status.configBackup}' | jq

Expected: repo1 repo2, a repo2-path of /<namespace>/<instance>/backups/db, a BackupRepo2Bootstrapped event, and a ConfigBackupCompleted event within a day.

A running bundle holds the instance's exec lock

The operator serializes pod execs per instance (the same lock occ commands take), and a config-bundle run holds it for the whole run — the path probe, the tar stream and the upload. On a large instance that can be a few minutes.

While it is held, anything else that execs into that instance queues behind it: an occ command, a NextcloudCommand, a maintenance task. Those paths already retry, so what you will see is a TemporaryError and a retry in the operator log, not a failure — the work happens once the lock frees.

This is a deliberate trade, not an oversight. Two concurrent execs into one Nextcloud pod can interleave, and occ is not safe to run concurrently with itself; a corrupted occ run during a backup would be a far worse outcome than a few minutes of queueing. If a bundle's timing is inconvenient for a particular instance, move its slot rather than working around the lock — see the daily jittered hour above.

Trigger a configuration backup now

There is no dedicated annotation. Clear the daily marker and force a reconcile:

kubectl -n "$NS" patch nci "$INST" --subresource=status --type=merge \
  -p '{"status":{"configBackup":{"lastSuccess":null}}}'
kubectl -n "$NS" annotate nci "$INST" k8s.bnerd.com/reconcile="$(date +%s)" --overwrite

The next maintenance tick uploads a bundle. This does not delete anything: the previous bundles stay in the bucket until retention prunes them.

Rotate the backup credentials or the encryption key

kubectl -n nextcloud-operator-system create secret generic db-backup-s3 \
  --from-literal=accessKey='...' --from-literal=secretKey='...' \
  --dry-run=client -o yaml | kubectl apply -f -

Every instance namespace re-syncs its derived <instance>-pgbackrest-s3 Secret on the next reconcile tick — the operator compares a content hash, so nothing else is rewritten and no repo-host pod is rolled unnecessarily.

Keep retired encryption keys

Rotating db-backup-cipher means new backups use the new key while existing backups still need the old one. Retain every retired key at least as long as the backups it can open, and record which key covers which date range.

Opt one instance out

spec:
  database:
    postgres:
      backup:
        enabled: false     # no repo2, no derived Secret, no bootstrap
  backups:
    config:
      enabled: false       # no configuration bundle

database.postgres.backup.enabled: false wins over a centrally-configured bucket: a fleet default never resurrects backups on an instance that deliberately opted out, and no credentials Secret is written into its namespace.

When a configuration backup fails

kubectl -n "$NS" get nci "$INST" -o jsonpath='{.status.configBackup.lastFailureReason}{"\n"}'
kubectl -n "$NS" get events --field-selector reason=ConfigBackupFailed

The reason is redacted of credentials before it is stored. Common causes:

Reason contains Cause Fix
config/ not found No config/ in the pod — the instance is not installed yet Wait for install to complete
tar exited The pod refused to read a path Check pod permissions; check the paths still exist
AccessDenied Bucket credentials or policy Verify the central Secret and the bucket policy
not found (yet) The central Secret has not synced Check the landscape sync for the operator namespace
exceeded the ... byte budget Only a warning, not a failure — custom_apps//themes/ were dropped Slim those directories, or accept the smaller bundle

A failure never clears configBackup.lastSuccess/lastKey: the last good bundle is still in the bucket.

Operator-Managed Labels (read-only)

The operator sets these labels on pool instances and related resources. They are useful for kubectl -l selectors and dashboards — do not edit them by hand.

Label Set on Purpose
k8s.bnerd.com/managed-by NextcloudInstance Always nextcloud-operator for pool-created instances
k8s.bnerd.com/pool NextcloudInstance Name of the NextcloudPool that created it
k8s.bnerd.com/assigned NextcloudInstance true / false — whether the instance is claimed by a Nextcloud
k8s.bnerd.com/nextcloud NextcloudInstance Name of the assigned Nextcloud (when assigned)
k8s.bnerd.com/nextcloud-ns NextcloudInstance Namespace of the assigned Nextcloud
k8s.bnerd.com/profile NextcloudInstance Profile the pool template referenced

Example queries:

# All unassigned instances in a pool
kubectl get nci -A -l k8s.bnerd.com/pool=my-pool,k8s.bnerd.com/assigned=false

# All instances assigned to a specific Nextcloud
kubectl get nci -A -l k8s.bnerd.com/nextcloud=my-tenant

Operator-Managed Annotations (read-only)

The operator writes these annotations to record credential provenance. Unlike the annotations in the Summary table above, you don't set these yourself — do not edit them by hand.

Annotation Value Set on Purpose
k8s.bnerd.com/admin-credential generated NextcloudInstance The instance's admin credential (spec.admin.credentialsSecret) was operator-generated rather than user-supplied — see Admin Credential Management → Generated credentials. Written the same reconcile the credentialsSecret ref is written back.
k8s.bnerd.com/inline-admin-advisory sent NextcloudInstance The one-time InlineAdminPasswordDeprecated advisory event has already fired for this instance (non-pool instance with an inline spec.admin.password). Prevents the advisory from repeating on every reconcile.