Deletion & Cleanup¶
This guide describes how the operator tears down a Nextcloud instance and everything it provisioned — the HelmRelease, S3 data and bucket, the managed database, secrets, PVCs and the namespace — and how to verify that nothing is left behind.
Deletion is destructive and irreversible by default
Deleting a Nextcloud removes its data (S3 objects, database, volumes). Verify backups and any
data-retention requirements before you delete. There is no undo unless you configure
reclaimPolicy: Retain on the pool — see reclaimPolicy: Retain below.
What gets deleted¶
Deleting the logical Nextcloud (nc) resource cascades to its NextcloudInstance (nci), whose
deletion handler runs an ordered, blocking cleanup. The nci (and its namespace) only disappear once
every step has succeeded — the finalizer keeps the resource in Terminating and the operator retries
until cleanup completes.
| Step | Resource | Removed when |
|---|---|---|
| 1 | HelmRelease (cascades pods, services, ingress) | always |
| 2 | S3 objects + bucket | bucket was auto-created by the operator (status.appliedS3Config.autoCreated: true) and status.appliedS3Config.bucketDeleted is confirmed true |
| 2.5 | S3Backup (data backup) |
spec.backups.data.deleteOnCleanup: true |
| 3 | Managed PostgreSQL (PerconaPGCluster) |
spec.database.deleteOnCleanup: true or the instance owns its namespace |
| 4 | Secrets (admin, db, redis, s3, mail, …) | after Step 2 is confirmed (bucketDeleted: true) — the S3 secret is preserved until the bucket is gone |
| 4.5 | Orphaned PVCs | only when the namespace is preserved (not owned) |
| 5 | Namespace | the instance owns it (k8s.bnerd.com/instance label matches) — only after Step 2 is confirmed |
Cleanup flags¶
- S3 — only auto-created buckets are emptied and deleted. A bucket you supplied yourself
(
spec.s3.bucket/credentialsSecret) is preserved; the operator never deletes user buckets. All objects, including non-current versions and delete markers, are paginated and removed before the bucket is dropped. Deletion is gated on a durablestatus.appliedS3Config.bucketDeletedflag that is set only once the bucket is provably gone. The S3 credentials Secret and the owned namespace are not removed until that flag is set — this ensures that if the delete fails partway through, the next kopf retry can still authenticate and finish emptying the bucket. See S3 teardown — stuck deletion runbook below. spec.database.deleteOnCleanup(defaultfalse) — delete the managedPerconaPGCluster. See the note below about owned namespaces.spec.backups.data.deleteOnCleanup(defaultfalse) — delete theS3Backupresource (and its repository) instead of preserving it.
Owned namespaces always remove the database
When the instance owns its namespace (the normal case for operator-provisioned instances),
namespace deletion removes the managed database and all PVCs regardless of deleteOnCleanup —
there is nothing left to preserve once the namespace is gone. To keep a database, deploy it outside
the instance's namespace and reference it. The operator deletes the PerconaPGCluster before
tearing down the namespace and waits for its teardown to finish, so the namespace does not stall in
Terminating on Percona finalizers.
PVCs¶
All PVCs the operator provisions (Nextcloud data, Redis, PostgreSQL data and pgBackRest) live inside
the instance namespace, so deleting the namespace (Step 5) cascades and removes them. The explicit
PVC sweep (Step 4.5) only runs in the edge case where the namespace is preserved (e.g. an instance
pointed at a shared namespace via spec.instanceRef) — there it best-effort deletes PVCs labelled for
this instance's HelmRelease (app.kubernetes.io/instance=<name>-nextcloud) and its PostgreSQL cluster
(postgres-operator.crunchydata.com/cluster=<name>-pg).
reclaimPolicy: Retain¶
NextcloudPool.spec.lifecycle.reclaimPolicy (default Delete) controls what happens to live data
when the assigned Nextcloud is deleted. Setting it to Retain preserves the tenant's data for
manual inspection or migration instead of destroying it.
What is preserved vs. removed¶
| Resource | Delete (default) |
Retain |
|---|---|---|
| S3 bucket (auto-created) | Emptied and deleted | Preserved — recorded in audit log only (no K8s label) |
Managed PostgreSQL (PerconaPGCluster) |
Deleted | Preserved — inside retained namespace; audit log line emitted |
| PVCs (Nextcloud data, Redis, pg data) | Deleted via namespace cascade | Preserved — inside retained namespace |
| Owned namespace | Deleted | Preserved — receives label k8s.bnerd.com/reclaim=retained + annotation |
| HelmRelease (cascades pods, services, ingress) | Deleted | Deleted |
| Operator-owned Secrets (admin, db, redis, …) | Deleted | Deleted |
NextcloudInstance (nci) object |
Removed | Removed |
In short: control-plane resources are always cleaned up; live data is preserved. Only the namespace receives the K8s label — the S3 bucket is traceable via the audit log only.
How retained resources are marked¶
Only the owned namespace receives a Kubernetes label and annotation:
- Label:
k8s.bnerd.com/reclaim=retained - Annotation:
k8s.bnerd.com/reclaim-note: released-for-manual-reclamation
The PVCs and managed PostgreSQL cluster live inside the retained namespace, so they are discoverable via the namespace label. The S3 bucket is an external resource — it carries no Kubernetes label. All three are recorded in the operator audit log with an INFO-level line:
S3 reclaim audit: instance=<ns>/<name> bucket=<bucket> reclaim=retained released-for-manual-reclamation (Retain policy — bucket preserved)
PG reclaim audit: instance=<ns>/<name> pgcluster=<name>-pg reclaim=retained released-for-manual-reclamation (Retain policy — managed DB preserved)
Namespace reclaim audit: instance=<ns>/<name> namespace=<ns> reclaim=retained released-for-manual-reclamation (Retain policy — namespace preserved)
The audit log is the authoritative record for the S3 bucket name — since the bucket carries no tag, the log is the only machine-readable trace linking it to the deleted instance.
S3Backup is orthogonal¶
reclaimPolicy does not gate the S3Backup resource. The backup follows its own flag:
spec.backups.data.deleteOnCleanup (default false). Under Retain, the S3 data bucket is
preserved AND the backup is also preserved (unless deleteOnCleanup: true is set independently).
Resolution order and the stamped annotation¶
The policy is resolved at instance-deletion time in this order:
- Live pool CR (authoritative) — the instance's
k8s.bnerd.com/poollabel names theNextcloudPool; itsspec.lifecycle.reclaimPolicyapplies whenever the pool is readable. - Stamped annotation — when the pool cannot be resolved (label missing, pool CR deleted, API
error), the instance's own
k8s.bnerd.com/reclaim-policyannotation applies. Pool deletion stamps this annotation onto every still-assigned instance (see Pool Provisioning → Decommissioning a pool), so deleting a pool does not silently revoke aRetainguarantee. - Fail-safe
Delete— any remaining resolution miss (no stamp, invalid value) causes the operator to fail safe toDelete.Retainis only applied when it is unambiguously and explicitly configured (live pool or stamp). A lookup failure can never silently preserve data or widen deletion.
Configuring reclaimPolicy¶
apiVersion: k8s.bnerd.com/v1alpha1
kind: NextcloudPool
metadata:
name: production
spec:
replicas: 5
lifecycle:
reclaimPolicy: Retain # default: Delete
Runbook: manually reclaiming retained resources¶
After a Retain-policy deletion, the live resources remain in the cluster. The owned namespace
is labeled k8s.bnerd.com/reclaim=retained — use that to locate everything that was preserved.
The S3 bucket name is recorded only in the operator audit log.
1. Find retained namespaces¶
2. Inspect resources inside a retained namespace¶
The PVCs, managed PostgreSQL cluster, and remaining workloads live inside the retained namespace. The namespace label is the entry point — not the resources themselves:
3. Find the S3 bucket name via the audit log¶
The S3 bucket carries no Kubernetes tag. Its name is recorded in the operator log:
kubectl logs -n <operator-ns> deploy/nextcloud-operator \
| grep -E "S3 reclaim audit.*reclaim=retained"
The log line includes the bucket name: bucket=<bucket-name>.
4. Manual cleanup sequence¶
Once you have verified or migrated the data, clean up in this order:
-
Delete S3 bucket — use your S3 client (the operator has released ownership):
-
Delete the managed PostgreSQL cluster:
-
Delete residual PVCs (if any survive the pg cluster teardown):
-
Delete the retained namespace:
Wait for Percona finalizers
Deleting the PerconaPGCluster first and waiting for it to finish prevents the namespace from
stalling in Terminating on Percona's finalizers. Check with
kubectl get perconapgcluster -n <ns> before deleting the namespace.
Recreate safety¶
You can delete a Nextcloud and immediately recreate one with the same name without producing a
duplicate namespace. The operator binds each instance to the Nextcloud's uid; on recreate it detects
the previous instance still terminating and waits for it to finish before provisioning a fresh one.
The old, orphaned instance is never re-blocked by its same-named successor and always completes its own
cleanup.
Audit trail¶
S3 teardown emits one consolidated, INFO-level audit line — one per outcome transition, not
one per retry tick — in the same format on every deletion path — happy path, Retain,
S3-disabled, and unresolvable credentials alike:
S3 cleanup audit: instance=<namespace>/<instance> bucket=<bucket> objects=<N> versions=<M> status=<status>
status |
Meaning |
|---|---|
deleted |
Bucket existed, was emptied (including non-current versions), and deleted. |
absent |
Bucket already didn't exist — idempotent success. |
retained |
reclaimPolicy: Retain — bucket intentionally preserved, not touched. |
skipped |
Two distinct triggers, same status string — tell them apart via bucket=: S3 not enabled on this instance at all (bucket=-), or S3 enabled but the bucket was user-supplied, not auto-created (bucket=<name>) — nothing of the operator's to delete either way. |
no_credentials |
Auto-created bucket, but S3 credentials couldn't be resolved. Not confirmed gone — deletion retries until resolved; see S3 teardown — stuck deletion runbook below. |
failed |
A genuine drain/delete error. Not confirmed gone — blocks secret/namespace deletion; see S3 teardown — stuck deletion runbook below. |
Only deleted, absent, and skipped (either trigger) count as confirmed gone and let
teardown proceed to the secret/namespace steps; no_credentials and failed (surfaced on the
instance as reason: S3CredentialsUnavailable / a generic failure in the Deleting
condition, distinct from the audit line's own status value) trigger the fail-fast retry.
Because the audit line only fires on an outcome transition, repeated retries of the same
still-failing outcome do not re-log on every tick — expect one no_credentials or failed
line, then one more once the underlying issue is fixed and the next attempt changes the
outcome.
The same bucket name, outcome, and object/version counts are also posted as a Kubernetes
event on the NextcloudInstance (reason: S3BucketCleanup) whenever a bucket-delete is
attempted, so kubectl describe nci <name> surfaces them without grepping operator pod logs:
PVC sweeps log PVC cleanup audit: instance=<namespace>/<instance> deleted pvc=<name>. Capture
these and the S3 audit lines from the operator pod for compliance:
kubectl logs -n <operator-ns> deploy/nextcloud-operator | grep -E "S3 cleanup audit|PVC cleanup audit"
Runbook: delete a Nextcloud instance¶
Pre-flight checklist¶
Confirm all of these before deleting:
- Sign-off — deletion is approved by whoever owns the tenant relationship. This is
destructive by default (see the warning above) — there is no "are you sure?" prompt once
kubectl deleteruns. - Retention policy — any contractual/legal data-retention period for this tenant has expired, or you have an explicit exception on record.
- Backups verified — if
spec.backups.data.enabledwas set, confirm a recent, restorable backup exists. Do this regardless of whichreclaimPolicyyou're about to use —Retainpreserves the live resources, it is not itself a backup. - No active users — check activity (e.g.
occ user:lastseenvia theNextcloudCommandrunner) rather than assuming the tenant is idle. - Reclaim policy decided —
Delete(default, destroys live data on completion) orRetain(preserves the S3 bucket, managed DB, PVCs, and namespace for manual reclamation) — see reclaimPolicy: Retain below. If you needRetain, make sure it's set (via the pool'sspec.lifecycle.reclaimPolicy, or the instance's stampedk8s.bnerd.com/reclaim-policyannotation) before you delete. - Bucket name noted, for later verification against your S3 client:
Steps¶
- Delete the resource:
- Monitor progress:
kubectl get nc <name> -n <namespace> # eventually NotFound kubectl get ns | grep <instance-namespace> # shows Terminating, then gone kubectl describe nci <name> -n <namespace> | grep -A2 S3BucketCleanup # bucket outcome + counts kubectl logs -n <operator-ns> deploy/nextcloud-operator | grep -E "CLEANUP|S3 cleanup audit" - Verify cleanup is complete:
Troubleshooting¶
Namespace stuck in Terminating¶
A namespace usually finishes terminating within a few minutes. If it stalls:
kubectl describe namespace <name> | grep -iA3 finalizers
kubectl get nci -n <name> -o jsonpath='{.items[*].metadata.finalizers}'
Most often a NextcloudInstance cleanup step is still failing (e.g. a managed DB whose operator is
unreachable, or an S3 endpoint that is down) and the operator is retrying — check the operator logs for
the blocking step. If the Deleting condition shows reason: S3BucketCleanupPending or
reason: S3CredentialsUnavailable, see the S3 teardown — stuck deletion runbook.
See also Troubleshooting → Finalizer blocks namespace deletion.
Last resort — force-clearing namespace finalizers orphans storage
Force-clearing a namespace's finalizers abandons whatever the finalizer owner was cleaning up.
If the NextcloudInstance inside was mid-S3-teardown, the bucket will be orphaned. Only do
this if the finalizer's owner is permanently gone and you have confirmed the S3 bucket is
already gone (or are prepared to reclaim it manually):
Instance refuses to delete (assigned to an active Nextcloud)¶
A pool-assigned instance is protected from direct deletion. Delete the Nextcloud instead, or use the
force-delete escape hatch — see Operations → Force Delete.
S3 bucket not deleted¶
The operator only deletes auto-created buckets. If you supplied the bucket yourself
(spec.s3.bucket / credentialsSecret), delete it yourself — the operator leaves it
completely untouched. If an auto-created bucket should have been deleted but survives,
check the operator logs for S3 cleanup audit: … status=failed and the preceding error
(credentials, endpoint reachability, or bucket policy). See the
S3 teardown — stuck deletion runbook for the full
diagnostic and recovery procedure.
Runbook: S3 teardown — stuck deletion (#7833)¶
Symptom: A NextcloudInstance is stuck in Terminating and kubectl get nci $NAME -n $NS
shows Deleting=False with reason: S3BucketCleanupPending or reason: S3CredentialsUnavailable.
What is happening:
The operator tears down S3 in a fail-fast, fail-safe sequence:
- It tries to empty and delete the auto-created bucket.
- Only when the bucket is confirmed gone (deleted, already absent, or not the operator's to
delete) does it set the durable
status.appliedS3Config.bucketDeleted: trueflag. - Steps 4 (secret deletion) and 5 (namespace deletion) do not run until that flag is set.
If the bucket is not confirmed gone, the operator raises a TemporaryError and kopf retries —
with both the S3 credentials Secret and the namespace intact so the next retry can still
authenticate and finish the teardown. This prevents the bucket from being orphaned (the pre-0.19.2
bug where the namespace was deleted before the bucket was confirmed gone, taking the credentials
secret with it and deadlocking the retry).
Check the current state:
# Deleting condition (S3BucketCleanupPending or S3CredentialsUnavailable)
kubectl get nci $NAME -n $NS \
-o jsonpath='{.status.conditions[?(@.type=="Deleting")]}' | jq
# Durable bucketDeleted flag
kubectl get nci $NAME -n $NS \
-o jsonpath='{.status.appliedS3Config.bucketDeleted}{"\n"}'
# Operator logs for this instance
kubectl logs -n nextcloud-operator-system deploy/nextcloud-operator \
| grep -E "S3 cleanup audit|S3 bucket|bucketDeleted|S3CredentialsUnavailable|S3BucketCleanupPending"
reason: S3BucketCleanupPending — the bucket delete attempt failed:
The bucket exists and the operator has credentials, but emptying or deleting it raised an error. Common causes:
| Log pattern | Likely cause | Remediation |
|---|---|---|
delete_objects reported N error(s) … AccessDenied |
Bucket policy or IAM role blocks DeleteObject |
Add s3:DeleteObject permission, then let kopf retry |
delete_objects reported N error(s) … InternalError |
Transient S3 endpoint issue | Kopf retries automatically every 30 s |
S3 cleanup audit: … status=failed |
Generic drain/delete failure | Check the preceding error line for the root cause |
Once the underlying issue is fixed the operator retries automatically. You can also force an immediate retry:
# The instance is already in Terminating — kopf retries on its own schedule,
# but a forced reconcile on the logical Nextcloud (if it still exists) accelerates it.
# If the nci is being deleted directly, kopf's TemporaryError retry is the mechanism.
kubectl logs -n nextcloud-operator-system deploy/nextcloud-operator \
| grep -E "Cleanup incomplete|S3 bucket cleanup"
reason: S3CredentialsUnavailable — the credentials Secret is missing:
This means the S3 credentials Secret (<name>-nextcloud-s3) was deleted before the bucket was
confirmed gone — for example, if the secret was deleted manually, or if this instance was
upgrading from a pre-0.19.2 operator that had the deadlock bug and left the secret deleted.
The operator will not proceed and will not orphan the bucket. Your options:
- Restore the credentials secret and let the operator retry. The secret must contain valid credentials for the bucket's endpoint. If you have the credentials at hand:
kubectl create secret generic $NAME-nextcloud-s3 \
-n $NS \
--from-literal=endpoint=<endpoint> \
--from-literal=bucket=<bucket-name> \
--from-literal=accessKey=<access-key> \
--from-literal=secretKey=<secret-key>
Once the secret exists the operator can authenticate, delete the bucket, set
bucketDeleted: true, and complete teardown automatically.
- Delete the bucket manually on the object store, then patch the durable flag so the operator treats the bucket as gone and proceeds:
# Delete the bucket via your S3 client first, then:
kubectl patch nci $NAME -n $NS \
--type=merge \
--subresource=status \
-p '{"status":{"appliedS3Config":{"bucketDeleted":true}}}'
The operator will detect bucketDeleted: true on the next retry and skip the bucket step,
proceeding to secret and namespace teardown.
Last resort — force-clearing the finalizer:
Orphans the bucket — only use when the bucket is already confirmed gone
Force-clearing the kopf finalizer drops the TemporaryError retry loop and immediately
removes the NextcloudInstance object. If the bucket still exists, it will be orphaned.
Only do this after confirming (via your S3 client) that the bucket is gone or was never
created:
Audit log confirmation:
Once the operator successfully deletes the bucket you will see:
S3 cleanup audit: instance=<namespace>/<instance> bucket=<bucket> objects=<N> versions=<M> status=deleted
And status.appliedS3Config.bucketDeleted will be true.
Orphaned PVCs¶
If an instance did not own its namespace and PVCs remain, delete them by instance label:
kubectl get pvc -n <namespace> -l app.kubernetes.io/instance=<name>-nextcloud
kubectl get pvc -n <namespace> -l postgres-operator.crunchydata.com/cluster=<name>-pg
0.20.0 upgrade note: previously-stuck deletions may now complete (destructive)¶
Before 0.20.0, a profile that centralized S3 credentials (e.g. defaults.s3.{accessKey,
secretKey,endpoint,region} shared across every profile-backed instance, with each instance
only setting its own s3.{enabled,bucket}) had those credentials silently dropped during
deletion — a bug in on_delete's own profile-merge step, the same class of defect described
in Profiles → How Defaults Merge but in a fourth, separate
copy of the merge logic. The instance's effective s3 config resolved to empty credentials,
delete_auto_created_bucket returned no_credentials, and the fail-fast gate above raised
TemporaryError forever — every retry re-ran the identical broken merge and hit the
identical empty credentials, with no self-correction possible. This is structurally the same
shape as a stuck-Terminating namespace with no working retry path.
0.20.0 fixes the merge. Consequence at upgrade: any instance that was stuck Terminating specifically because of this bug resumes and completes its deletion on the operator's next reconcile attempt for that instance — including the S3 bucket delete, which is destructive and irreversible.
Before upgrading a fleet with any stuck-Terminating instances:
- List instances currently stuck deleting:
- For every row with a bucket not yet deleted (
bucketDeleted=false), decide before upgrading whether that bucket's data may still be needed. If it might be: complete a manual data audit first, or set the instance's reclaim policy toRetain(see reclaimPolicy: Retain) so the fixed path retains the bucket instead of deleting it. - Don't assume a stuck instance will stay stuck after the upgrade — this fix is exactly what un-sticks it, and it un-sticks it by finishing the deletion, not by pausing it.