Troubleshooting¶
Day-2 reference for diagnosing a NextcloudInstance, Nextcloud, or NextcloudPool that isn't behaving as expected. Start with State inspection, then Logs, then match your symptom in the Common errors table.
State inspection — always start here¶
NS=nextcloud-demo
NAME=demo
# 1. CRD status + events (events often contain the real reason)
kubectl describe nci $NAME -n $NS
# 2. Conditions (ready / failed / why)
kubectl get nci $NAME -n $NS -o jsonpath='{.status.conditions}' | jq
# 3. Resolved Helm chart version
kubectl get nci $NAME -n $NS -o jsonpath='{.status.versionResolution}' | jq
# 4. Downstream HelmRelease
kubectl describe helmrelease -n $NS
# 5. Pods (anything CrashLoopBackOff / ImagePullBackOff / Pending?)
kubectl get pods -n $NS
kubectl describe pod -n $NS <pod-name>
# 6. Managed database (only if spec.database.managed=true)
kubectl get perconapgcluster -n $NS
kubectl describe perconapgcluster -n $NS
# 7. Secrets created by the operator
kubectl get secret -n $NS -l app.kubernetes.io/managed-by=nextcloud-operator
For the logical Nextcloud tenant resource:
kubectl describe nc $NAME -n $NS
kubectl get nc $NAME -n $NS -o jsonpath='{.status.instanceRef}' | jq
Where are the logs?¶
| Component | Command |
|---|---|
| Operator (kopf handlers) | kubectl logs -n nextcloud-operator-system -l app.kubernetes.io/name=nextcloud-operator --tail=500 |
| Operator — specific handler errors | Same, then grep -i "error\|PermanentError\|TemporaryError" |
| Flux HelmController (who actually installs the chart) | kubectl logs -n flux-system -l app=helm-controller --tail=200 |
Flux SourceController (resolves HelmRepository) |
kubectl logs -n flux-system -l app=source-controller --tail=200 |
| Nextcloud PHP | kubectl logs -n $NS deploy/<release>-nextcloud -c nextcloud --tail=200 |
Nextcloud occ output (on-demand) |
NextcloudCommand CRD; result in .status.results |
| Percona PG operator | kubectl logs -n pgo -l app.kubernetes.io/name=percona-postgresql-operator --tail=200 |
| PostgreSQL itself | kubectl logs -n $NS <pgcluster>-instance1-xxxx-0 -c database --tail=200 |
Make the operator verbose
The operator is started with kopf run --verbose by default in our chart. If you run it differently, increase verbosity with --verbose or --debug to see handler-level reconciliation traces.
Force reconciliation¶
When state looks wrong but no error is visible, nudge the operator to re-run its handlers:
# Force reconcile (operator re-runs all field handlers)
kubectl annotate nci $NAME -n $NS k8s.bnerd.com/reconcile=$(date +%s) --overwrite
# Force maintenance tasks on demand
kubectl annotate nci $NAME -n $NS k8s.bnerd.com/run-maintenance=$(date +%s) --overwrite
See Operations & Annotations for all supported annotations, including k8s.bnerd.com/force-delete for bypassing deletion protection.
Common errors¶
Instance stuck in Pending or Creating¶
Symptom: kubectl get nci shows Pending or Creating for more than 5 minutes.
Check in order:
- Operator logs for validation errors — a
PermanentErrormeans the spec is wrong and no amount of retrying will help. Example: missingspec.admin, invalidspec.database.type, unknownspec.version. - HelmRepository + HelmRelease exist —
kubectl get helmrelease,helmrepository -n $NS. If missing, Flux isn't installed or the operator couldn't reach the K8s API. - Pods
Pending— usually means no defaultStorageClassor the cluster is out of resources.kubectl describe podsurfaces the scheduler's reason. - Pods
ImagePullBackOff— the registry isn't reachable or the image/tag doesn't exist. Checkspec.image/ resolved chart version.
Database not ready yet — managed PG never becomes ready¶
Symptom: Operator logs repeat TemporaryError: Database not ready yet until the 20-minute timeout, then the instance goes Failed.
Causes:
- Percona PG Operator not installed —
kubectl get crd perconapgclusters.pgv2.percona.com. Install viahelm install pgo percona/pg-operator -n pgo --create-namespace. - Percona operator is running but has no RBAC in the target namespace — see its logs:
kubectl logs -n pgo -l app.kubernetes.io/name=percona-postgresql-operator. - No
StorageClass— the PG cluster can't provision PVCs. - Resources too tight — PG instance pods pending because no node has enough CPU/memory.
To recover after fixing the underlying issue, annotate the instance to force reconcile:
NetworkPolicy blocks Patroni's API-server access (managed PostgreSQL)¶
Symptom: the PostgreSQL pod(s) for a managed database report a 3/4 ready
container count and never converge; Patroni's own logs show connection timeouts or
K8sConnectionFailed while talking to its DCS (Distributed Configuration Store) —
the Kubernetes API server itself, Percona PG's default DCS backend on Kubernetes.
Left long enough, this looks identical to the generic Database not ready yet
timeout above, but no amount of waiting resolves it.
Cause: an egress-restricting NetworkPolicy in the instance's namespace — often
delivered via a profile's or instance's helm.values.extraManifests block — blocks
the pod's egress to the API server. Patroni requires direct pod→API-server access for
leader election; without it, the cluster never reports healthy. The operator does not
detect or warn about this today (extraManifests is an opaque Helm-values
passthrough, and the operator has no NetworkPolicy-aware logic anywhere) — this is a
config-side diagnosis, not an operator condition.
Fix: ensure any egress NetworkPolicy applied to the managed-database namespace
allows:
- TCP 6443 to the control-plane node(s) / API-server endpoints (the real endpoint
IPs, not just the in-cluster
kubernetesService) - TCP 443 to the API server's ClusterIP (
kubernetes.default.svc)
Diagnose:
# Confirm the ready count and look for the stuck container
kubectl get pods -n $NS -l postgres-operator.crunchydata.com/cluster=<name>-pg
# READY 3/4 is the signature
# Patroni's own logs, in the `database` container
kubectl logs -n $NS <pg-pod> -c database | grep -i "k8sconnectionfailed\|dcs\|connection.*timed out"
# What's actually restricting egress
kubectl get networkpolicy -n $NS -o yaml
See Managed PostgreSQL → Networking Requirements for the full requirement.
ConflictingDatabaseConfig warning event (managed DB with a credentialsSecret)¶
Symptom: kubectl get events on the instance shows a Warning event with
reason: ConflictingDatabaseConfig, e.g.:
spec.database.managed=true and spec.database.credentialsSecret='my-db-credentials'
are both set. The operator's own managed-database connection info is used;
credentialsSecret is ignored for connection info.
Cause: spec.database.managed: true and spec.database.credentialsSecret are
both set. This is not a failure — the instance is not blocked and does not go
Failed — but it is a spec smell: the two fields name different databases in
intent (a Percona cluster the operator provisions vs. an external one), and only
one of them is actually live. The operator always prefers its own managed-database
connection info; credentialsSecret is ignored for connection purposes. The event
fires once per instance (deduped via status.databaseConfigConflictWarned), not on
every reconcile.
Fix: decide which database the instance should actually use, then clean up the spec accordingly:
- Keep managed — remove
spec.database.credentialsSecret. No behavior change (it was already being ignored), just removes the stale field and the warning. - Actually want external — set
spec.database.managed: false(or omit it) socredentialsSecrettakes effect for real. This does not migrate data between databases; only do this if the instance was never actually provisioned against the managed cluster, or you have already migrated the data yourself.
See Managed PostgreSQL → managed: true always wins over a
credentialsSecret
for the full precedence rule.
HelmRelease stuck or in Failed state¶
Symptom: kubectl get helmrelease -n $NS shows Ready=False.
Diagnose:
kubectl describe helmrelease -n $NS
kubectl logs -n flux-system -l app=helm-controller --tail=200 | grep $NAME
Common HelmRelease failures:
chart "nextcloud" version "x.y.z" not found— the resolved chart version doesn't exist in the repository. Checkstatus.versionResolutionon the instance; if you pinnedspec.helm.version, verify that tag exists in the upstream Helm repo.values don't validate against schema— usually from customspec.helm.values. Test your values locally withhelm template.timed out waiting for the condition— chart installed but pods never went ready. Inspect the pods directly.
To force Flux to retry: flux reconcile helmrelease <release-name> -n $NS --with-source.
Pool instance never gets assigned¶
Symptom: Nextcloud (logical) stays in Assigning phase; NextcloudPool.status.unassigned is 0.
Causes:
- Pool is drained — all instances already assigned. Increase
spec.replicason the pool or wait for the pool reconciler to replenish. - Labels don't match —
spec.poolSelector.matchLabelson theNextclouddoesn't matchtemplate.metadata.labelson the pool. Compare withkubectl get nci -A --show-labels | grep pool. - Pool instances stuck
Pending— the pool is creating replacements but they can't become ready (see "Instance stuck in Pending" above). - Instances exist but aren't
Readyyet — only a fully installed instance (phase: Ready,status.installed: true) is assignable. The operator will not hand a partially-provisioned instance to a tenant. Checkkubectl get nci -n $NSfor the pool's instances; if they sit inDeployingwith reasonWaitingForInstall, see the next section. Once they reachReady, the pendingNextcloudis assigned automatically on the next reconcile.
Instance stuck in Deploying with the Installed condition False¶
Symptom: Pods are up and HelmRelease/workload are ready, but the instance never reaches Ready. The Installed condition is False and status.installed is false.
The operator treats "Ready" as installed, not just "pods are running". After the HelmRelease and workload come up it runs occ status once and only advances to Ready when Nextcloud reports installed: true. While it waits, the Installed condition's reason tells you what is happening:
Already-Ready instances are also checked once (backfill)
Instances that reached phase: Ready under an older operator version — before the install check existed — were never verified (status.installCheckedAt is unset, status.installed is empty). The operator now runs the install check once for these too. This backfill is read-only and flag-only: if such an instance turns out to be not installed it is flagged with the WedgedNeedsManualInstall reason (below), never auto-installed. The check fires only while installCheckedAt is unset and stops once the result is cached.
kubectl get nci $NAME -n $NS -o jsonpath='{.status.conditions[?(@.type=="Installed")]}' | jq
kubectl get nci $NAME -n $NS -o jsonpath='{.status.installed}{"\n"}'
Installed reason |
Meaning | What to do |
|---|---|---|
WaitingForInstall |
Pods are up; Nextcloud is finishing its first install. | Normal for the first 1–3 min after pods become ready. No action. |
HealInstallTriggered |
The first install had wedged (started but never finished) with an empty database, so the operator re-ran occ maintenance:install in place to un-wedge it. |
None — the next reconcile re-checks and should flip to Ready. The operator did this only after proving the database was empty. |
DbProbePending |
The DB-empty safety probe could not complete on the last few reconciles (up to 10 consecutive). No install was attempted. | Usually transient — the Postgres pod may still be starting. Check kubectl get pods -n $NS. Clears automatically once the probe succeeds. |
DbProbeUnreachable |
The DB-empty safety probe has been UNDETERMINED for 10 or more consecutive reconciles (~5 min at the default 30 s timer). The operator cannot confirm the database is empty and will not auto-install (fail-closed by design). | Alertable — see the DbProbeUnreachable runbook below. |
WedgedWithData |
Nextcloud reports not installed, but its database already contains tables. The operator will never auto-install over existing data (it would destroy it). | Manual investigation required — see below. |
RegressedAfterInstall |
The instance previously reported installed and now reports not-installed (a regression, not a first-install failure). The operator will not auto-install. | Manual investigation required — see below. |
HealAttemptsExhausted |
The operator tried the safe self-heal install the bounded number of times and it still isn't installed. | Manual investigation required — see below. |
WedgedNeedsManualInstall |
The install-check backfill found a pre-existing instance (already phase: Ready under an older operator version) that reports not installed. Because the instance was long-Ready its database state is unknown, so the operator flags it rather than auto-installing. |
Confirm the database state, then run occ maintenance:install manually — see below. |
The operator never auto-installs over a non-empty database
The automatic self-heal runs occ maintenance:install only when it can prove there is nothing to lose: the instance is not installed, has never been installed before, and the database is verifiably empty of tables. If any of those is not true — or the operator cannot verify the database is empty — it stops and surfaces a Warning condition instead of touching the instance. This is deliberate: re-installing over a data-bearing instance would wipe it.
Manual path for WedgedWithData / RegressedAfterInstall / HealAttemptsExhausted / WedgedNeedsManualInstall:
- Look at what the instance's database actually contains:
- If the database is genuinely empty and you want the operator to (re)install, it is safe to clear the wedged state — recreate the instance (delete the
NextcloudInstance; the pool reconciler creates a fresh one) so it provisions cleanly. ForWedgedNeedsManualInstallspecifically (an already-Readyinstance flagged by the backfill), prefer installing in place once you have confirmed the database is empty: - If the database has real tenant data, do not re-install. This is a recovery case: restore from backup or repair the Nextcloud install in place with the appropriate
occcommands. Escalate to b'nerd if you are unsure — re-installing here loses data.
Runbook: DbProbeUnreachable — instance stuck not-installed (#7832)¶
Symptom: kubectl get nci $NAME -n $NS -o jsonpath='{.status.conditions[?(@.type=="Installed")]}'
returns reason: DbProbeUnreachable. A Warning event is also recorded on the NextcloudInstance
object.
What this means:
The operator runs a DB-empty safety probe before it will auto-install Nextcloud: it connects to the
Postgres primary pod and counts tables via psql. This probe returning UNDETERMINED (non-zero
exit, no pod, network policy blocked, wrong credentials, …) for 10 or more consecutive reconciles
(roughly 5 minutes at the default 30 s timer) escalates from the benign DbProbePending condition
to an alertable DbProbeUnreachable condition plus a Warning event.
This is fail-closed by design. The operator will not auto-install while the probe result is unknown — even if it has been unknown for a long time. This is not a data-loss risk: no install command has been run, nothing has been mutated. The instance is simply held in a pre-install holding pattern until the connectivity problem is fixed.
Diagnose the probe failure:
# See the full condition message (includes failure count)
kubectl get nci $NAME -n $NS \
-o jsonpath='{.status.conditions[?(@.type=="Installed")]}' | jq
# See the Warning event
kubectl get events -n $NS --field-selector reason=DbProbeUnreachable
# Check probe failure counters in status
kubectl get nci $NAME -n $NS \
-o jsonpath='{.status.dbProbeFailures}{"\n"}{.status.dbProbeFirstFailedAt}{"\n"}'
# Is the Postgres primary pod running?
kubectl get pods -n $NS -l postgres-operator.crunchydata.com/role=master
# or, for Percona-managed clusters:
kubectl get pods -n $NS | grep -i pg
# Has the PG cluster actually finished bootstrapping? The #7832 incident's exact
# signature was a Postgres primary pod that looked Running while the cluster's
# pgBackRest replica-create bootstrap job was still stuck (or had failed) — the
# instance/data directory never fully initialized, so the primary answered exec
# but psql connections behaved as if the DB wasn't there yet.
kubectl get jobs -n $NS -l postgres-operator.crunchydata.com/pgbackrest-backup=replica-create
kubectl logs -n $NS job/<name>-repl-create-xxxxx # from the job listed above
# pg_hba: confirm the `nextcloud` role is actually allowed to connect locally
# (the probe execs into the primary pod and connects to 127.0.0.1, so this is a
# local, not remote, pg_hba entry)
kubectl exec -n $NS <pgcluster>-instance1-xxxx-0 -c database -- \
psql -U postgres -c "SELECT rolname FROM pg_roles WHERE rolname = 'nextcloud';"
kubectl exec -n $NS <pgcluster>-instance1-xxxx-0 -c database -- \
cat /pgdata/pg*/pg_hba.conf | grep -i nextcloud
# Operator logs — look for "DB-empty probe" lines
kubectl logs -n nextcloud-operator-system deploy/nextcloud-operator \
| grep -i "DB-empty probe"
Common causes and remediation:
| Cause | Remediation |
|---|---|
| Postgres primary pod not running / still starting | Wait for the cluster to become healthy (kubectl get perconapgcluster -n $NS), or fix the underlying infrastructure issue. The operator retries automatically. |
PG cluster bootstrap stuck — pgBackRest replica-create job never completed (the #7832 incident's actual root cause) |
The primary pod can be Running and still exec-reachable while the cluster's storage/stanza initialization hasn't finished. Check the bootstrap job's status and logs (commands above); if it's stuck or Failed, investigate the underlying storage/S3-repo issue (or delete the failed job to let the Percona operator retry it — the PG cluster itself, not the probe, needs fixing here). |
Wrong DB credentials in the <name>-nextcloud-db Secret |
Correct the credentials. If using spec.database.credentialsSecret, verify the referenced secret contains the correct password. |
pg_hba.conf does not allow the nextcloud role to log in locally |
The DB-empty probe connects to 127.0.0.1 inside the Postgres pod. Confirm the role exists and pg_hba.conf has a local entry for it (commands above). |
| Role / database not yet created | On an external (non-managed) database, ensure the role and database exist before creating the instance. |
| Network policy blocks pod exec from operator | The probe uses kubectl exec (WebSocket) into the Postgres pod. Ensure the Kubernetes API server can exec into pods in the instance namespace. |
(historical, fixed in 0.21.0) DB password contained a shell metacharacter (e.g. a ') |
Before 0.21.0 the probe interpolated the password directly into the exec script; a quote in the password broke the script's quoting, the probe exited non-zero, and a perfectly healthy DB wedged as permanently UNDETERMINED. As of 0.21.0 the password travels via exec stdin, never the script text — this class of failure can no longer occur. If you're still on an older operator version and suspect this, the workaround is to avoid ' in generated DB passwords, or upgrade. |
Recovery:
Once DB connectivity is restored and the probe can run, the operator auto-recovers on the next
reconcile — no manual intervention is needed. If the probe returns db_empty=True and all other
preconditions hold (never previously installed, attempts remaining), occ maintenance:install
runs automatically and the instance proceeds to Ready. This holds regardless of how many
consecutive failures were counted before recovery — the escalation to DbProbeUnreachable is
purely an observability signal; it does not change the gate's behavior once the probe starts
succeeding again, and the failure counter resets to zero on that first success.
If you want to accelerate recovery rather than waiting for the next 30 s timer tick:
kubectl annotate nci $NAME -n $NS \
k8s.bnerd.com/reconcile=$(date -u +"%Y-%m-%dT%H:%M:%SZ") \
--overwrite
Last resort: recycle the instance. If the underlying PG cluster is unrecoverable (a wedged
bootstrap that won't complete, storage corruption, etc.), the safest path is to delete the
NextcloudInstance and let the pool reconciler or a fresh apply provision a clean one — never
attempt to force an install against a cluster whose state you can't verify. As of 0.19.2/0.20.0's
teardown fixes (#7833), instance deletion — including the managed-PG and S3 cleanup paths — no
longer deadlocks, so this is now a safe, ordinary operation rather than something that risks
leaving orphaned resources. Reach for it only after the diagnosis steps above have ruled out a
fixable connectivity/credentials issue — this condition never mutates or destroys data on its
own, so there's no urgency to recycle before you've actually looked.
PromQL alert hint:
To alert on this condition before it sits unnoticed, query the Kubernetes events API or the
operator's structured log. If you use kube-state-metrics with a kube_event_count rule:
# Alert when a DbProbeUnreachable Warning event is newer than 10 minutes
count by (namespace, name) (
kube_event_created_at{reason="DbProbeUnreachable", type="Warning"} >
(time() - 600)
) > 0
Alternatively, watch the Installed condition reason via the operator's Prometheus metrics (if
you have CRD-status scraping) or your log-based alerting on the operator pod:
# Log line emitted each time the alertable condition fires
kubectl logs -n nextcloud-operator-system deploy/nextcloud-operator \
| grep "DbProbeUnreachable"
Cannot load API key for SignalingServer / RecordingServer¶
Symptom: The operator retries every 60s with this TemporaryError.
The referenced credentialsSecret is missing or lacks the expected key. Check:
The secret must contain the API key under the key specified in the SignalingServer/RecordingServer spec.
Authentication failed for <api-endpoint>: 401¶
Symptom: PermanentError in operator logs when registering a backend.
The API key in the credentialsSecret is wrong or the backend API is rejecting it. Verify the key on the backend (signaling or recording server), update the secret, then delete the SignalingServer / RecordingServer CR to re-register cleanly.
NextcloudCommand times out¶
Symptom: A NextcloudCommand finishes with phase: Failed and status.results[].stderr shows a timeout.
- A single command exceeded
spec.perCommandTimeoutSeconds(default 300s). For expensive migrations likeocc db:convert-filecache-bigint, raise the per-command timeout to e.g.3600. - The overall job exceeded
spec.timeoutSeconds. Raise it or split the commands across multipleNextcloudCommandresources. - No running Nextcloud pod was available — the instance is not
Ready. Checkspec.targetRefpoints to a healthyNextcloudInstance/Nextcloud.
S3 data backup enabled but no repository configured¶
Symptom: PermanentError at instance creation time.
spec.backups.data.enabled: true requires either an S3Backup CRD (from the bnerd backup operator) to be installed, or the backup repository to be configured. Install the backup operator, or disable the feature.
CrashLoopBackOff on Nextcloud pod after an upgrade¶
Symptom: After bumping spec.version, pods crash-loop with migration errors.
Since 0.23.0 the operator runs db:add-missing-indices and a fast maintenance:repair inside the upgrade window itself, in addition to the pre-existing expensive set (which the maintenance timer now runs for every instance once it notices the version changed, not only those with a windowStart) — check status.upgrade.steps[] first to see whether the in-window pass already failed, and why. You can still trigger the maintenance task set manually:
Watch the operator logs to confirm the maintenance tasks ran. If migrations like add-missing-indices or convert-filecache-bigint fail, inspect the output and run them as a dedicated NextcloudCommand with a longer timeout.
Instance stays in maintenance mode after an upgrade¶
Symptom: Users see Nextcloud's maintenance page long after the version bump rolled out. status.phase is Ready.
This is deliberate, not a wedge. Since 0.23.0 the operator holds the maintenance window across the whole upgrade — including app reconvergence — and keeps it closed when a critical app is still disabled. The critical set is {user_oidc} while spec.oidc.enabled is set: an instance whose only login path is dead is presented as "in maintenance" rather than as "up" and unusable.
kubectl get nci $NAME -n $NS -o jsonpath='{.status.upgrade}' | jq '{phase, pendingApps, maintenanceHeld, attempts}'
kubectl get nci $NAME -n $NS -o jsonpath='{.status.upgrade.steps}' | jq '.[] | select(.success == false)'
phase: AppsPending+maintenanceHeld: true→ the operator is retrying and will lift the window by itself as soon as the app can be enabled (typically when the app store publishes a compatible release).- To bring the instance back up now, with the affected app still broken:
phase: Convergingwith a failedmaintenanceOffstep → the operator converged the apps but could not reach the pod to lift the window. It retries; check pod health.UpgradeStuckcondition present → no progress for two hours. Start from the failed step'smessage, which carries occ's own stderr.
Full mechanism and runbook: Upgrades → Post-upgrade app convergence.
Apps disabled after an upgrade (PostUpgradeAppsDisabled)¶
Symptom: AppsHealthy is False with reason PostUpgradeAppsDisabled, and nextcloud_operator_upgrade_apps_pending is 1 for the instance.
The upgrade disabled an app that has no compatible release available, and the operator's convergence loop is retrying (no attempt cap, ~15 min apart). Waiting is usually the correct action — it heals itself. Do not hand-run occ app:enable via kubectl exec: that bypasses the per-instance occ lock and races the loop.
If the app is gone for good, stop declaring it (spec.apps.<app>.enabled: false) — the loop converges as soon as the declared set matches reality.
Finalizer blocks namespace deletion¶
Symptom: kubectl delete namespace hangs; the namespace stays in Terminating.
A NextcloudInstance still has a finalizer because cleanup is incomplete (typically a managed DB that won't delete). To unblock:
# Check what's blocking
kubectl get nci -n $NS -o jsonpath='{.items[*].metadata.finalizers}'
# Force-delete (skips operator cleanup — only use when you accept losing state)
kubectl annotate nci $NAME -n $NS k8s.bnerd.com/force-delete=true --overwrite
kubectl delete nci $NAME -n $NS
For the full teardown order, audit log, and recreate-safety behaviour, see Deletion & Cleanup.
When to file a bug vs. keep debugging¶
File a bug if:
- The operator panics or the pod crash-loops (
kubectl logsshows a Python traceback without a clearPermanentError). - A
TemporaryErrorrepeats indefinitely even after the underlying cause is fixed. -
Status fields contradict reality (e.g.
phase: Readybut pods areCrashLoopBackOff).Readyonly gets set when the HelmRelease, Deployment, Endpoints, and Ingress all report ready and Nextcloud reportsinstalled: true(status.installed) — so this combination should no longer be reachable. If you see it, file a bug withkubectl get nci $NAME -n $NS -o yamlattached. -
Instance is stuck in
Deploying: inspectstatus.workloadand theReadycondition'sreason.WaitingForHelmRelease→ check the Flux HelmRelease (kubectl describe helmrelease ...).WaitingForPods→ describe the Nextcloud Deployment; usually a values misconfiguration or PVC problem.WaitingForEndpoints→ the Service has no ready backends (pod ready probe failing).WaitingForIngress→ the cluster's ingress controller hasn't assigned a load-balancer address yet. -
Instance is stuck in
Creatingwithdatabase.managed: true: inspectstatus.databaseand theDatabaseReadycondition.reason=Initializingis normal — Percona PG cluster startup typically takes 1–5 min.reason=ProvisioningFailed→ checkkubectl describe perconapgcluster $NAME-pg -n $NSand look at events.reason=Timeout→ 20 min have passed and the cluster still isn't ready; phase will transition toFailedbut the operator keeps retrying. Common causes: missing StorageClass, pg-operator pod not running, image pull failure. Fix the underlying infra issue and the next 60 s retry should self-heal back toCreating → Deploying → Readywithout operator intervention. -
Instance is at
phase: Failedwithstatus.database.reason=PgOperatorNotFound: the Percona PG operator CRD is not installed in the cluster. This is aPermanentError— install the pg-operator (make install-pg-operatoror your usual Flux/Helm flow) and then re-trigger viakubectl annotate nci $NAME -n $NS k8s.bnerd.com/reconcile=$(date +%s) --overwrite.
Keep debugging yourself if:
- The error message explicitly says what's wrong (missing secret, invalid field, wrong version). The operator is telling you — believe it.
- The HelmRelease is failing — that's a chart/values issue, not an operator bug.
- Kubernetes primitives are broken (no
StorageClass, no ingress, no DNS). Those aren't the operator's job to fix.
See also:
- Operations & Annotations — reconcile, run-maintenance, force-delete
- Monitoring — metrics and dashboards for proactive detection
- Managed PostgreSQL — deeper DB-specific troubleshooting