Monitoring¶
The operator exports Prometheus metrics about its own reconciliation loops and the state of every managed Nextcloud, NextcloudInstance, NextcloudPool and HelmRelease. A ServiceMonitor and two Grafana dashboards (Overview + Detail) ship with the chart and as optional flat-YAML manifests.
Metrics are opt-in on both install paths.
Quick start¶
Helm install¶
helm upgrade --install nextcloud-operator ./chart \
--set metrics.enabled=true \
--set metrics.serviceMonitor.enabled=true \
--set metrics.grafana.enabled=true
This creates:
- Container port
metrics/9090and environment variables on the Deployment - A
Serviceexposing port9090 - A
ServiceMonitor(requires the Prometheus Operator CRDs) - Two
ConfigMaps, one per dashboard, labelledgrafana_dashboard: "1"so the Grafana sidecar picks them up automatically
Flat YAML install¶
# 1. Enable metrics in the operator deployment (edit METRICS_ENABLED to "true")
kubectl -n nextcloud-operator-system set env deploy/nextcloud-operator METRICS_ENABLED=true
# 2. Apply the monitoring stack
kubectl apply -f deploy/monitoring.yaml \
-f deploy/dashboard-overview.yaml \
-f deploy/dashboard-detail.yaml
deploy/monitoring.yaml contains the Service and ServiceMonitor. The two
dashboard ConfigMaps are regenerated from the chart via make sync-dashboards.
Prerequisites¶
| Component | Required for |
|---|---|
| Prometheus + prometheus-operator CRDs | ServiceMonitor scraping |
| Grafana with the sidecar dashboard loader | Auto-loading the shipped dashboards |
If you run a stand-alone Grafana, import the JSON from chart/dashboards/
directly.
Exposed metrics¶
All metrics are prefixed with nextcloud_operator_, with one deliberate exception: pgbackrest_last_backup_completion_timestamp_seconds (see Backup freshness (pgBackRest) below) is unprefixed on purpose, so it matches the client's existing *pgback* alert discovery.
State gauges (updated by the background collector every 60s)¶
| Metric | Type | Labels | Meaning |
|---|---|---|---|
nextclouds_total |
Gauge | phase |
Nextcloud CR count by phase |
nextcloudinstances_total |
Gauge | phase |
NextcloudInstance CR count by phase |
nextcloudpools_total |
Gauge | phase |
NextcloudPool count by phase |
nextcloudprofiles_total |
Gauge | — | NextcloudProfile count |
pool_replicas |
Gauge | pool, type (desired/ready/unassigned/assigned) |
Per-pool replica breakdown |
pool_instance_info |
Gauge | pool, instance, namespace, phase, assigned (true/false) |
One series per pool replica (status.instances[]) |
nextcloud_info |
Gauge | name, namespace, phase, instance_name, instance_namespace, profile, url |
Per-CR identity (always 1) |
nextcloudinstance_info |
Gauge | name, namespace, phase, assigned_to (owning Nextcloud name, empty when spare), pool, profile, url |
Per-NCI identity (always 1) |
nextcloudpool_info |
Gauge | name, phase, desired, ready, unassigned, assigned |
Per-pool identity (always 1) |
assignment_info |
Gauge | nextcloud_namespace, nextcloud_name, nextcloud_url, instance_namespace, instance_name, profile, pool, state (assigned/spare) |
One series per NextcloudInstance. Lets dashboards JOIN tenant identifiers (URL, name) onto K8s workload metrics by instance_namespace. Spare pool instances appear with empty nextcloud_* labels. |
nextcloud_condition |
Gauge | name, namespace, type |
Condition state (1/0/-1 = True/False/Unknown) |
nextcloudinstance_condition |
Gauge | name, namespace, type |
Condition state (1/0/-1 = True/False/Unknown) |
helmrelease_ready |
Gauge | namespace, name |
HelmRelease Ready condition (1/0/-1) |
pool_replicas and nextcloudpool_info — fixed in 0.21.0 (#7865). Both read
status.ready/status.unassigned/status.assigned (the actual NextcloudPool CRD
status field names) from the collector's 60s sweep. Before 0.21.0 the collector read
status.readyReplicas/unassignedReplicas/assignedReplicas, which don't exist on the
CRD — so the ready/unassigned/assigned label values and pool_replicas series were
always 0, regardless of the pool's actual state. desired (from spec.replicas) was
unaffected. If a CapacityPoolStarved-style alert on these gauges never fired even
during a real pool exhaustion, this is why; it's now sourced correctly.
pool_instance_info — new in 0.21.0. Emits one series per entry in
status.instances[], labelled with that replica's phase and whether it's currently
assigned to a Nextcloud. Unlike pool_replicas' aggregate counts, this lets you
distinguish an idle-Ready spare (healthy, waiting to be matched) from an idle
replica stuck in a non-Ready phase (wedged) — the failure mode pool_replicas alone
can't surface, since it only counts, it doesn't name. Cleared and repopulated wholesale
every collector sweep, same lifecycle as pool_replicas and nextcloudpool_info — a
replica removed from the pool (scaled down, deleted) drops out of the metric on the
next sweep, no manual cleanup needed.
Example query — pool replicas that are unassigned and not Ready (idle-wedged,
consuming pool capacity without being usable):
Event counters¶
| Metric | Type | Labels | Meaning |
|---|---|---|---|
reconcile_total |
Counter | resource, result (success/error/temporary_error/permanent_error) |
One increment per kopf handler invocation |
errors_total |
Counter | resource, stage (validation/db_provision/helmrelease/occ/maintenance/…) |
Categorised error accounting |
pool_scale_total |
Counter | pool, direction (up/down) |
Pool instance create/delete events |
pool_assignment_total |
Counter | pool, result (success/conflict/no_match) |
Outcome of pool match attempts |
maintenance_task_total |
Counter | task, result (success/error) |
Periodic and post-upgrade OCC task runs |
Latency histograms¶
| Metric | Labels | Covers |
|---|---|---|
operation_duration_seconds |
operation, result |
Generic operation timer; used for db_provision, helmrelease_create_or_update, occ_command, maintenance_task |
instance_ready_duration_seconds |
profile |
Seconds from NCI creation to phase=Ready |
nextcloud_assignment_duration_seconds |
pool |
Seconds from Nextcloud creation to first pool assignment |
Backup freshness (pgBackRest)¶
Unix timestamp of the most recent Succeeded pgBackRest backup completion, one series per (repo, type) per instance. type is full/incremental/differential. This is the metric OpenProject #7835 exists to expose — the Percona operator runs pgBackRest itself, but nothing previously read its results back into Prometheus, so a PgBackRestBackupStale alert had no data to fire on.
Source and update cadence — different from the other gauges on this page, worth calling out explicitly:
- Set on the per-instance maintenance timer (
TIMER_MAINTENANCE_INTERVAL, default 15 min), not the 60s fleet-wide background collector. - Runs before the
spec.maintenance.windowStartgate — a windowless instance still gets this metric updated, same "visibility, not application" rationale asUpdateAvailableand the apps-health check (this feature is pure observation of Percona's ownPerconaPGBackupCRs; it never mutates anything). - Managed-database instances only (
spec.database.managed: true) — unmanaged/external databases have noPerconaPGBackupCRs to read, so the check is skipped entirely for them; no series, no error. - No
phase: Readygate — backup history is Percona's own state, unrelated to the Nextcloud application's own readiness. - For each
(repo, type), the value is the max completion timestamp amongSucceededbackups only —Failed/Runningbackups are ignored regardless of their own timestamp, and a same-(repo, type)pair never regresses to an older value just because of list ordering. - A list failure (API error, Percona CRD briefly unavailable) leaves the metric at its last-known value, logs a warning, and does not block the rest of that maintenance tick.
Series lifecycle:
- The series are removed when the instance is deleted (
on_delete) — this is the first per-instance Prometheus series cleanup in this operator; every other per-instance gauge on this page (nextcloudinstance_info,*_condition,helmrelease_ready) is instead swept wholesale by the periodic background collector rather than cleared on deletion, so don't expect that same wholesale-sweep behavior here. - A managed-DB instance that has never had a
Succeededbackup has no series at all — there's no "value of 0" or placeholder. Absence is itself the signal: alert onabsent()for the fleet-wide case (below), not on a low/zero value, since a low value never occurs — the choice is "a real timestamp" or "no series."
Backup freshness (configuration bundle)¶
Unix timestamp of the last successful upload of the encrypted configuration-files bundle. One series per instance — no repo/type split, because there is exactly one bundle per instance per day.
- Set on the same per-instance maintenance timer, ahead of the
windowStartgate, for the same "visibility, not application" reason as the pgBackRest gauge above. - Applies to every instance, managed database or not: the configuration bundle is about
config/, the encryption keys and the operator-owned Secrets, none of which depend on who runs the database. - No series at all until an instance has had one successful bundle — absence is the signal, exactly as above. An instance with no resolvable backup bucket, or no encryption key, never produces one.
- A failed run does not move or clear the value: the last good bundle is still in the bucket. Watch the
ConfigBackupFailedevent andstatus.configBackup.lastFailureReasonfor the failure itself. - Series are removed when the instance is deleted, and also swept by the background collector against the live instance set — so a deletion force-completed by removing the finalizer cannot leave a frozen timestamp firing a stale-backup alert forever.
- alert: NextcloudConfigBackupStale
expr: time() - nextcloud_operator_config_backup_timestamp_seconds > 48 * 3600
for: 1h
labels: {severity: warning}
annotations:
summary: "No configuration backup for {{ $labels.namespace }}/{{ $labels.instance }} in 48h"
description: "A database backup without config/ (instanceid, secret, passwordsalt) cannot be restored."
Post-upgrade app convergence¶
1 while a completed version upgrade still has declared apps that are not enabled — the operator's convergence loop is retrying (see Upgrades → Post-upgrade app convergence). This is the alertable signal for that state: the retry loop has no attempt cap and can legitimately run for days while the app store catches up, so the thing an on-call human needs is a list of which instances are still waiting, not a log line that scrolled past at 03:00.
Series lifecycle — the same "absence is the signal" shape as the backup gauges above:
- No series at all for an instance that is not waiting. The series is removed on
Completed, on acceptance via the annotation, and on instance deletion — never zeroed. A stuck-at-0 series still matches most alert expressions and would accumulate one row per instance that ever upgraded. - Re-seeded from
status.upgradeby the 60s background collector. The in-process registry starts empty after an operator restart while the pending state lives in the CR, so without this a restart would silently clear every pending alert in the fleet. The sweep is the authority in both directions: an instance that completed while the operator was down loses its stale series on the next tick. - Also pruned against the live instance set, so a deletion force-completed by removing the finalizer cannot leave the alert firing for a resource that no longer exists.
Pair it with the condition, which carries the why:
kubectl get nci $NAME -n $NS -o jsonpath='{.status.conditions[?(@.type=="AppsHealthy")]}' | jq
kubectl get nci $NAME -n $NS -o jsonpath='{.status.upgrade.pendingApps}' | jq
- alert: NextcloudUpgradeAppsPending
expr: nextcloud_operator_upgrade_apps_pending == 1
for: 6h
labels: {severity: warning}
annotations:
summary: "Apps still disabled after upgrade on {{ $labels.namespace }}/{{ $labels.instance }}"
description: >-
The operator is retrying automatically and heals as soon as a compatible release
is published — 6h of patience is deliberate. Check status.upgrade.maintenanceHeld:
if true, a critical app (SSO) is pending and users are still locked out.
# Users locked out: the flow has held maintenance mode for longer than its own bound.
- alert: NextcloudUpgradeStuck
expr: nextcloud_operator_nextcloudinstance_condition{type="UpgradeStuck"} == 1
for: 15m
labels: {severity: critical}
annotations:
summary: "Upgrade not progressing on {{ $labels.namespace }}/{{ $labels.name }}"
Example alert: PgBackRestBackupStale¶
The client's proposed threshold rule (>36h without a successful backup), plus an absent() companion for the "never had a successful backup at all" case:
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: pgbackrest-backup-freshness
namespace: monitoring
spec:
groups:
- name: pgbackrest-backup-freshness
rules:
- alert: PgBackRestBackupStale
expr: time() - max by (namespace) (pgbackrest_last_backup_completion_timestamp_seconds) > 36 * 3600
for: 15m
labels:
severity: critical
annotations:
summary: "pgBackRest backup is stale for namespace {{ $labels.namespace }}"
description: "No successful pgBackRest backup completed in over 36h for {{ $labels.namespace }}."
- alert: PgBackRestMetricAbsent
expr: absent(pgbackrest_last_backup_completion_timestamp_seconds)
for: 15m
labels:
severity: critical
annotations:
summary: "No pgbackrest_last_backup_completion_timestamp_seconds series exist at all"
description: "Either no managed-DB instance has ever completed a successful pgBackRest backup, or the metric pipeline itself is broken (RBAC, scrape config, operator down). Check operator logs and the perconapgbackups RBAC grant before assuming every backup genuinely failed."
A known limitation of the absent() companion, stated plainly rather than glossed over: absent() can only tell you the metric has zero series anywhere — it cannot enumerate which specific managed-DB instances are missing a series, because it has no way to know the expected instance set. If one managed-DB instance among many has simply never completed a backup, its (namespace, instance) combination is silently absent from the metric with no per-instance alert firing on that fact alone — only the max by (namespace) threshold rule above would eventually catch it (once whatever did age past 36h, or once you notice the panel is empty for that instance). Closing this gap generically would need a join against a metric that enumerates "which instances have database.managed: true" — the operator doesn't currently expose one; nextcloudinstance_info's labels stop at profile, not database mode. Worth a follow-up if per-instance "backup never ran even once" detection turns out to matter in practice, but out of scope for this metric alone.
# Quick check against a specific instance without waiting for the alert
kubectl get nci my-instance -n my-namespace -o jsonpath='{.spec.database.managed}'
curl -s http://<operator-pod>:9090/metrics | grep pgbackrest_last_backup
Dashboards¶
Overview (uid: nextcloud-operator-overview)¶
Fleet-wide view. Panels:
- Totals (Nextclouds, NCIs, Ready, Failed)
- Phase distribution (pie)
- Reconciliation rate per resource/result
- Error rate per resource/stage
instance_ready_duration_secondsp50/p95/p99- Operation p95 by operation
- Pool replicas + assignment rate
- HelmRelease Ready/Not-Ready/Unknown counts
- Failed NextcloudInstances table
- Tenant ↔ Instance Assignment table (joins tenant URL/name to instance via
assignment_info)
Detail (uid: nextcloud-operator-detail)¶
Per-instance drill-down. Template variables: $namespace, $instance.
Panels are filtered by these variables where per-instance labels exist (info
and condition gauges, HelmRelease gauge). Reconcile and error series are shown
at the operator level for context — they don't carry name labels by design
to keep cardinality bounded.
Tuning¶
metrics.collectorInterval(Helm) /METRICS_COLLECTOR_INTERVAL(env): Background collector frequency in seconds. Default 60. Drop to 15–30 if you want faster dashboard refresh; raise to 300 for large fleets.metrics.serviceMonitor.interval: Prometheus scrape interval. Default 60s.
Cardinality notes¶
Per-instance labels exist only on state gauges (*_info, *_condition,
helmrelease_ready, assignment_info) and on pgbackrest_last_backup_completion_timestamp_seconds.
pool_instance_info is per-instance too, but only for pool replicas — its cardinality is
bounded by total pool capacity across pools (spec.replicas summed), not overall fleet
size, and it carries no tenant-identifying labels (no nextcloud_url/nextcloud_name).
Counters and histograms carry
resource / operation / stage / pool / profile — bounded by the number
of pools, profiles and operation kinds, not by the number of tenants. This keeps
scrape volume flat as fleet size grows.
assignment_info emits exactly one series per NextcloudInstance, so its
cardinality equals fleet size. The nextcloud_url label is high-cardinality but
bounded by tenant count and only changes when a customer renames their URL —
suitable for table panels and ad-hoc joins, not for use in PromQL by ()
groupings.
nextcloud_operator_config_backup_timestamp_seconds emits at most one series per
instance, so its cardinality is bounded by fleet size and carries no
tenant-identifying labels.
pgbackrest_last_backup_completion_timestamp_seconds is bounded by managed-DB
instance count × (repo, type) combinations per instance (from 0.22.0, two
repos — the local volume repo1 and the off-cluster repo2 — × 1-2 types) — smaller than fleet size, not larger, since unmanaged-DB instances
contribute nothing. Unlike every other per-instance gauge on this page, it's
also explicitly cleared on instance deletion rather than swept by the
background collector — see Backup freshness (pgBackRest).
Troubleshooting¶
- Metrics endpoint returns 404:
METRICS_ENABLEDis nottrueon the Deployment, or the collector server failed to start (check operator logs for "Prometheus metrics server started"). - ServiceMonitor not picked up: check that the label selector on the
Service (
app.kubernetes.io/name: nextcloud-operator) matches what the ServiceMonitor expects, and that the Prometheus instance'sserviceMonitorSelectorallows the chart's labels. - Dashboards don't appear in Grafana: the sidecar only scans ConfigMaps
labelled
grafana_dashboard: "1". Override withmetrics.grafana.dashboardLabelsif your Grafana uses a different label.