Skip to content

Monitoring

The operator exports Prometheus metrics about its own reconciliation loops and the state of every managed Nextcloud, NextcloudInstance, NextcloudPool and HelmRelease. A ServiceMonitor and two Grafana dashboards (Overview + Detail) ship with the chart and as optional flat-YAML manifests.

Metrics are opt-in on both install paths.

Quick start

Helm install

helm upgrade --install nextcloud-operator ./chart \
  --set metrics.enabled=true \
  --set metrics.serviceMonitor.enabled=true \
  --set metrics.grafana.enabled=true

This creates:

  • Container port metrics/9090 and environment variables on the Deployment
  • A Service exposing port 9090
  • A ServiceMonitor (requires the Prometheus Operator CRDs)
  • Two ConfigMaps, one per dashboard, labelled grafana_dashboard: "1" so the Grafana sidecar picks them up automatically

Flat YAML install

# 1. Enable metrics in the operator deployment (edit METRICS_ENABLED to "true")
kubectl -n nextcloud-operator-system set env deploy/nextcloud-operator METRICS_ENABLED=true

# 2. Apply the monitoring stack
kubectl apply -f deploy/monitoring.yaml \
              -f deploy/dashboard-overview.yaml \
              -f deploy/dashboard-detail.yaml

deploy/monitoring.yaml contains the Service and ServiceMonitor. The two dashboard ConfigMaps are regenerated from the chart via make sync-dashboards.

Prerequisites

Component Required for
Prometheus + prometheus-operator CRDs ServiceMonitor scraping
Grafana with the sidecar dashboard loader Auto-loading the shipped dashboards

If you run a stand-alone Grafana, import the JSON from chart/dashboards/ directly.

Exposed metrics

All metrics are prefixed with nextcloud_operator_, with one deliberate exception: pgbackrest_last_backup_completion_timestamp_seconds (see Backup freshness (pgBackRest) below) is unprefixed on purpose, so it matches the client's existing *pgback* alert discovery.

State gauges (updated by the background collector every 60s)

Metric Type Labels Meaning
nextclouds_total Gauge phase Nextcloud CR count by phase
nextcloudinstances_total Gauge phase NextcloudInstance CR count by phase
nextcloudpools_total Gauge phase NextcloudPool count by phase
nextcloudprofiles_total Gauge — NextcloudProfile count
pool_replicas Gauge pool, type (desired/ready/unassigned/assigned) Per-pool replica breakdown
pool_instance_info Gauge pool, instance, namespace, phase, assigned (true/false) One series per pool replica (status.instances[])
nextcloud_info Gauge name, namespace, phase, instance_name, instance_namespace, profile, url Per-CR identity (always 1)
nextcloudinstance_info Gauge name, namespace, phase, assigned_to (owning Nextcloud name, empty when spare), pool, profile, url Per-NCI identity (always 1)
nextcloudpool_info Gauge name, phase, desired, ready, unassigned, assigned Per-pool identity (always 1)
assignment_info Gauge nextcloud_namespace, nextcloud_name, nextcloud_url, instance_namespace, instance_name, profile, pool, state (assigned/spare) One series per NextcloudInstance. Lets dashboards JOIN tenant identifiers (URL, name) onto K8s workload metrics by instance_namespace. Spare pool instances appear with empty nextcloud_* labels.
nextcloud_condition Gauge name, namespace, type Condition state (1/0/-1 = True/False/Unknown)
nextcloudinstance_condition Gauge name, namespace, type Condition state (1/0/-1 = True/False/Unknown)
helmrelease_ready Gauge namespace, name HelmRelease Ready condition (1/0/-1)

pool_replicas and nextcloudpool_info — fixed in 0.21.0 (#7865). Both read status.ready/status.unassigned/status.assigned (the actual NextcloudPool CRD status field names) from the collector's 60s sweep. Before 0.21.0 the collector read status.readyReplicas/unassignedReplicas/assignedReplicas, which don't exist on the CRD — so the ready/unassigned/assigned label values and pool_replicas series were always 0, regardless of the pool's actual state. desired (from spec.replicas) was unaffected. If a CapacityPoolStarved-style alert on these gauges never fired even during a real pool exhaustion, this is why; it's now sourced correctly.

pool_instance_info — new in 0.21.0. Emits one series per entry in status.instances[], labelled with that replica's phase and whether it's currently assigned to a Nextcloud. Unlike pool_replicas' aggregate counts, this lets you distinguish an idle-Ready spare (healthy, waiting to be matched) from an idle replica stuck in a non-Ready phase (wedged) — the failure mode pool_replicas alone can't surface, since it only counts, it doesn't name. Cleared and repopulated wholesale every collector sweep, same lifecycle as pool_replicas and nextcloudpool_info — a replica removed from the pool (scaled down, deleted) drops out of the metric on the next sweep, no manual cleanup needed.

Example query — pool replicas that are unassigned and not Ready (idle-wedged, consuming pool capacity without being usable):

nextcloud_operator_pool_instance_info{assigned="false", phase!="Ready"}

Event counters

Metric Type Labels Meaning
reconcile_total Counter resource, result (success/error/temporary_error/permanent_error) One increment per kopf handler invocation
errors_total Counter resource, stage (validation/db_provision/helmrelease/occ/maintenance/…) Categorised error accounting
pool_scale_total Counter pool, direction (up/down) Pool instance create/delete events
pool_assignment_total Counter pool, result (success/conflict/no_match) Outcome of pool match attempts
maintenance_task_total Counter task, result (success/error) Periodic and post-upgrade OCC task runs

Latency histograms

Metric Labels Covers
operation_duration_seconds operation, result Generic operation timer; used for db_provision, helmrelease_create_or_update, occ_command, maintenance_task
instance_ready_duration_seconds profile Seconds from NCI creation to phase=Ready
nextcloud_assignment_duration_seconds pool Seconds from Nextcloud creation to first pool assignment

Backup freshness (pgBackRest)

pgbackrest_last_backup_completion_timestamp_seconds{namespace, instance, repo, type}

Unix timestamp of the most recent Succeeded pgBackRest backup completion, one series per (repo, type) per instance. type is full/incremental/differential. This is the metric OpenProject #7835 exists to expose — the Percona operator runs pgBackRest itself, but nothing previously read its results back into Prometheus, so a PgBackRestBackupStale alert had no data to fire on.

Source and update cadence — different from the other gauges on this page, worth calling out explicitly:

  • Set on the per-instance maintenance timer (TIMER_MAINTENANCE_INTERVAL, default 15 min), not the 60s fleet-wide background collector.
  • Runs before the spec.maintenance.windowStart gate — a windowless instance still gets this metric updated, same "visibility, not application" rationale as UpdateAvailable and the apps-health check (this feature is pure observation of Percona's own PerconaPGBackup CRs; it never mutates anything).
  • Managed-database instances only (spec.database.managed: true) — unmanaged/external databases have no PerconaPGBackup CRs to read, so the check is skipped entirely for them; no series, no error.
  • No phase: Ready gate — backup history is Percona's own state, unrelated to the Nextcloud application's own readiness.
  • For each (repo, type), the value is the max completion timestamp among Succeeded backups only — Failed/Running backups are ignored regardless of their own timestamp, and a same-(repo, type) pair never regresses to an older value just because of list ordering.
  • A list failure (API error, Percona CRD briefly unavailable) leaves the metric at its last-known value, logs a warning, and does not block the rest of that maintenance tick.

Series lifecycle:

  • The series are removed when the instance is deleted (on_delete) — this is the first per-instance Prometheus series cleanup in this operator; every other per-instance gauge on this page (nextcloudinstance_info, *_condition, helmrelease_ready) is instead swept wholesale by the periodic background collector rather than cleared on deletion, so don't expect that same wholesale-sweep behavior here.
  • A managed-DB instance that has never had a Succeeded backup has no series at all — there's no "value of 0" or placeholder. Absence is itself the signal: alert on absent() for the fleet-wide case (below), not on a low/zero value, since a low value never occurs — the choice is "a real timestamp" or "no series."

Backup freshness (configuration bundle)

nextcloud_operator_config_backup_timestamp_seconds{namespace, instance}

Unix timestamp of the last successful upload of the encrypted configuration-files bundle. One series per instance — no repo/type split, because there is exactly one bundle per instance per day.

  • Set on the same per-instance maintenance timer, ahead of the windowStart gate, for the same "visibility, not application" reason as the pgBackRest gauge above.
  • Applies to every instance, managed database or not: the configuration bundle is about config/, the encryption keys and the operator-owned Secrets, none of which depend on who runs the database.
  • No series at all until an instance has had one successful bundle — absence is the signal, exactly as above. An instance with no resolvable backup bucket, or no encryption key, never produces one.
  • A failed run does not move or clear the value: the last good bundle is still in the bucket. Watch the ConfigBackupFailed event and status.configBackup.lastFailureReason for the failure itself.
  • Series are removed when the instance is deleted, and also swept by the background collector against the live instance set — so a deletion force-completed by removing the finalizer cannot leave a frozen timestamp firing a stale-backup alert forever.
- alert: NextcloudConfigBackupStale
  expr: time() - nextcloud_operator_config_backup_timestamp_seconds > 48 * 3600
  for: 1h
  labels: {severity: warning}
  annotations:
    summary: "No configuration backup for {{ $labels.namespace }}/{{ $labels.instance }} in 48h"
    description: "A database backup without config/ (instanceid, secret, passwordsalt) cannot be restored."

Post-upgrade app convergence

nextcloud_operator_upgrade_apps_pending{namespace, instance}

1 while a completed version upgrade still has declared apps that are not enabled — the operator's convergence loop is retrying (see Upgrades → Post-upgrade app convergence). This is the alertable signal for that state: the retry loop has no attempt cap and can legitimately run for days while the app store catches up, so the thing an on-call human needs is a list of which instances are still waiting, not a log line that scrolled past at 03:00.

Series lifecycle — the same "absence is the signal" shape as the backup gauges above:

  • No series at all for an instance that is not waiting. The series is removed on Completed, on acceptance via the annotation, and on instance deletion — never zeroed. A stuck-at-0 series still matches most alert expressions and would accumulate one row per instance that ever upgraded.
  • Re-seeded from status.upgrade by the 60s background collector. The in-process registry starts empty after an operator restart while the pending state lives in the CR, so without this a restart would silently clear every pending alert in the fleet. The sweep is the authority in both directions: an instance that completed while the operator was down loses its stale series on the next tick.
  • Also pruned against the live instance set, so a deletion force-completed by removing the finalizer cannot leave the alert firing for a resource that no longer exists.

Pair it with the condition, which carries the why:

kubectl get nci $NAME -n $NS -o jsonpath='{.status.conditions[?(@.type=="AppsHealthy")]}' | jq
kubectl get nci $NAME -n $NS -o jsonpath='{.status.upgrade.pendingApps}' | jq
- alert: NextcloudUpgradeAppsPending
  expr: nextcloud_operator_upgrade_apps_pending == 1
  for: 6h
  labels: {severity: warning}
  annotations:
    summary: "Apps still disabled after upgrade on {{ $labels.namespace }}/{{ $labels.instance }}"
    description: >-
      The operator is retrying automatically and heals as soon as a compatible release
      is published — 6h of patience is deliberate. Check status.upgrade.maintenanceHeld:
      if true, a critical app (SSO) is pending and users are still locked out.

# Users locked out: the flow has held maintenance mode for longer than its own bound.
- alert: NextcloudUpgradeStuck
  expr: nextcloud_operator_nextcloudinstance_condition{type="UpgradeStuck"} == 1
  for: 15m
  labels: {severity: critical}
  annotations:
    summary: "Upgrade not progressing on {{ $labels.namespace }}/{{ $labels.name }}"

Example alert: PgBackRestBackupStale

The client's proposed threshold rule (>36h without a successful backup), plus an absent() companion for the "never had a successful backup at all" case:

apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: pgbackrest-backup-freshness
  namespace: monitoring
spec:
  groups:
    - name: pgbackrest-backup-freshness
      rules:
        - alert: PgBackRestBackupStale
          expr: time() - max by (namespace) (pgbackrest_last_backup_completion_timestamp_seconds) > 36 * 3600
          for: 15m
          labels:
            severity: critical
          annotations:
            summary: "pgBackRest backup is stale for namespace {{ $labels.namespace }}"
            description: "No successful pgBackRest backup completed in over 36h for {{ $labels.namespace }}."

        - alert: PgBackRestMetricAbsent
          expr: absent(pgbackrest_last_backup_completion_timestamp_seconds)
          for: 15m
          labels:
            severity: critical
          annotations:
            summary: "No pgbackrest_last_backup_completion_timestamp_seconds series exist at all"
            description: "Either no managed-DB instance has ever completed a successful pgBackRest backup, or the metric pipeline itself is broken (RBAC, scrape config, operator down). Check operator logs and the perconapgbackups RBAC grant before assuming every backup genuinely failed."

A known limitation of the absent() companion, stated plainly rather than glossed over: absent() can only tell you the metric has zero series anywhere — it cannot enumerate which specific managed-DB instances are missing a series, because it has no way to know the expected instance set. If one managed-DB instance among many has simply never completed a backup, its (namespace, instance) combination is silently absent from the metric with no per-instance alert firing on that fact alone — only the max by (namespace) threshold rule above would eventually catch it (once whatever did age past 36h, or once you notice the panel is empty for that instance). Closing this gap generically would need a join against a metric that enumerates "which instances have database.managed: true" — the operator doesn't currently expose one; nextcloudinstance_info's labels stop at profile, not database mode. Worth a follow-up if per-instance "backup never ran even once" detection turns out to matter in practice, but out of scope for this metric alone.

# Quick check against a specific instance without waiting for the alert
kubectl get nci my-instance -n my-namespace -o jsonpath='{.spec.database.managed}'
curl -s http://<operator-pod>:9090/metrics | grep pgbackrest_last_backup

Dashboards

Overview (uid: nextcloud-operator-overview)

Fleet-wide view. Panels:

  • Totals (Nextclouds, NCIs, Ready, Failed)
  • Phase distribution (pie)
  • Reconciliation rate per resource/result
  • Error rate per resource/stage
  • instance_ready_duration_seconds p50/p95/p99
  • Operation p95 by operation
  • Pool replicas + assignment rate
  • HelmRelease Ready/Not-Ready/Unknown counts
  • Failed NextcloudInstances table
  • Tenant ↔ Instance Assignment table (joins tenant URL/name to instance via assignment_info)

Detail (uid: nextcloud-operator-detail)

Per-instance drill-down. Template variables: $namespace, $instance. Panels are filtered by these variables where per-instance labels exist (info and condition gauges, HelmRelease gauge). Reconcile and error series are shown at the operator level for context — they don't carry name labels by design to keep cardinality bounded.

Tuning

  • metrics.collectorInterval (Helm) / METRICS_COLLECTOR_INTERVAL (env): Background collector frequency in seconds. Default 60. Drop to 15–30 if you want faster dashboard refresh; raise to 300 for large fleets.
  • metrics.serviceMonitor.interval: Prometheus scrape interval. Default 60s.

Cardinality notes

Per-instance labels exist only on state gauges (*_info, *_condition, helmrelease_ready, assignment_info) and on pgbackrest_last_backup_completion_timestamp_seconds. pool_instance_info is per-instance too, but only for pool replicas — its cardinality is bounded by total pool capacity across pools (spec.replicas summed), not overall fleet size, and it carries no tenant-identifying labels (no nextcloud_url/nextcloud_name). Counters and histograms carry resource / operation / stage / pool / profile — bounded by the number of pools, profiles and operation kinds, not by the number of tenants. This keeps scrape volume flat as fleet size grows.

assignment_info emits exactly one series per NextcloudInstance, so its cardinality equals fleet size. The nextcloud_url label is high-cardinality but bounded by tenant count and only changes when a customer renames their URL — suitable for table panels and ad-hoc joins, not for use in PromQL by () groupings.

nextcloud_operator_config_backup_timestamp_seconds emits at most one series per instance, so its cardinality is bounded by fleet size and carries no tenant-identifying labels.

pgbackrest_last_backup_completion_timestamp_seconds is bounded by managed-DB instance count × (repo, type) combinations per instance (from 0.22.0, two repos — the local volume repo1 and the off-cluster repo2 — × 1-2 types) — smaller than fleet size, not larger, since unmanaged-DB instances contribute nothing. Unlike every other per-instance gauge on this page, it's also explicitly cleared on instance deletion rather than swept by the background collector — see Backup freshness (pgBackRest).

Troubleshooting

  • Metrics endpoint returns 404: METRICS_ENABLED is not true on the Deployment, or the collector server failed to start (check operator logs for "Prometheus metrics server started").
  • ServiceMonitor not picked up: check that the label selector on the Service (app.kubernetes.io/name: nextcloud-operator) matches what the ServiceMonitor expects, and that the Prometheus instance's serviceMonitorSelector allows the chart's labels.
  • Dashboards don't appear in Grafana: the sidecar only scans ConfigMaps labelled grafana_dashboard: "1". Override with metrics.grafana.dashboardLabels if your Grafana uses a different label.