---
title: "Monitoring Setup - Pilot Docs"
description: "Set up Prometheus and Grafana monitoring for Pilot"
url: "https://docs.calinora.io/guides/monitoring-setup/"
---

# Monitoring Setup

Pilot exposes Prometheus metrics at `/metrics`. This guide covers measurement coverage in Pilot and setting up Prometheus scraping and Grafana dashboards.

## Storage and Follower Lag Coverage

When one broker is offline, Pilot continues refreshing valid storage and follower-lag measurements from the other brokers. Coverage states how many expected brokers supplied usable measurements, for example **5/6 brokers**. A missing measurement alone does not establish that its broker is offline; the broker health observation supplies that status.

| Measurement | Display when coverage is incomplete |
| - | - |
| Broker with a valid measurement | Current value |
| Broker without a usable measurement | Unavailable |
| Cluster total | Unavailable, with incomplete coverage |
| Topic or partition aggregate | Available only when its required broker measurements are complete |

Storage and follower lag can have different coverage. Lag also depends on usable replica and leader information; a broker can provide valid storage measurements while its lag remains unavailable. Missing readings never count as zero.

Complete storage byte coverage does not guarantee a segment-count estimate: unknown segment counts remain unavailable (`null` in the API) while valid byte measurements stay visible.

If collection fails for the whole request, retained values are marked stale. Source timestamps identify when measurements and topology were observed, rather than when the page refreshed. Partial monitoring data does not satisfy the complete-data checks used for sampling or reassignment safety.

## Prometheus Configuration

Add Pilot as a scrape target in your `prometheus.yml`:

```yaml
scrape_configs:
  - job_name: pilot
    scrape_interval: 15s
    static_configs:
      - targets: ['pilot:8080']
```

For Kubernetes with service discovery:

```yaml
scrape_configs:
  - job_name: pilot
    kubernetes_sd_configs:
      - role: pod
    relabel_configs:
      - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]
        action: keep
        regex: true
      - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_port]
        action: replace
        target_label: __address__
        regex: (.+)
        replacement: ${1}
```

## Key Metrics to Watch

### Cluster Health

| Metric | Alert Condition | Description |
| - | - | - |
| `pilot_cluster_under_replicated_partitions` | `> 0` for 5m | Partitions with ISR < replicas |
| `pilot_cluster_offline_partitions` | `> 0` for 1m | Partitions with no leader |
| `pilot_cluster_brokers_unavailable` | `> 0` for 5m, with a current observation | Known brokers absent from Kafka’s broker roster |
| `pilot_cluster_broker_observation_current` | `== 0` for 1m | Broker discovery is unavailable |
| `pilot_cluster_broker_observation_expires_timestamp_seconds` | `< time()` for 1m | Broker observation expired |

Broker counts retain their last successful values when discovery fails. Gate broker-count alerts on both `pilot_cluster_broker_observation_current == 1` and `pilot_cluster_broker_observation_expires_timestamp_seconds > time()`. Keep a separate alert for failed or expired observations so a monitoring problem cannot appear healthy.

### Operations

| Metric | Alert Condition | Description |
| - | - | - |
| `pilot_reassignment_active` | `> 0` for 2h | Reassignment stuck |
| `pilot_sampler_last_successful_collection_timestamp` | stale > 60s | Sampler not collecting |
| `pilot_license_expiry_timestamp_seconds` | < 7 days from now | License expiring soon |
| `pilot_audit_events_failure_total{error_type="dropped"}` | increase > 0 in 15m | Audit events lost |

### Performance

| Metric | Alert Condition | Description |
| - | - | - |
| `pilot_http_requests_total{status_code=~"5.."}` | error rate > 1% | API errors |
| `pilot_http_request_duration_seconds` | P99 > 5s | Slow API responses |

## Alert Rules

Example Prometheus alert rules:

```yaml
groups:
  - name: pilot
    rules:
      - alert: KafkaUnderReplicatedPartitions
        expr: pilot_cluster_under_replicated_partitions > 0
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "Under-replicated partitions detected"
          description: "{{ $value }} partitions have ISR < replicas"

      - alert: KafkaOfflinePartitions
        expr: pilot_cluster_offline_partitions > 0
        for: 1m
        labels:
          severity: critical
        annotations:
          summary: "Offline partitions detected"
          description: "{{ $value }} partitions have no leader"

      - alert: PilotSamplerStale
        expr: time() - pilot_sampler_last_successful_collection_timestamp > 60
        for: 2m
        labels:
          severity: warning
        annotations:
          summary: "Pilot sampler not collecting"
          description: "No successful collection in over 60 seconds"

      - alert: PilotLicenseExpiring
        expr: pilot_license_expiry_timestamp_seconds - time() < 7 * 24 * 3600
        labels:
          severity: warning
        annotations:
          summary: "Pilot license expiring within 7 days"

      - alert: PilotReassignmentStuck
        expr: pilot_reassignment_active > 0
        for: 2h
        labels:
          severity: warning
        annotations:
          summary: "Partition reassignment running for over 2 hours"

      - alert: PilotSelfHealingNotRunning
        expr: |
          (time() - pilot_selfhealing_last_success_timestamp_seconds{type="critical"} > 1800)
          and pilot_cluster_under_replicated_partitions > 0
        labels:
          severity: warning
        annotations:
          summary: "Critical self-healing not running while URPs exist"

      - alert: PilotAuditEventsLost
        expr: increase(pilot_audit_events_failure_total{error_type="dropped"}[15m]) > 0
        labels:
          severity: warning
        annotations:
          summary: "Audit events were lost"
```

See [Prometheus Metrics](https://docs.calinora.io/reference/metrics/#example-promql-queries) for PromQL query examples.
