Monitoring Setup
Pilot exposes Prometheus metrics at /metrics. This guide covers measurement coverage in Pilot and setting up Prometheus scraping and Grafana dashboards.
Storage and Follower Lag Coverage
When one broker is offline, Pilot continues refreshing valid storage and follower-lag measurements from the other brokers. Coverage states how many expected brokers supplied usable measurements, for example 5/6 brokers. A missing measurement alone does not establish that its broker is offline; the broker health observation supplies that status.
| Measurement | Display when coverage is incomplete |
|---|---|
| Broker with a valid measurement | Current value |
| Broker without a usable measurement | Unavailable |
| Cluster total | Unavailable, with incomplete coverage |
| Topic or partition aggregate | Available only when its required broker measurements are complete |
Storage and follower lag can have different coverage. Lag also depends on usable replica and leader information; a broker can provide valid storage measurements while its lag remains unavailable. Missing readings never count as zero.
Complete storage byte coverage does not guarantee a segment-count estimate: unknown segment counts remain unavailable (null in the API) while valid byte measurements stay visible.
If collection fails for the whole request, retained values are marked stale. Source timestamps identify when measurements and topology were observed, rather than when the page refreshed. Partial monitoring data does not satisfy the complete-data checks used for sampling or reassignment safety.
Prometheus Configuration
Add Pilot as a scrape target in your prometheus.yml:
scrape_configs:
- job_name: pilot
scrape_interval: 15s
static_configs:
- targets: ['pilot:8080']For Kubernetes with service discovery:
scrape_configs:
- job_name: pilot
kubernetes_sd_configs:
- role: pod
relabel_configs:
- source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]
action: keep
regex: true
- source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_port]
action: replace
target_label: __address__
regex: (.+)
replacement: ${1}Key Metrics to Watch
Cluster Health
| Metric | Alert Condition | Description |
|---|---|---|
pilot_cluster_under_replicated_partitions | > 0 for 5m | Partitions with ISR < replicas |
pilot_cluster_offline_partitions | > 0 for 1m | Partitions with no leader |
pilot_cluster_brokers_unavailable | > 0 for 5m, with a current observation | Known brokers absent from Kafka’s broker roster |
pilot_cluster_broker_observation_current | == 0 for 1m | Broker discovery is unavailable |
pilot_cluster_broker_observation_expires_timestamp_seconds | < time() for 1m | Broker observation expired |
Broker counts retain their last successful values when discovery fails. Gate broker-count alerts on both pilot_cluster_broker_observation_current == 1 and pilot_cluster_broker_observation_expires_timestamp_seconds > time(). Keep a separate alert for failed or expired observations so a monitoring problem cannot appear healthy.
Operations
| Metric | Alert Condition | Description |
|---|---|---|
pilot_reassignment_active | > 0 for 2h | Reassignment stuck |
pilot_sampler_last_successful_collection_timestamp | stale > 60s | Sampler not collecting |
pilot_license_expiry_timestamp_seconds | < 7 days from now | License expiring soon |
pilot_audit_events_failure_total{error_type="dropped"} | increase > 0 in 15m | Audit events lost |
Performance
| Metric | Alert Condition | Description |
|---|---|---|
pilot_http_requests_total{status_code=~"5.."} | error rate > 1% | API errors |
pilot_http_request_duration_seconds | P99 > 5s | Slow API responses |
Alert Rules
Example Prometheus alert rules:
groups:
- name: pilot
rules:
- alert: KafkaUnderReplicatedPartitions
expr: pilot_cluster_under_replicated_partitions > 0
for: 5m
labels:
severity: warning
annotations:
summary: "Under-replicated partitions detected"
description: "{{ $value }} partitions have ISR < replicas"
- alert: KafkaOfflinePartitions
expr: pilot_cluster_offline_partitions > 0
for: 1m
labels:
severity: critical
annotations:
summary: "Offline partitions detected"
description: "{{ $value }} partitions have no leader"
- alert: PilotSamplerStale
expr: time() - pilot_sampler_last_successful_collection_timestamp > 60
for: 2m
labels:
severity: warning
annotations:
summary: "Pilot sampler not collecting"
description: "No successful collection in over 60 seconds"
- alert: PilotLicenseExpiring
expr: pilot_license_expiry_timestamp_seconds - time() < 7 * 24 * 3600
labels:
severity: warning
annotations:
summary: "Pilot license expiring within 7 days"
- alert: PilotReassignmentStuck
expr: pilot_reassignment_active > 0
for: 2h
labels:
severity: warning
annotations:
summary: "Partition reassignment running for over 2 hours"
- alert: PilotSelfHealingNotRunning
expr: |
(time() - pilot_selfhealing_last_success_timestamp_seconds{type="critical"} > 1800)
and pilot_cluster_under_replicated_partitions > 0
labels:
severity: warning
annotations:
summary: "Critical self-healing not running while URPs exist"
- alert: PilotAuditEventsLost
expr: increase(pilot_audit_events_failure_total{error_type="dropped"}[15m]) > 0
labels:
severity: warning
annotations:
summary: "Audit events were lost"See Prometheus Metrics for PromQL query examples.