Skip to Content
GuidesMonitoring Setup

Monitoring Setup

Pilot exposes Prometheus metrics at /metrics. This guide covers measurement coverage in Pilot and setting up Prometheus scraping and Grafana dashboards.

Storage and Follower Lag Coverage

When one broker is offline, Pilot continues refreshing valid storage and follower-lag measurements from the other brokers. Coverage states how many expected brokers supplied usable measurements, for example 5/6 brokers. A missing measurement alone does not establish that its broker is offline; the broker health observation supplies that status.

MeasurementDisplay when coverage is incomplete
Broker with a valid measurementCurrent value
Broker without a usable measurementUnavailable
Cluster totalUnavailable, with incomplete coverage
Topic or partition aggregateAvailable only when its required broker measurements are complete

Storage and follower lag can have different coverage. Lag also depends on usable replica and leader information; a broker can provide valid storage measurements while its lag remains unavailable. Missing readings never count as zero.

Complete storage byte coverage does not guarantee a segment-count estimate: unknown segment counts remain unavailable (null in the API) while valid byte measurements stay visible.

If collection fails for the whole request, retained values are marked stale. Source timestamps identify when measurements and topology were observed, rather than when the page refreshed. Partial monitoring data does not satisfy the complete-data checks used for sampling or reassignment safety.

Prometheus Configuration

Add Pilot as a scrape target in your prometheus.yml:

scrape_configs: - job_name: pilot scrape_interval: 15s static_configs: - targets: ['pilot:8080']

For Kubernetes with service discovery:

scrape_configs: - job_name: pilot kubernetes_sd_configs: - role: pod relabel_configs: - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape] action: keep regex: true - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_port] action: replace target_label: __address__ regex: (.+) replacement: ${1}

Key Metrics to Watch

Cluster Health

MetricAlert ConditionDescription
pilot_cluster_under_replicated_partitions> 0 for 5mPartitions with ISR < replicas
pilot_cluster_offline_partitions> 0 for 1mPartitions with no leader
pilot_cluster_brokers_unavailable> 0 for 5m, with a current observationKnown brokers absent from Kafka’s broker roster
pilot_cluster_broker_observation_current== 0 for 1mBroker discovery is unavailable
pilot_cluster_broker_observation_expires_timestamp_seconds< time() for 1mBroker observation expired

Broker counts retain their last successful values when discovery fails. Gate broker-count alerts on both pilot_cluster_broker_observation_current == 1 and pilot_cluster_broker_observation_expires_timestamp_seconds > time(). Keep a separate alert for failed or expired observations so a monitoring problem cannot appear healthy.

Operations

MetricAlert ConditionDescription
pilot_reassignment_active> 0 for 2hReassignment stuck
pilot_sampler_last_successful_collection_timestampstale > 60sSampler not collecting
pilot_license_expiry_timestamp_seconds< 7 days from nowLicense expiring soon
pilot_audit_events_failure_total{error_type="dropped"}increase > 0 in 15mAudit events lost

Performance

MetricAlert ConditionDescription
pilot_http_requests_total{status_code=~"5.."}error rate > 1%API errors
pilot_http_request_duration_secondsP99 > 5sSlow API responses

Alert Rules

Example Prometheus alert rules:

groups: - name: pilot rules: - alert: KafkaUnderReplicatedPartitions expr: pilot_cluster_under_replicated_partitions > 0 for: 5m labels: severity: warning annotations: summary: "Under-replicated partitions detected" description: "{{ $value }} partitions have ISR < replicas" - alert: KafkaOfflinePartitions expr: pilot_cluster_offline_partitions > 0 for: 1m labels: severity: critical annotations: summary: "Offline partitions detected" description: "{{ $value }} partitions have no leader" - alert: PilotSamplerStale expr: time() - pilot_sampler_last_successful_collection_timestamp > 60 for: 2m labels: severity: warning annotations: summary: "Pilot sampler not collecting" description: "No successful collection in over 60 seconds" - alert: PilotLicenseExpiring expr: pilot_license_expiry_timestamp_seconds - time() < 7 * 24 * 3600 labels: severity: warning annotations: summary: "Pilot license expiring within 7 days" - alert: PilotReassignmentStuck expr: pilot_reassignment_active > 0 for: 2h labels: severity: warning annotations: summary: "Partition reassignment running for over 2 hours" - alert: PilotSelfHealingNotRunning expr: | (time() - pilot_selfhealing_last_success_timestamp_seconds{type="critical"} > 1800) and pilot_cluster_under_replicated_partitions > 0 labels: severity: warning annotations: summary: "Critical self-healing not running while URPs exist" - alert: PilotAuditEventsLost expr: increase(pilot_audit_events_failure_total{error_type="dropped"}[15m]) > 0 labels: severity: warning annotations: summary: "Audit events were lost"

See Prometheus Metrics for PromQL query examples.

Last updated on