Skip to Content
ReferencePrometheus Metrics

Prometheus Metrics

Pilot exposes Prometheus metrics at the /metrics endpoint. All application metrics use the pilot_ namespace prefix.

Standard go_* runtime metrics (goroutines, GC, memory) and process_* metrics (CPU, RSS, open file descriptors) are included automatically.

HTTP Layer

RED (Rate, Errors, Duration) metrics captured by middleware for every API request. The route label uses the route template (e.g. /api/v1/topics/{topic}/config) to keep cardinality bounded.

MetricTypeLabelsDescription
pilot_http_request_duration_secondsHistogrammethod, route, status_codeDuration of HTTP requests. Buckets: 1ms-10s.
pilot_http_requests_totalCountermethod, route, status_codeTotal number of HTTP requests.
pilot_http_request_size_bytesHistogrammethod, routeSize of HTTP request bodies. Buckets: 100B-1MB.
pilot_http_response_size_bytesHistogrammethod, routeSize of HTTP response bodies. Buckets: 1KB-10MB.
pilot_http_inflight_requestsGauge-Number of currently in-flight HTTP requests.

Kafka Operations

Latency and error tracking for Kafka admin/client calls. Instrumented operations include get_cluster_info, get_watermarks, describe_log_dirs, describe_config, update_topic_config, get_broker_config, alter_partition_reassignments, preferred_leader_election, describe_client_quotas, alter_client_quotas, and alter_replica_log_dirs.

MetricTypeLabelsDescription
pilot_kafka_operation_duration_secondsHistogramoperation, statusDuration of a Kafka operation. status is success or error. Buckets: 10ms-60s.
pilot_kafka_operation_errors_totalCounteroperation, error_typeTotal Kafka operation errors. error_type: timeout, auth, connection, kafka_error, other.

Cluster Health

The background health monitor updates these gauges on each check cycle (default 30s). The /api/v1/ready readiness probe requires a current broker observation with at least one available broker. A failed or expired observation makes readiness unavailable without declaring that every broker is offline. Rack assessment status does not change readiness.

MetricTypeLabelsDescription
pilot_cluster_brokers_totalGauge-Total known brokers in the cluster.
pilot_cluster_brokers_availableGauge-Brokers present in the last successful discovery response.
pilot_cluster_brokers_unavailableGauge-Known brokers absent from the last successful discovery response.
pilot_cluster_broker_observation_currentGauge-Whether the latest broker discovery attempt succeeded: 1 or 0. Also check the expiry timestamp.
pilot_cluster_broker_observation_timestamp_secondsGauge-Unix timestamp of the last successful broker observation, or 0 before the first success.
pilot_cluster_broker_observation_expires_timestamp_secondsGauge-Unix timestamp when that observation expires, or 0 before the first success.
pilot_cluster_brokers_maintenanceGauge-Brokers currently in maintenance mode.
pilot_cluster_under_replicated_partitionsGauge-Partitions where len(ISR) < len(Replicas).
pilot_cluster_offline_partitionsGauge-Partitions with no leader or empty ISR.
pilot_cluster_urps_from_offline_brokersGauge-Under-replicated partitions caused by offline brokers.
pilot_cluster_urps_from_follower_lagGauge-Under-replicated partitions caused by follower lag (all replicas online but not in ISR).
pilot_cluster_not_rack_aware_partitionsGauge-Exact count of assigned partitions that cannot retain minimum ISR after any single rack loss. Emitted only for a complete, current assessment.
pilot_cluster_rack_assessment_statusGaugestatusOne for the active status, zero for the others: complete, partial, unavailable.
pilot_cluster_rack_assessment_partitionsGaugecoveragePartition counts for total, evaluated, and unknown. All three series are omitted when the total scope is unknown.
pilot_cluster_rack_assessment_known_violationsGauge-Proven violations among evaluated partitions. A lower bound when partial; omitted when unavailable.
pilot_cluster_rack_assessment_unknown_partitionsGaugereasonPositive counts of unknown partition evidence by bounded reason code. Use coverage series for the assessment denominator.
pilot_cluster_rack_assessment_timestamp_secondsGaugesourceUnix seconds for assessment, topology, config_oldest, config_newest, and valid_until. Unknown timestamps are omitted.
pilot_cluster_rack_assessment_config_refresh_failuresGauge-Topic configuration reads that failed or lacked a valid effective minimum ISR in the completed assessment. Omitted before any assessment.
pilot_cluster_health_check_duration_secondsHistogram-Duration of each health check cycle. Buckets: 100ms-10s.

Broker discovery is independent of partition leader-election errors. A failed discovery retains the last successful broker counts and timestamps. Interpret these counts as current only when pilot_cluster_broker_observation_current == 1 and pilot_cluster_broker_observation_expires_timestamp_seconds > time(). Alert separately when observations fail or expire. Discovery reports Kafka’s broker roster, not a direct connectivity test to every broker.

Rack health checks assigned replicas against effective minimum ISR after losing the rack containing the most replicas. It does not establish current ISR availability. Missing configuration, rack labels, assignments or usable topology remain unknown. A partial assessment can contain proven violations alongside unknown partitions.

An absent exact-count series means the exact count is unavailable. Do not replace it with zero in queries or alerts. Check pilot_cluster_rack_assessment_status{status="complete"} before interpreting zero as a complete assessment without known violations. Expiry removes the exact series at scrape time even if no new health pass has completed.

Source timestamps retain the original observations. valid_until is a deadline, not an observation time. Configuration validity follows the cache refresh TTL plus one health interval and the five-second rack stage budget (65s by default). Topology validity follows the configured metadata refresh interval plus request and retry-backoff budget (65s with the application’s default 10s metadata interval). A failed configuration refresh leaves affected partitions unknown even if historical values exist.

The /api/v1/cluster/health response includes brokerObservation with status current, stale, or unavailable, plus the last successful observation and expiry timestamps when known. Counts in a stale observation describe the last successful check. The response also carries the rack assessment and returns notRackAwarePartitions: null unless it is complete and current. An initialized monitor returns HTTP 200 with its evidence status, including unavailable before the first successful broker discovery. An unavailable health provider returns HTTP 503. See the OpenAPI specification  for the response schema.

Follower Lag

Gauges derived from the OffsetLag field in Kafka’s DescribeLogDirs response, updated from the shared storage snapshot. Its cache lifetime is the greater of 60 seconds and four times METADATA_UPDATE_INTERVAL, with refresh scheduled before expiry. This is independent of the watermark sampling cadence.

MetricTypeLabelsDescription
pilot_follower_lag_totalGauge-Total offset lag across all follower replicas in the cluster.
pilot_follower_lag_maxGauge-Maximum offset lag of any single follower replica.
pilot_follower_lag_partitions_with_lagGauge-Partitions with non-zero follower lag.

Follower-lag API responses include collectedAt for the storage measurements and topologyObservedAt for the replica metadata. They also report measurement coverage: valid broker readings continue refreshing during a broker outage, while incomplete aggregates are null. See Measurement Coverage.

The Prometheus lag gauges update only from complete snapshots and retain their previous values during incomplete collection. A fresh scrape alone does not establish fresh lag measurements. Compare matching replicas, coverage and source observation windows before treating a difference against a JMX or exporter sample as an error.

Consumer Groups

Consumer collection runs every PILOT_CONSUMER_GROUP_COLLECTION_INTERVAL (default: 15s). Consumption rates measure committed-offset progress, not live fetch traffic. Per-group rate gauges use recent offset history followed by an exponential average with a 300-second time constant. They can remain positive while an inactive group’s commits are flat. Do not apply rate() to these already-calculated rate gauges.

Pilot uses metadata heuristics to exclude likely forward offset resets and skips over records removed by retention from consumption rates. Consuming an entire backlog after an idle period can look identical to an offset reset, so bursty consumption can be undercounted.

MetricTypeLabelsDescription
pilot_consumer_group_lagGaugeconsumer_group, topic, partitionCollected offset lag.
pilot_consumer_group_current_offsetGaugeconsumer_group, topic, partitionCommitted offset.
pilot_consumer_group_log_end_offsetGaugetopic, partitionCollected partition high watermark.
pilot_consumer_group_consume_rateGaugeconsumer_group, topic, partitionSmoothed committed-offset advancement per second.
pilot_consumer_group_lag_velocityGaugeconsumer_group, topic, partitionLag change per second; positive means growing lag.
pilot_consumer_group_lag_msGaugeconsumer_group, topic, partitionLag divided by a positive smoothed consumption rate, in milliseconds; an estimate, not record age.
pilot_consumer_group_time_to_close_secondsGaugeconsumer_group, topic, partitionEstimated time to close lag; -1 when no decreasing-lag estimate is available.
pilot_consumer_group_lag_retention_percentGaugeconsumer_group, topic, partitionLag relative to the retained offset range.
pilot_consumer_group_stateGaugeconsumer_group, stateOne for the group’s collected state.
pilot_consumer_group_membersGaugeconsumer_groupCollected member count.
pilot_consumer_group_totalGauge-Collected group count.
pilot_consumer_group_hot_partitionGaugeconsumer_group, topic, partitionOne when partition lag is more than two standard deviations above its group mean.
pilot_consumer_group_topic_partitionsGaugeconsumer_group, topicPartition count for a subscribed topic.
pilot_consumer_group_observation_started_timestamp_secondsGauge-Start of the last successful Kafka offset and watermark collection, in Unix seconds. Zero before the first success.
pilot_consumer_group_observation_completed_timestamp_secondsGauge-Completion of that successful collection, in Unix seconds. Retained on failure.
pilot_consumer_group_collection_duration_secondsHistogram-Consumer collection duration.
pilot_consumer_group_collection_errors_totalCounterphaseCollection errors; phase is list or lag.
pilot_consumer_group_collection_totalCounterresultCollection cycles; result is success or error.

The observation timestamps describe a collection interval, not an atomic snapshot or the later Prometheus scrape time. HTTP lag and consumption summaries expose the same bounds as observationStartedAt and observationCompletedAt. Failed collection retains previous values and timestamps; always check freshness.

Topic consumption estimates use the maximum group rate per partition before smoothing and summing partitions. They are not the sum of per-group gauges. For independent validation, match source offsets and time windows rather than expecting Pilot’s smoothed rates to equal another exporter’s shorter-window calculation. See Rates and Activity Scoring.

Partition Reassignments

Metrics from the reassignment executor and completion tracker.

MetricTypeLabelsDescription
pilot_reassignment_activeGauge-Partition reassignments currently submitted to Kafka.
pilot_reassignment_pendingGauge-Partition reassignments queued, waiting for broker capacity.
pilot_reassignment_completed_totalCounterreasonCumulative completed reassignments.
pilot_reassignment_failed_totalCounterreason, failure_typeRegistered but not yet reported; do not alert on it.
pilot_reassignment_duration_secondsHistogramreasonTime from submission to completion per partition. Buckets: 1s-1h.
pilot_reassignment_broker_active_movesGaugebroker_idActive moves for a specific broker.

Label values:

  • reason - currently always other
  • broker_id - Kafka broker ID (bounded by cluster size)

Self-Healing

Pilot runs three independent self-healing loops: activity (default: 30m), critical (default: 5m), and rf (default: 15m).

MetricTypeLabelsDescription
pilot_selfhealing_runs_totalCountertype, resultTotal self-healing loop executions.
pilot_selfhealing_run_duration_secondsHistogramtypeDuration of a self-healing execution. Buckets: 0.5s-5m.
pilot_selfhealing_last_success_timestamp_secondsGaugetypeUnix timestamp of the last successful healing run.
pilot_selfhealing_skipped_totalCountertype, skip_reasonSelf-healing runs that were skipped.
pilot_selfhealing_reassignments_applied_totalCountertypePartition reassignments self-healing submitted to Kafka in successful applies.
pilot_selfhealing_enabledGaugetypeWhether self-healing is enabled (1) or disabled (0).
pilot_selfhealing_interval_secondsGaugetypeConfigured interval in seconds per type.
pilot_selfhealing_proceeded_with_attributed_blockers_totalCountertypeRuns that proceeded despite hard-health blockers attributed to the offline brokers being evacuated.

Label values:

  • type - activity, critical, rf
  • result - success, error, skipped, no_action
  • skip_reason - outside_window, cooldown, partitions_in_cooldown, license_invalid, active_reassignments, reassignment_in_progress, higher_priority_running, yielded_to_higher_priority, no_proposal, no_reassignments, no_matching_reassignments, data_not_ready, benefit_unconfirmed, insufficient_benefit, not_worth_applying, stale_proposal, plan_changed, hard_health_blocker, unsafe_to_apply, dry_run

Proposal Generation

MetricTypeLabelsDescription
pilot_proposal_generation_duration_secondsHistogramtriggerDuration of a proposal generation run. Buckets: 0.5s-2m.
pilot_proposal_generation_totalCountertrigger, resultTotal proposal generation attempts.
pilot_proposal_imbalance_largest_deviationGaugemetricLargest per-broker difference from the equal share in the current placement, in the metric’s unit (count, bytes, msg/s or bytes/s).
pilot_proposal_imbalance_materialGaugemetric1 when the metric’s imbalance counts for balancing in the current placement, 0 when the metric has essentially no traffic or, with PILOT_BALANCE_FLOOR set, its largest per-broker difference is below that minimum difference.
pilot_proposal_quality_scoreGauge-Cluster imbalance variance (%) from latest proposal. Lower is better.
pilot_proposal_reassignment_countGauge-Number of partition moves in the latest proposal.
pilot_proposal_variance_currentGaugemetricRelative imbalance (%) of the current placement, as metrics.current. Includes metrics with no traffic or below the minimum difference.
pilot_proposal_variance_projectedGaugemetricRelative imbalance (%) after the latest proposal, as metrics.projected.
pilot_proposal_window_cancellations_totalCounter-Proposals cancelled due to execution time window expiry.
pilot_proposal_estimated_bytesGauge-Estimated bytes the current proposal would move.
pilot_proposal_leader_movesGauge-Leader moves in the current proposal.
pilot_proposal_replica_movesGauge-Replica moves in the current proposal.
pilot_proposal_cross_rack_movesGauge-Cross-rack moves in the current proposal.
pilot_proposal_excluded_partitionsGauge-Partitions excluded from balancing (topic regex, in-flight, cooldown).
pilot_proposal_candidates_rejected_totalCounterreasonCandidates rejected during scoring, by reason.
pilot_proposal_stage_a_duration_secondsHistogram-Time spent reaching balance. Buckets: 10ms-10s.
pilot_proposal_stage_b_duration_secondsHistogram-Time spent reducing movement cost. Buckets: 10ms-10s.
pilot_proposal_stage_b_moves_prunedGauge-Moves removed by cost pruning.
pilot_proposal_stage_b_bytes_savedGauge-Bytes saved by cost pruning.
pilot_proposal_safe_to_applyGauge-1 if the current proposal passed the safety checks, 0 if not.
pilot_proposal_confidence_scoreGauge-Current proposal confidence score (0.0-1.0), a measurement diagnostic. See Proposal Confidence.
pilot_proposal_structural_limitsGaugemetricStructural limit value per metric (0 if not at limit).
pilot_cluster_profile_total_bytesGauge-Total cluster data bytes.
pilot_cluster_profile_median_partition_bytesGauge-Median partition size in bytes.
pilot_cluster_profile_p90_partition_bytesGauge-90th percentile partition size in bytes.

Label values:

  • trigger - startup, scheduled, manual, fingerprint_change, config_change, metric_drift, post_reassignment, post_apply, proposal_expired, broker_eligibility_changed
  • result - success, error, skipped (current proposal reused), kept (current plan re-checked and kept)
  • metric - leader, follower, disk, producerMsgRate, producerByteRate, consumerMsgRate, consumerByteRate

Metadata Sampler

The metadata backend runs a collection loop every ~10s (configurable via METADATA_UPDATE_INTERVAL).

MetricTypeLabelsDescription
pilot_sampler_collection_duration_secondsHistogram-Duration of each metadata collection cycle. Buckets: 100ms-10s.
pilot_sampler_collection_errors_totalCounterphaseErrors encountered during collection.
pilot_sampler_partitions_trackedGauge-Number of partitions being monitored.
pilot_sampler_topics_trackedGauge-Number of topics being monitored.
pilot_sampler_last_successful_collection_timestampGauge-Unix timestamp of last successful collection.
pilot_sampler_cosampled_partitionsGauge-Partitions with a disk size and retained-offset pair measured in the same cycle, used for message-size estimates.
pilot_sampler_logdir_cosampling_totalCounteroutcomeCollection cycles by whether the log-dir snapshot could be paired with the cycle’s watermark read.

Label values:

  • phase - cluster_info, watermarks, logdir, collection
  • outcome - cache_miss (the sampler’s own sweep), forced (a sweep taken because the shared cache kept serving older snapshots), held (no pairable snapshot, previous pair retained)

Rolling Restart

MetricTypeLabelsDescription
pilot_rolling_restart_activeGauge-1 while a rolling restart is in progress, 0 otherwise.
pilot_rolling_restart_duration_secondsHistogram-Overall duration of a rolling restart. Buckets: 1m-2h.
pilot_rolling_restart_broker_phase_duration_secondsHistogrambroker_id, phaseDuration of each phase per broker. Buckets: 1s-10m.
pilot_rolling_restart_brokers_totalCounterresultBrokers processed during rolling restarts.
pilot_rolling_restart_totalCounterresultRolling restarts by result.

Label values:

  • phase - drain, stop, start, ready_wait, isr_catchup, restore
  • result - completed, failed, cancelled for pilot_rolling_restart_total; success, failed for pilot_rolling_restart_brokers_total

Audit Logger

Metrics for the Kafka producer used for audit event delivery to __pilot_audit_log.

MetricTypeLabelsDescription
pilot_audit_events_totalCounter-Total audit events attempted.
pilot_audit_events_success_totalCounter-Audit events successfully delivered.
pilot_audit_events_failure_totalCountererror_typeFailed first delivery attempts, which are retried, and events lost for good (dropped).
pilot_audit_delivery_duration_secondsHistogram-Latency of audit event Kafka delivery. Buckets: 1ms-5s.

Label values: error_type - delivery_error, timeout, context_cancelled, dropped. Alert on dropped; see Delivery.

State Store

Pilot persists broker state to a compacted Kafka topic (__pilot_broker_state).

MetricTypeLabelsDescription
pilot_statestore_operations_totalCounterresultState store produce operations.
pilot_statestore_delivery_duration_secondsHistogram-Produce delivery latency. Buckets: 1ms-5s.
pilot_statestore_consumed_messages_totalCounter-Messages consumed from the state topic.
pilot_statestore_brokers_storedGauge-Brokers in the in-memory state snapshot.

License

MetricTypeLabelsDescription
pilot_license_validGauge-1 if licensed, 0 if expired or invalid.
pilot_license_expiry_timestamp_secondsGauge-Unix timestamp when the license expires.
pilot_license_fetch_totalCounterresultLicense auto-fetch attempts. result: success, error.

Response Cache

In-memory cache for expensive API responses with a 15-second TTL.

MetricTypeLabelsDescription
pilot_cache_hits_totalCountercache_keyResponse cache hits.
pilot_cache_misses_totalCountercache_keyResponse cache misses.

Label values: cache_key - partitions_data, logdirs_data, logdirs_summary_data, comprehensive_cluster, etc.

Authentication

Only populated when authentication is enabled.

MetricTypeLabelsDescription
pilot_auth_logins_totalCounterprovider, resultOAuth login attempts.
pilot_auth_token_refresh_totalCounterresultAccess token refresh attempts.
pilot_auth_active_sessionsGauge-Number of active authenticated sessions.

External HTTP Clients

Outbound HTTP calls to external services.

MetricTypeLabelsDescription
pilot_external_http_duration_secondsHistogramtarget, status_codeDuration of outbound HTTP requests. Buckets: 50ms-10s.
pilot_external_http_errors_totalCountertarget, error_typeFailed outbound HTTP requests.

Label values:

  • target - license_server, oauth_provider, kafka_oauth
  • error_type - connection, timeout, other

Metadata Refresher

Background metadata refresh keeps broker rack information and topology current. The application uses METADATA_UPDATE_INTERVAL (default: 10s); the separate sampler discovery cadence does not define rack topology validity.

MetricTypeLabelsDescription
pilot_metadata_refresher_duration_secondsHistogram-Duration of metadata refresh. Buckets: 100ms-30s.
pilot_metadata_refresher_runs_totalCounterresultTotal refresh operations. result: success, error.

Reassignment Monitor

MetricTypeLabelsDescription
pilot_reassignment_monitor_check_duration_secondsHistogram-Duration of each monitor check cycle. Buckets: 100ms-10s.
pilot_reassignment_monitor_checks_totalCounter-Total check cycles executed.

Build Info

MetricTypeLabelsDescription
pilot_infoGaugeversion, go_versionAlways 1. Labels carry the Pilot version and Go runtime version.

Agent Host Metrics

Re-exported from heartbeats sent by Pilot Agents every 5s. All series are gauges, labelled with broker_id (Kafka broker ID) and node_id (the host’s hostname as reported by the agent). When an agent disconnects, every pilot_agent_* series for that broker is removed via DeletePartialMatch so a decommissioned broker does not leave stale gauges. The transition is gRPC-driven (the stream tearing down), not on a fixed sweeper, so series clear within seconds of the connection ending.

Network and disk-io values are cumulative-since-broker-boot. They are modelled as gauges, not counters, to avoid per-broker delta bookkeeping on the Pilot side. rate() works correctly; the only caveat is that a broker process restart will surface as a brief negative blip rather than an automatic counter-reset detection.

MetricTypeLabelsDescription
pilot_agent_connectedGaugebroker_id, node_id1 while the agent is connected. Series is removed on disconnect.
pilot_agent_broker_runningGaugebroker_id, node_id1 if the Kafka broker process is running on the host, 0 otherwise. Distinguishes broker-down from agent-down.
pilot_agent_uptime_secondsGaugebroker_id, node_idUptime of the agent process in seconds. Step decreases indicate agent restarts.
pilot_agent_cert_expires_at_timestamp_secondsGaugebroker_id, node_idUnix timestamp at which the agent’s mTLS certificate expires.
pilot_agent_upgrade_availableGaugebroker_id, node_id1 if a newer agent binary is available than what the agent is running, 0 otherwise.
pilot_agent_discovery_confidenceGaugebroker_id, node_idAgent discovery confidence (0-1). Low values mean Pilot guessed at install paths and lifecycle commands.
pilot_agent_infoGaugebroker_id, node_id, agent_version, kafka_version, cluster_id, kraft_role, vendorAlways 1. Inventory labels carry agent and broker metadata. Single sample per broker; old labelsets are dropped on change.
pilot_agent_cpu_percentGaugebroker_id, node_idBroker host CPU utilisation (0-100).
pilot_agent_memory_used_bytesGaugebroker_id, node_idMemory used on the broker host.
pilot_agent_memory_total_bytesGaugebroker_id, node_idTotal memory on the broker host.
pilot_agent_open_file_descriptorsGaugebroker_id, node_idOpen file descriptors held by the broker process.
pilot_agent_max_file_descriptorsGaugebroker_id, node_idMaximum file descriptors allowed for the broker process.
pilot_agent_disk_used_bytesGaugebroker_id, node_id, mount_pointDisk space used per mount point.
pilot_agent_disk_total_bytesGaugebroker_id, node_id, mount_pointDisk space total per mount point.
pilot_agent_disk_available_bytesGaugebroker_id, node_id, mount_pointDisk space available per mount point. Differs from total minus used on filesystems with reserved blocks.
pilot_agent_disk_usage_percentGaugebroker_id, node_id, mount_pointDisk usage percent per mount point (0-100).
pilot_agent_disk_inode_usedGaugebroker_id, node_id, mount_pointInodes used per mount point.
pilot_agent_disk_inode_totalGaugebroker_id, node_id, mount_pointInodes total per mount point.
pilot_agent_log_dir_used_bytesGaugebroker_id, node_id, log_dirBytes used by a Kafka log directory.
pilot_agent_log_dir_total_bytesGaugebroker_id, node_id, log_dirBytes total for the filesystem holding a Kafka log directory.
pilot_agent_log_dir_available_bytesGaugebroker_id, node_id, log_dirBytes available on the filesystem holding a Kafka log directory.
pilot_agent_log_dir_usage_percentGaugebroker_id, node_id, log_dirUsage percent for the filesystem holding a Kafka log directory (0-100).
pilot_agent_network_bytesGaugebroker_id, node_id, directionCumulative-since-boot network bytes. Use rate().
pilot_agent_network_packetsGaugebroker_id, node_id, directionCumulative-since-boot network packets. Use rate().
pilot_agent_network_errorsGaugebroker_id, node_id, directionCumulative-since-boot network errors. Use rate().
pilot_agent_network_dropsGaugebroker_id, node_id, directionCumulative-since-boot network drops. Use rate().
pilot_agent_network_link_speed_mbpsGaugebroker_id, node_idMaximum NIC link speed observed across the broker host’s network interfaces, in Mbps.
pilot_agent_disk_io_bytesGaugebroker_id, node_id, opCumulative-since-boot disk I/O bytes. Use rate().
pilot_agent_disk_io_opsGaugebroker_id, node_id, opCumulative-since-boot disk I/O operations. Use rate().
pilot_agent_disk_io_time_msGaugebroker_id, node_idCumulative-since-boot disk I/O time in milliseconds. Use rate().
pilot_agent_disk_io_weighted_time_msGaugebroker_id, node_idCumulative-since-boot weighted I/O time in milliseconds (saturation signal). Use rate().

Label values:

  • direction - tx (sent / out) or rx (received / in)
  • op - read or write
  • mount_point, log_dir - filesystem paths reported by the agent
  • node_id - hostname reported by the agent at registration

Example PromQL Queries

Error rate (5xx responses over 5 minutes)

sum(rate(pilot_http_requests_total{status_code=~"5.."}[5m])) / sum(rate(pilot_http_requests_total[5m]))

P99 request latency by route

histogram_quantile(0.99, sum(rate(pilot_http_request_duration_seconds_bucket[5m])) by (le, route) )

Kafka operation error rate

sum(rate(pilot_kafka_operation_duration_seconds_count{status="error"}[5m])) by (operation) / sum(rate(pilot_kafka_operation_duration_seconds_count[5m])) by (operation)

Under-replicated partitions alert

pilot_cluster_under_replicated_partitions > 0

Reassignment stuck (active > 0 for over 2 hours)

pilot_reassignment_active > 0

Sampler stale (no success in 60s)

time() - pilot_sampler_last_successful_collection_timestamp > 60

Self-healing not running

Critical healing hasn’t succeeded in 30m while URPs exist:

(time() - pilot_selfhealing_last_success_timestamp_seconds{type="critical"} > 1800) and pilot_cluster_under_replicated_partitions > 0

License expiring within 7 days

pilot_license_expiry_timestamp_seconds - time() < 7 * 24 * 3600

Audit events lost

increase(pilot_audit_events_failure_total{error_type="dropped"}[15m]) > 0

Audit delivery retry rate

sum(rate(pilot_audit_events_failure_total{error_type!="dropped"}[5m])) / sum(rate(pilot_audit_events_total[5m]))

Cache hit ratio

sum(rate(pilot_cache_hits_total[5m])) / (sum(rate(pilot_cache_hits_total[5m])) + sum(rate(pilot_cache_misses_total[5m])))

Follower lag warning (total lag exceeds threshold)

pilot_follower_lag_total > 1000

Single replica falling far behind

pilot_follower_lag_max > 10000

Agent disconnected (no series for a known broker)

absent(pilot_agent_connected{broker_id="1"})

Broker host log directory above 85% full

pilot_agent_log_dir_usage_percent > 85

Broker host network throughput per direction

rate(pilot_agent_network_bytes[5m])

Agent online but broker process is down

pilot_agent_connected == 1 and pilot_agent_broker_running == 0

Agent mTLS certificate expiring within 7 days

pilot_agent_cert_expires_at_timestamp_seconds - time() < 7 * 24 * 3600

Inventory by Kafka version

count by (kafka_version) (pilot_agent_info)

Agent restarted (uptime stepped down) within the last 5 minutes

delta(pilot_agent_uptime_seconds[5m]) < 0

Low-confidence discovery (Pilot guessed at install paths)

pilot_agent_discovery_confidence < 0.7
Last updated on