Prometheus Metrics
Pilot exposes Prometheus metrics at the /metrics endpoint. All application metrics use the pilot_ namespace prefix.
Standard go_* runtime metrics (goroutines, GC, memory) and process_* metrics (CPU, RSS, open file descriptors) are included automatically.
HTTP Layer
RED (Rate, Errors, Duration) metrics captured by middleware for every API request. The route label uses the route template (e.g. /api/v1/topics/{topic}/config) to keep cardinality bounded.
| Metric | Type | Labels | Description |
|---|---|---|---|
pilot_http_request_duration_seconds | Histogram | method, route, status_code | Duration of HTTP requests. Buckets: 1ms-10s. |
pilot_http_requests_total | Counter | method, route, status_code | Total number of HTTP requests. |
pilot_http_request_size_bytes | Histogram | method, route | Size of HTTP request bodies. Buckets: 100B-1MB. |
pilot_http_response_size_bytes | Histogram | method, route | Size of HTTP response bodies. Buckets: 1KB-10MB. |
pilot_http_inflight_requests | Gauge | - | Number of currently in-flight HTTP requests. |
Kafka Operations
Latency and error tracking for Kafka admin/client calls. Instrumented operations include get_cluster_info, get_watermarks, describe_log_dirs, describe_config, update_topic_config, get_broker_config, alter_partition_reassignments, preferred_leader_election, describe_client_quotas, alter_client_quotas, and alter_replica_log_dirs.
| Metric | Type | Labels | Description |
|---|---|---|---|
pilot_kafka_operation_duration_seconds | Histogram | operation, status | Duration of a Kafka operation. status is success or error. Buckets: 10ms-60s. |
pilot_kafka_operation_errors_total | Counter | operation, error_type | Total Kafka operation errors. error_type: timeout, auth, connection, kafka_error, other. |
Cluster Health
The background health monitor updates these gauges on each check cycle (default 30s). The /api/v1/ready readiness probe requires a current broker observation with at least one available broker. A failed or expired observation makes readiness unavailable without declaring that every broker is offline. Rack assessment status does not change readiness.
| Metric | Type | Labels | Description |
|---|---|---|---|
pilot_cluster_brokers_total | Gauge | - | Total known brokers in the cluster. |
pilot_cluster_brokers_available | Gauge | - | Brokers present in the last successful discovery response. |
pilot_cluster_brokers_unavailable | Gauge | - | Known brokers absent from the last successful discovery response. |
pilot_cluster_broker_observation_current | Gauge | - | Whether the latest broker discovery attempt succeeded: 1 or 0. Also check the expiry timestamp. |
pilot_cluster_broker_observation_timestamp_seconds | Gauge | - | Unix timestamp of the last successful broker observation, or 0 before the first success. |
pilot_cluster_broker_observation_expires_timestamp_seconds | Gauge | - | Unix timestamp when that observation expires, or 0 before the first success. |
pilot_cluster_brokers_maintenance | Gauge | - | Brokers currently in maintenance mode. |
pilot_cluster_under_replicated_partitions | Gauge | - | Partitions where len(ISR) < len(Replicas). |
pilot_cluster_offline_partitions | Gauge | - | Partitions with no leader or empty ISR. |
pilot_cluster_urps_from_offline_brokers | Gauge | - | Under-replicated partitions caused by offline brokers. |
pilot_cluster_urps_from_follower_lag | Gauge | - | Under-replicated partitions caused by follower lag (all replicas online but not in ISR). |
pilot_cluster_not_rack_aware_partitions | Gauge | - | Exact count of assigned partitions that cannot retain minimum ISR after any single rack loss. Emitted only for a complete, current assessment. |
pilot_cluster_rack_assessment_status | Gauge | status | One for the active status, zero for the others: complete, partial, unavailable. |
pilot_cluster_rack_assessment_partitions | Gauge | coverage | Partition counts for total, evaluated, and unknown. All three series are omitted when the total scope is unknown. |
pilot_cluster_rack_assessment_known_violations | Gauge | - | Proven violations among evaluated partitions. A lower bound when partial; omitted when unavailable. |
pilot_cluster_rack_assessment_unknown_partitions | Gauge | reason | Positive counts of unknown partition evidence by bounded reason code. Use coverage series for the assessment denominator. |
pilot_cluster_rack_assessment_timestamp_seconds | Gauge | source | Unix seconds for assessment, topology, config_oldest, config_newest, and valid_until. Unknown timestamps are omitted. |
pilot_cluster_rack_assessment_config_refresh_failures | Gauge | - | Topic configuration reads that failed or lacked a valid effective minimum ISR in the completed assessment. Omitted before any assessment. |
pilot_cluster_health_check_duration_seconds | Histogram | - | Duration of each health check cycle. Buckets: 100ms-10s. |
Broker discovery is independent of partition leader-election errors. A failed discovery retains the last successful broker counts and timestamps. Interpret these counts as current only when pilot_cluster_broker_observation_current == 1 and pilot_cluster_broker_observation_expires_timestamp_seconds > time(). Alert separately when observations fail or expire. Discovery reports Kafka’s broker roster, not a direct connectivity test to every broker.
Rack health checks assigned replicas against effective minimum ISR after losing the rack containing the most replicas. It does not establish current ISR availability. Missing configuration, rack labels, assignments or usable topology remain unknown. A partial assessment can contain proven violations alongside unknown partitions.
An absent exact-count series means the exact count is unavailable. Do not replace it with zero in queries or alerts. Check pilot_cluster_rack_assessment_status{status="complete"} before interpreting zero as a complete assessment without known violations. Expiry removes the exact series at scrape time even if no new health pass has completed.
Source timestamps retain the original observations. valid_until is a deadline, not an observation time. Configuration validity follows the cache refresh TTL plus one health interval and the five-second rack stage budget (65s by default). Topology validity follows the configured metadata refresh interval plus request and retry-backoff budget (65s with the application’s default 10s metadata interval). A failed configuration refresh leaves affected partitions unknown even if historical values exist.
The /api/v1/cluster/health response includes brokerObservation with status current, stale, or unavailable, plus the last successful observation and expiry timestamps when known. Counts in a stale observation describe the last successful check. The response also carries the rack assessment and returns notRackAwarePartitions: null unless it is complete and current. An initialized monitor returns HTTP 200 with its evidence status, including unavailable before the first successful broker discovery. An unavailable health provider returns HTTP 503. See the OpenAPI specification for the response schema.
Follower Lag
Gauges derived from the OffsetLag field in Kafka’s DescribeLogDirs response, updated from the shared storage snapshot. Its cache lifetime is the greater of 60 seconds and four times METADATA_UPDATE_INTERVAL, with refresh scheduled before expiry. This is independent of the watermark sampling cadence.
| Metric | Type | Labels | Description |
|---|---|---|---|
pilot_follower_lag_total | Gauge | - | Total offset lag across all follower replicas in the cluster. |
pilot_follower_lag_max | Gauge | - | Maximum offset lag of any single follower replica. |
pilot_follower_lag_partitions_with_lag | Gauge | - | Partitions with non-zero follower lag. |
Follower-lag API responses include collectedAt for the storage measurements and topologyObservedAt for the replica metadata. They also report measurement coverage: valid broker readings continue refreshing during a broker outage, while incomplete aggregates are null. See Measurement Coverage.
The Prometheus lag gauges update only from complete snapshots and retain their previous values during incomplete collection. A fresh scrape alone does not establish fresh lag measurements. Compare matching replicas, coverage and source observation windows before treating a difference against a JMX or exporter sample as an error.
Consumer Groups
Consumer collection runs every PILOT_CONSUMER_GROUP_COLLECTION_INTERVAL (default: 15s). Consumption rates measure committed-offset progress, not live fetch traffic. Per-group rate gauges use recent offset history followed by an exponential average with a 300-second time constant. They can remain positive while an inactive group’s commits are flat. Do not apply rate() to these already-calculated rate gauges.
Pilot uses metadata heuristics to exclude likely forward offset resets and skips over records removed by retention from consumption rates. Consuming an entire backlog after an idle period can look identical to an offset reset, so bursty consumption can be undercounted.
| Metric | Type | Labels | Description |
|---|---|---|---|
pilot_consumer_group_lag | Gauge | consumer_group, topic, partition | Collected offset lag. |
pilot_consumer_group_current_offset | Gauge | consumer_group, topic, partition | Committed offset. |
pilot_consumer_group_log_end_offset | Gauge | topic, partition | Collected partition high watermark. |
pilot_consumer_group_consume_rate | Gauge | consumer_group, topic, partition | Smoothed committed-offset advancement per second. |
pilot_consumer_group_lag_velocity | Gauge | consumer_group, topic, partition | Lag change per second; positive means growing lag. |
pilot_consumer_group_lag_ms | Gauge | consumer_group, topic, partition | Lag divided by a positive smoothed consumption rate, in milliseconds; an estimate, not record age. |
pilot_consumer_group_time_to_close_seconds | Gauge | consumer_group, topic, partition | Estimated time to close lag; -1 when no decreasing-lag estimate is available. |
pilot_consumer_group_lag_retention_percent | Gauge | consumer_group, topic, partition | Lag relative to the retained offset range. |
pilot_consumer_group_state | Gauge | consumer_group, state | One for the group’s collected state. |
pilot_consumer_group_members | Gauge | consumer_group | Collected member count. |
pilot_consumer_group_total | Gauge | - | Collected group count. |
pilot_consumer_group_hot_partition | Gauge | consumer_group, topic, partition | One when partition lag is more than two standard deviations above its group mean. |
pilot_consumer_group_topic_partitions | Gauge | consumer_group, topic | Partition count for a subscribed topic. |
pilot_consumer_group_observation_started_timestamp_seconds | Gauge | - | Start of the last successful Kafka offset and watermark collection, in Unix seconds. Zero before the first success. |
pilot_consumer_group_observation_completed_timestamp_seconds | Gauge | - | Completion of that successful collection, in Unix seconds. Retained on failure. |
pilot_consumer_group_collection_duration_seconds | Histogram | - | Consumer collection duration. |
pilot_consumer_group_collection_errors_total | Counter | phase | Collection errors; phase is list or lag. |
pilot_consumer_group_collection_total | Counter | result | Collection cycles; result is success or error. |
The observation timestamps describe a collection interval, not an atomic snapshot or the later Prometheus scrape time. HTTP lag and consumption summaries expose the same bounds as observationStartedAt and observationCompletedAt. Failed collection retains previous values and timestamps; always check freshness.
Topic consumption estimates use the maximum group rate per partition before smoothing and summing partitions. They are not the sum of per-group gauges. For independent validation, match source offsets and time windows rather than expecting Pilot’s smoothed rates to equal another exporter’s shorter-window calculation. See Rates and Activity Scoring.
Partition Reassignments
Metrics from the reassignment executor and completion tracker.
| Metric | Type | Labels | Description |
|---|---|---|---|
pilot_reassignment_active | Gauge | - | Partition reassignments currently submitted to Kafka. |
pilot_reassignment_pending | Gauge | - | Partition reassignments queued, waiting for broker capacity. |
pilot_reassignment_completed_total | Counter | reason | Cumulative completed reassignments. |
pilot_reassignment_failed_total | Counter | reason, failure_type | Registered but not yet reported; do not alert on it. |
pilot_reassignment_duration_seconds | Histogram | reason | Time from submission to completion per partition. Buckets: 1s-1h. |
pilot_reassignment_broker_active_moves | Gauge | broker_id | Active moves for a specific broker. |
Label values:
reason- currently alwaysotherbroker_id- Kafka broker ID (bounded by cluster size)
Self-Healing
Pilot runs three independent self-healing loops: activity (default: 30m), critical (default: 5m), and rf (default: 15m).
| Metric | Type | Labels | Description |
|---|---|---|---|
pilot_selfhealing_runs_total | Counter | type, result | Total self-healing loop executions. |
pilot_selfhealing_run_duration_seconds | Histogram | type | Duration of a self-healing execution. Buckets: 0.5s-5m. |
pilot_selfhealing_last_success_timestamp_seconds | Gauge | type | Unix timestamp of the last successful healing run. |
pilot_selfhealing_skipped_total | Counter | type, skip_reason | Self-healing runs that were skipped. |
pilot_selfhealing_reassignments_applied_total | Counter | type | Partition reassignments self-healing submitted to Kafka in successful applies. |
pilot_selfhealing_enabled | Gauge | type | Whether self-healing is enabled (1) or disabled (0). |
pilot_selfhealing_interval_seconds | Gauge | type | Configured interval in seconds per type. |
pilot_selfhealing_proceeded_with_attributed_blockers_total | Counter | type | Runs that proceeded despite hard-health blockers attributed to the offline brokers being evacuated. |
Label values:
type-activity,critical,rfresult-success,error,skipped,no_actionskip_reason-outside_window,cooldown,partitions_in_cooldown,license_invalid,active_reassignments,reassignment_in_progress,higher_priority_running,yielded_to_higher_priority,no_proposal,no_reassignments,no_matching_reassignments,data_not_ready,benefit_unconfirmed,insufficient_benefit,not_worth_applying,stale_proposal,plan_changed,hard_health_blocker,unsafe_to_apply,dry_run
Proposal Generation
| Metric | Type | Labels | Description |
|---|---|---|---|
pilot_proposal_generation_duration_seconds | Histogram | trigger | Duration of a proposal generation run. Buckets: 0.5s-2m. |
pilot_proposal_generation_total | Counter | trigger, result | Total proposal generation attempts. |
pilot_proposal_imbalance_largest_deviation | Gauge | metric | Largest per-broker difference from the equal share in the current placement, in the metric’s unit (count, bytes, msg/s or bytes/s). |
pilot_proposal_imbalance_material | Gauge | metric | 1 when the metric’s imbalance counts for balancing in the current placement, 0 when the metric has essentially no traffic or, with PILOT_BALANCE_FLOOR set, its largest per-broker difference is below that minimum difference. |
pilot_proposal_quality_score | Gauge | - | Cluster imbalance variance (%) from latest proposal. Lower is better. |
pilot_proposal_reassignment_count | Gauge | - | Number of partition moves in the latest proposal. |
pilot_proposal_variance_current | Gauge | metric | Relative imbalance (%) of the current placement, as metrics.current. Includes metrics with no traffic or below the minimum difference. |
pilot_proposal_variance_projected | Gauge | metric | Relative imbalance (%) after the latest proposal, as metrics.projected. |
pilot_proposal_window_cancellations_total | Counter | - | Proposals cancelled due to execution time window expiry. |
pilot_proposal_estimated_bytes | Gauge | - | Estimated bytes the current proposal would move. |
pilot_proposal_leader_moves | Gauge | - | Leader moves in the current proposal. |
pilot_proposal_replica_moves | Gauge | - | Replica moves in the current proposal. |
pilot_proposal_cross_rack_moves | Gauge | - | Cross-rack moves in the current proposal. |
pilot_proposal_excluded_partitions | Gauge | - | Partitions excluded from balancing (topic regex, in-flight, cooldown). |
pilot_proposal_candidates_rejected_total | Counter | reason | Candidates rejected during scoring, by reason. |
pilot_proposal_stage_a_duration_seconds | Histogram | - | Time spent reaching balance. Buckets: 10ms-10s. |
pilot_proposal_stage_b_duration_seconds | Histogram | - | Time spent reducing movement cost. Buckets: 10ms-10s. |
pilot_proposal_stage_b_moves_pruned | Gauge | - | Moves removed by cost pruning. |
pilot_proposal_stage_b_bytes_saved | Gauge | - | Bytes saved by cost pruning. |
pilot_proposal_safe_to_apply | Gauge | - | 1 if the current proposal passed the safety checks, 0 if not. |
pilot_proposal_confidence_score | Gauge | - | Current proposal confidence score (0.0-1.0), a measurement diagnostic. See Proposal Confidence. |
pilot_proposal_structural_limits | Gauge | metric | Structural limit value per metric (0 if not at limit). |
pilot_cluster_profile_total_bytes | Gauge | - | Total cluster data bytes. |
pilot_cluster_profile_median_partition_bytes | Gauge | - | Median partition size in bytes. |
pilot_cluster_profile_p90_partition_bytes | Gauge | - | 90th percentile partition size in bytes. |
Label values:
trigger-startup,scheduled,manual,fingerprint_change,config_change,metric_drift,post_reassignment,post_apply,proposal_expired,broker_eligibility_changedresult-success,error,skipped(current proposal reused),kept(current plan re-checked and kept)metric-leader,follower,disk,producerMsgRate,producerByteRate,consumerMsgRate,consumerByteRate
Metadata Sampler
The metadata backend runs a collection loop every ~10s (configurable via METADATA_UPDATE_INTERVAL).
| Metric | Type | Labels | Description |
|---|---|---|---|
pilot_sampler_collection_duration_seconds | Histogram | - | Duration of each metadata collection cycle. Buckets: 100ms-10s. |
pilot_sampler_collection_errors_total | Counter | phase | Errors encountered during collection. |
pilot_sampler_partitions_tracked | Gauge | - | Number of partitions being monitored. |
pilot_sampler_topics_tracked | Gauge | - | Number of topics being monitored. |
pilot_sampler_last_successful_collection_timestamp | Gauge | - | Unix timestamp of last successful collection. |
pilot_sampler_cosampled_partitions | Gauge | - | Partitions with a disk size and retained-offset pair measured in the same cycle, used for message-size estimates. |
pilot_sampler_logdir_cosampling_total | Counter | outcome | Collection cycles by whether the log-dir snapshot could be paired with the cycle’s watermark read. |
Label values:
phase-cluster_info,watermarks,logdir,collectionoutcome-cache_miss(the sampler’s own sweep),forced(a sweep taken because the shared cache kept serving older snapshots),held(no pairable snapshot, previous pair retained)
Rolling Restart
| Metric | Type | Labels | Description |
|---|---|---|---|
pilot_rolling_restart_active | Gauge | - | 1 while a rolling restart is in progress, 0 otherwise. |
pilot_rolling_restart_duration_seconds | Histogram | - | Overall duration of a rolling restart. Buckets: 1m-2h. |
pilot_rolling_restart_broker_phase_duration_seconds | Histogram | broker_id, phase | Duration of each phase per broker. Buckets: 1s-10m. |
pilot_rolling_restart_brokers_total | Counter | result | Brokers processed during rolling restarts. |
pilot_rolling_restart_total | Counter | result | Rolling restarts by result. |
Label values:
phase-drain,stop,start,ready_wait,isr_catchup,restoreresult-completed,failed,cancelledforpilot_rolling_restart_total;success,failedforpilot_rolling_restart_brokers_total
Audit Logger
Metrics for the Kafka producer used for audit event delivery to __pilot_audit_log.
| Metric | Type | Labels | Description |
|---|---|---|---|
pilot_audit_events_total | Counter | - | Total audit events attempted. |
pilot_audit_events_success_total | Counter | - | Audit events successfully delivered. |
pilot_audit_events_failure_total | Counter | error_type | Failed first delivery attempts, which are retried, and events lost for good (dropped). |
pilot_audit_delivery_duration_seconds | Histogram | - | Latency of audit event Kafka delivery. Buckets: 1ms-5s. |
Label values: error_type - delivery_error, timeout, context_cancelled, dropped. Alert on dropped; see Delivery.
State Store
Pilot persists broker state to a compacted Kafka topic (__pilot_broker_state).
| Metric | Type | Labels | Description |
|---|---|---|---|
pilot_statestore_operations_total | Counter | result | State store produce operations. |
pilot_statestore_delivery_duration_seconds | Histogram | - | Produce delivery latency. Buckets: 1ms-5s. |
pilot_statestore_consumed_messages_total | Counter | - | Messages consumed from the state topic. |
pilot_statestore_brokers_stored | Gauge | - | Brokers in the in-memory state snapshot. |
License
| Metric | Type | Labels | Description |
|---|---|---|---|
pilot_license_valid | Gauge | - | 1 if licensed, 0 if expired or invalid. |
pilot_license_expiry_timestamp_seconds | Gauge | - | Unix timestamp when the license expires. |
pilot_license_fetch_total | Counter | result | License auto-fetch attempts. result: success, error. |
Response Cache
In-memory cache for expensive API responses with a 15-second TTL.
| Metric | Type | Labels | Description |
|---|---|---|---|
pilot_cache_hits_total | Counter | cache_key | Response cache hits. |
pilot_cache_misses_total | Counter | cache_key | Response cache misses. |
Label values: cache_key - partitions_data, logdirs_data, logdirs_summary_data, comprehensive_cluster, etc.
Authentication
Only populated when authentication is enabled.
| Metric | Type | Labels | Description |
|---|---|---|---|
pilot_auth_logins_total | Counter | provider, result | OAuth login attempts. |
pilot_auth_token_refresh_total | Counter | result | Access token refresh attempts. |
pilot_auth_active_sessions | Gauge | - | Number of active authenticated sessions. |
External HTTP Clients
Outbound HTTP calls to external services.
| Metric | Type | Labels | Description |
|---|---|---|---|
pilot_external_http_duration_seconds | Histogram | target, status_code | Duration of outbound HTTP requests. Buckets: 50ms-10s. |
pilot_external_http_errors_total | Counter | target, error_type | Failed outbound HTTP requests. |
Label values:
target-license_server,oauth_provider,kafka_oautherror_type-connection,timeout,other
Metadata Refresher
Background metadata refresh keeps broker rack information and topology current. The application uses METADATA_UPDATE_INTERVAL (default: 10s); the separate sampler discovery cadence does not define rack topology validity.
| Metric | Type | Labels | Description |
|---|---|---|---|
pilot_metadata_refresher_duration_seconds | Histogram | - | Duration of metadata refresh. Buckets: 100ms-30s. |
pilot_metadata_refresher_runs_total | Counter | result | Total refresh operations. result: success, error. |
Reassignment Monitor
| Metric | Type | Labels | Description |
|---|---|---|---|
pilot_reassignment_monitor_check_duration_seconds | Histogram | - | Duration of each monitor check cycle. Buckets: 100ms-10s. |
pilot_reassignment_monitor_checks_total | Counter | - | Total check cycles executed. |
Build Info
| Metric | Type | Labels | Description |
|---|---|---|---|
pilot_info | Gauge | version, go_version | Always 1. Labels carry the Pilot version and Go runtime version. |
Agent Host Metrics
Re-exported from heartbeats sent by Pilot Agents every 5s. All series are gauges, labelled with broker_id (Kafka broker ID) and node_id (the host’s hostname as reported by the agent). When an agent disconnects, every pilot_agent_* series for that broker is removed via DeletePartialMatch so a decommissioned broker does not leave stale gauges. The transition is gRPC-driven (the stream tearing down), not on a fixed sweeper, so series clear within seconds of the connection ending.
Network and disk-io values are cumulative-since-broker-boot. They are modelled as gauges, not counters, to avoid per-broker delta bookkeeping on the Pilot side. rate() works correctly; the only caveat is that a broker process restart will surface as a brief negative blip rather than an automatic counter-reset detection.
| Metric | Type | Labels | Description |
|---|---|---|---|
pilot_agent_connected | Gauge | broker_id, node_id | 1 while the agent is connected. Series is removed on disconnect. |
pilot_agent_broker_running | Gauge | broker_id, node_id | 1 if the Kafka broker process is running on the host, 0 otherwise. Distinguishes broker-down from agent-down. |
pilot_agent_uptime_seconds | Gauge | broker_id, node_id | Uptime of the agent process in seconds. Step decreases indicate agent restarts. |
pilot_agent_cert_expires_at_timestamp_seconds | Gauge | broker_id, node_id | Unix timestamp at which the agent’s mTLS certificate expires. |
pilot_agent_upgrade_available | Gauge | broker_id, node_id | 1 if a newer agent binary is available than what the agent is running, 0 otherwise. |
pilot_agent_discovery_confidence | Gauge | broker_id, node_id | Agent discovery confidence (0-1). Low values mean Pilot guessed at install paths and lifecycle commands. |
pilot_agent_info | Gauge | broker_id, node_id, agent_version, kafka_version, cluster_id, kraft_role, vendor | Always 1. Inventory labels carry agent and broker metadata. Single sample per broker; old labelsets are dropped on change. |
pilot_agent_cpu_percent | Gauge | broker_id, node_id | Broker host CPU utilisation (0-100). |
pilot_agent_memory_used_bytes | Gauge | broker_id, node_id | Memory used on the broker host. |
pilot_agent_memory_total_bytes | Gauge | broker_id, node_id | Total memory on the broker host. |
pilot_agent_open_file_descriptors | Gauge | broker_id, node_id | Open file descriptors held by the broker process. |
pilot_agent_max_file_descriptors | Gauge | broker_id, node_id | Maximum file descriptors allowed for the broker process. |
pilot_agent_disk_used_bytes | Gauge | broker_id, node_id, mount_point | Disk space used per mount point. |
pilot_agent_disk_total_bytes | Gauge | broker_id, node_id, mount_point | Disk space total per mount point. |
pilot_agent_disk_available_bytes | Gauge | broker_id, node_id, mount_point | Disk space available per mount point. Differs from total minus used on filesystems with reserved blocks. |
pilot_agent_disk_usage_percent | Gauge | broker_id, node_id, mount_point | Disk usage percent per mount point (0-100). |
pilot_agent_disk_inode_used | Gauge | broker_id, node_id, mount_point | Inodes used per mount point. |
pilot_agent_disk_inode_total | Gauge | broker_id, node_id, mount_point | Inodes total per mount point. |
pilot_agent_log_dir_used_bytes | Gauge | broker_id, node_id, log_dir | Bytes used by a Kafka log directory. |
pilot_agent_log_dir_total_bytes | Gauge | broker_id, node_id, log_dir | Bytes total for the filesystem holding a Kafka log directory. |
pilot_agent_log_dir_available_bytes | Gauge | broker_id, node_id, log_dir | Bytes available on the filesystem holding a Kafka log directory. |
pilot_agent_log_dir_usage_percent | Gauge | broker_id, node_id, log_dir | Usage percent for the filesystem holding a Kafka log directory (0-100). |
pilot_agent_network_bytes | Gauge | broker_id, node_id, direction | Cumulative-since-boot network bytes. Use rate(). |
pilot_agent_network_packets | Gauge | broker_id, node_id, direction | Cumulative-since-boot network packets. Use rate(). |
pilot_agent_network_errors | Gauge | broker_id, node_id, direction | Cumulative-since-boot network errors. Use rate(). |
pilot_agent_network_drops | Gauge | broker_id, node_id, direction | Cumulative-since-boot network drops. Use rate(). |
pilot_agent_network_link_speed_mbps | Gauge | broker_id, node_id | Maximum NIC link speed observed across the broker host’s network interfaces, in Mbps. |
pilot_agent_disk_io_bytes | Gauge | broker_id, node_id, op | Cumulative-since-boot disk I/O bytes. Use rate(). |
pilot_agent_disk_io_ops | Gauge | broker_id, node_id, op | Cumulative-since-boot disk I/O operations. Use rate(). |
pilot_agent_disk_io_time_ms | Gauge | broker_id, node_id | Cumulative-since-boot disk I/O time in milliseconds. Use rate(). |
pilot_agent_disk_io_weighted_time_ms | Gauge | broker_id, node_id | Cumulative-since-boot weighted I/O time in milliseconds (saturation signal). Use rate(). |
Label values:
direction-tx(sent / out) orrx(received / in)op-readorwritemount_point,log_dir- filesystem paths reported by the agentnode_id- hostname reported by the agent at registration
Example PromQL Queries
Error rate (5xx responses over 5 minutes)
sum(rate(pilot_http_requests_total{status_code=~"5.."}[5m]))
/ sum(rate(pilot_http_requests_total[5m]))P99 request latency by route
histogram_quantile(0.99,
sum(rate(pilot_http_request_duration_seconds_bucket[5m])) by (le, route)
)Kafka operation error rate
sum(rate(pilot_kafka_operation_duration_seconds_count{status="error"}[5m])) by (operation)
/ sum(rate(pilot_kafka_operation_duration_seconds_count[5m])) by (operation)Under-replicated partitions alert
pilot_cluster_under_replicated_partitions > 0Reassignment stuck (active > 0 for over 2 hours)
pilot_reassignment_active > 0Sampler stale (no success in 60s)
time() - pilot_sampler_last_successful_collection_timestamp > 60Self-healing not running
Critical healing hasn’t succeeded in 30m while URPs exist:
(time() - pilot_selfhealing_last_success_timestamp_seconds{type="critical"} > 1800)
and pilot_cluster_under_replicated_partitions > 0License expiring within 7 days
pilot_license_expiry_timestamp_seconds - time() < 7 * 24 * 3600Audit events lost
increase(pilot_audit_events_failure_total{error_type="dropped"}[15m]) > 0Audit delivery retry rate
sum(rate(pilot_audit_events_failure_total{error_type!="dropped"}[5m]))
/ sum(rate(pilot_audit_events_total[5m]))Cache hit ratio
sum(rate(pilot_cache_hits_total[5m]))
/ (sum(rate(pilot_cache_hits_total[5m])) + sum(rate(pilot_cache_misses_total[5m])))Follower lag warning (total lag exceeds threshold)
pilot_follower_lag_total > 1000Single replica falling far behind
pilot_follower_lag_max > 10000Agent disconnected (no series for a known broker)
absent(pilot_agent_connected{broker_id="1"})Broker host log directory above 85% full
pilot_agent_log_dir_usage_percent > 85Broker host network throughput per direction
rate(pilot_agent_network_bytes[5m])Agent online but broker process is down
pilot_agent_connected == 1 and pilot_agent_broker_running == 0Agent mTLS certificate expiring within 7 days
pilot_agent_cert_expires_at_timestamp_seconds - time() < 7 * 24 * 3600Inventory by Kafka version
count by (kafka_version) (pilot_agent_info)Agent restarted (uptime stepped down) within the last 5 minutes
delta(pilot_agent_uptime_seconds[5m]) < 0Low-confidence discovery (Pilot guessed at install paths)
pilot_agent_discovery_confidence < 0.7