Monitoring & Dashboard
Pilot provides real-time visibility into your Kafka cluster through a web dashboard and REST API. All monitoring features are free and do not require a license.
Dashboard Overview
Access the dashboard at http://localhost:8080 (or your configured PORT). The overview page displays:
- Cluster health - broker availability, under-replicated partitions, offline partitions
- Broker status - online/offline/maintenance state per broker with rack assignments
- Partition distribution - leader and replica counts per broker
- Activity scores - per-partition throughput estimates based on metadata sampling
Rates and Activity Scoring
Pilot estimates partition activity from broker metadata, without consuming any messages. Every METADATA_UPDATE_INTERVAL (default: 10s), it:
- Fetches partition watermarks - high/low offsets for all partitions
- Reads the shared storage snapshot - disk usage per partition per broker, refreshed separately from watermarks
- Computes activity rates - message rate from watermark deltas, byte rate from log size changes
The Topics page separates producer and consumer estimates:
| Measurement | Meaning |
|---|---|
| Producer message rate | High-watermark advancement per second, smoothed across observations. |
| Producer byte rate | Estimated log growth per second, using distinct disk observations. This is not broker network ingress. |
| Consumer message rate | Committed-offset advancement. The topic estimate takes the maximum group rate per partition, smooths it, then sums partitions. It does not sum every group’s traffic. |
| Consumer byte rate | Committed-offset progress multiplied by an estimated record size. This is not broker network egress. |
| Stored replicas | Storage across replicas, including replication copies, rather than one logical copy of the topic. |
Producer rates use a fast exponential average. Consumer rates use recent committed-offset history followed by an exponential average with a 300-second time constant. This is not a finite five-minute window: once the raw input is zero, about 37% of the old consumer estimate remains after five minutes and 5% after fifteen minutes. Stopping producers or consumers therefore need not immediately show zero. Consumer rates also depend on how frequently clients commit offsets, not just when they fetch or process records.
Proposal scoring uses separate averaged rate history. See Measurement Readiness and Sustained Benefit for the checks that determine when a rebalance can be applied.
Storage figures in the dashboard, the Brokers and Partitions views and /api/v1/cluster/logdirs are measured sizes. For topics whose cleanup.policy contains compact, proposals and what-if simulations weigh each partition at its typical size over its log-cleaner cycle instead, so their broker disk totals can differ from these views. See Compacted Topics.
Sample Availability
Sampled producer rates distinguish current observations from Partial sample, Last known and Unavailable. A missing observation is not measured zero. Failed reads retain the last valid values and their original source timestamps rather than presenting an unsuccessful attempt as a new measurement.
Partition API responses expose samplingStatus, rateObservedAt and diskMeasuredAt. The rate timestamp describes the message-rate observation; the disk timestamp describes storage. hasSamples means historical samples exist, not that the latest attempt succeeded. Inspect coverage and source ages when using aggregate rates.
Current limitation: byte rates can briefly show zero after startup before two distinct disk observations establish a rate. Message-rate availability does not establish byte-rate readiness. Do not interpret this initial zero as proof that no bytes are being produced.
Comparing with Prometheus
Match topic and partition identities, definitions, observation times and calculation windows. Pilot’s own Prometheus metrics expose its collected estimates and are not an independent reference. Broker JMX and Kafka exporter metrics have their own collection windows; a broker byte counter is not equivalent to sampled log growth. Retention can advance log starts without consumer activity.
Consumer observation timestamps identify the successful Kafka read interval. Follower-lag responses carry collectedAt for the storage measurements and topologyObservedAt for replica metadata; compare against those rather than the scrape time. See Consumer Group Metrics.
Broker Health Monitoring
A background health monitor checks broker availability every 30 seconds:
- Available - responding to metadata requests
- Unavailable - not responding (counted as
total - available) - Maintenance - explicitly placed in Maintenance Mode
The readiness probe at /api/v1/ready requires a health observation and at least one available Kafka broker. Rack assessment uncertainty does not change readiness.
Rack Topology
Pilot checks whether each partition’s assigned replicas can retain its effective min.insync.replicas after any single rack is lost. For example, four replicas distributed 2/1/1 across three racks can retain minimum ISR 2 after the largest rack is lost. Co-location alone does not establish a violation. This is an assignment assessment, not a guarantee that the current ISR can keep accepting writes.
The rackAwareness assessment reports complete, partial or unavailable, with evaluated and unknown partition counts, proven violations, source timestamps and an evidence deadline. Missing minimum-ISR configuration, rack labels, assignments or usable topology remain unknown. A partial assessment can have both proven violations and unknown partitions.
Expired evidence becomes unavailable even if the next health pass has not completed. The legacy HTTP notRackAwarePartitions count is null unless the assessment is complete and current; the corresponding exact Prometheus series is omitted. Never replace either with zero. Zero establishes no assigned-partition rack violations only when the assessment is complete and current. See Cluster Health Metrics.
API Endpoints
| Method | Path | Description |
|---|---|---|
GET | /api/v1/cluster | Basic cluster info |
GET | /api/v1/cluster/comprehensive | Full cluster data (use ?summary=true to omit partition arrays) |
GET | /api/v1/cluster/health | Broker and partition health |
GET | /api/v1/cluster/logdirs | Disk usage (use ?summary=true for broker-level only) |
GET | /api/v1/cluster/follower-lag | Follower lag (use ?summary=true, ?topic=, ?broker=) |
GET | /api/v1/partitions | Partition listing (supports ?minimal=true and ?fields=f1,f2) |
GET | /api/v1/brokers/racks | Rack topology |
GET | /api/v1/health | Liveness probe |
GET | /api/v1/ready | Readiness probe |
Performance Options
For large clusters, use query parameters to reduce payload size:
/api/v1/cluster/comprehensive?summary=true- strips partition arrays from topics (60-80% smaller)/api/v1/cluster/logdirs?summary=true- broker-level only (90%+ smaller)/api/v1/partitions?minimal=true- strips extra fields/api/v1/partitions?fields=topic,partition,leader,replicas- whitelist specific fields
Follower Lag Monitoring
Follower lag measures the offset gap between leader and follower replicas for each partition. A non-zero lag means followers have not yet replicated all messages from the leader. Persistent or growing lag can indicate slow brokers, network bottlenecks, or disk I/O pressure on follower nodes.
How It Works
Pilot collects follower lag by reading the OffsetLag field from Kafka’s DescribeLogDirs API response. This uses the shared log directory snapshot, so follower-lag requests do not require an additional sweep. The storage cache lifetime is the greater of 60 seconds and four times METADATA_UPDATE_INTERVAL; refresh starts before expiry. Storage and watermarks therefore have different source times. The raw per-replica lag values are aggregated into four granularity levels.
A caller’s timeout limits its wait without cancelling the shared refresh. Expired storage is not returned as current data; a request can still fail if no current snapshot becomes available within its budget.
Granularity Levels
The dashboard and API surface follower lag at four levels:
- Cluster total - sum of all follower offset lag across every replica in the cluster
- Per-broker - total follower lag for all replicas hosted on a given broker
- Per-topic - total follower lag across all partitions of a topic
- Per-partition - lag for each individual follower replica of a partition
URP Classification
When under-replicated partitions (URPs) are detected, Pilot classifies each URP by its root cause:
- Offline broker - the replica’s broker is not responding to metadata requests
- Follower lag - the broker is online but the replica has fallen behind the leader
This distinction is important for operations: offline-broker URPs require infrastructure attention, while follower-lag URPs often resolve on their own once the underlying I/O or network pressure subsides. The proposal engine correctly skips lag-caused URPs during critical fix generation, avoiding unnecessary partition moves for transient replication delays.
Proposal Engine Integration
Follower lag data is fed into the proposal engine to improve rebalancing decisions:
- Inter-broker move exclusion - brokers with active follower lag are excluded as targets for inter-broker moves (moves that add new replica data to a broker). A broker that can’t keep up with existing replication should not receive additional partitions.
- Leader-only moves allowed - leader-only shifts (changing which existing replica is the leader) are still permitted to lag brokers, since they don’t add new data to replicate.
- Critical fix preference - when repairing URPs or rack violations, the engine prefers brokers without follower lag as replacement targets. Lag brokers are still used as a last resort if no other candidates are available.
- Swap exclusion - bidirectional swap moves are skipped when either broker involved has follower lag, since swaps send new data in both directions.
API Endpoint
| Method | Path | Description |
|---|---|---|
GET | /api/v1/cluster/follower-lag | Follower lag data with optional filtering |
Query parameters:
| Parameter | Description |
|---|---|
topic | Filter to a specific topic |
broker | Filter to a specific broker ID |
summary | When true, returns only cluster and broker level totals (omits partition detail) |
Health Threshold
Set PILOT_FOLLOWER_LAG_WARNING_THRESHOLD (default: 1000) to control when total cluster follower lag triggers a warning on the health endpoint.
Prometheus Integration
Pilot exposes Prometheus metrics at /metrics. Key monitoring metrics include:
pilot_cluster_brokers_available/pilot_cluster_brokers_unavailablepilot_cluster_under_replicated_partitionspilot_cluster_offline_partitionspilot_sampler_partitions_trackedpilot_follower_lag_totalpilot_follower_lag_max
See Prometheus Metrics for the complete reference and Monitoring Setup for Prometheus/Grafana configuration.