Rebalancing Proposals
Pilot generates rebalancing proposals that balance cluster load while limiting partition movement. Proposals are reviewed before execution - nothing moves without explicit approval (unless self-healing is enabled).
How It Works
The proposal engine analyzes cluster state and generates partition reassignments that reduce imbalance across brokers while keeping movement cost low.
The optimizer balances seven metrics per broker:
- Leader count - number of partition leaders per broker
- Follower count - number of follower replicas per broker
- Disk usage - log directory sizes per broker; partitions of compacted topics count at their typical size over the log-cleaner cycle (see Compacted Topics)
- Producer message rate and producer byte rate - inbound traffic per broker
- Consumer message rate and consumer byte rate - consumer traffic per broker
Each broker should carry its fair share of every load. See Even Balance for the rule. This page quotes its default numbers, 10% and 5%, which follow PILOT_BALANCE_THRESHOLD.
Proposal Lifecycle
- Generate - manually via API/UI, or automatically on a schedule
- Review - inspect proposed moves, cost, and the expected effect on each load
- Apply - execute the proposal (requires license)
- Monitor - track reassignment progress in real-time
An empty move list does not establish that the cluster is healthy or balanced. Check the Balance Summary, generation blockers and advisories. A load at its limit, such as one with too few partitions to spread, can leave a broker more than 10% from its fair share when no further moves are proposed. The review separates leader-only changes from replica membership changes and shows any replication-factor changes.
Proposal Status
The main status describes whether Pilot recommends taking action:
| Status | Meaning |
|---|---|
| Collecting measurements | Pilot needs more measurement history before evaluating optional balancing |
| Partition moves in progress | Wait for the current moves to finish |
| Settling after moves | Recent partition moves are still affecting measurements; the remaining wait is shown when it can be estimated |
| Balancing paused | A broker is down and not in maintenance. See Unavailable Brokers |
| No rebalance needed | Every broker is within 10% of its fair share and no topic is concentrated, or no partition change is available |
| Balanced, at its limit | A broker stays more than 10% from its fair share, or a topic stays concentrated, but no plan worth applying brings it closer. See At Its Limit |
| Monitoring | Pilot is checking that a broker stays more than 10% from its fair share or a topic stays concentrated, or no plan improves balance enough to be worth applying |
| Rebalance recommended | The plan is confirmed and passes the current application checks |
| Repairs recommended | The proposal contains only correctness repairs, such as replication-factor or rack-placement fixes |
| Refresh proposal | The plan was not checked within the last two minutes, or the placement changed |
The message names the broker and the load, for example:
Broker 4 is 28% below its fair share of leaders. The plan brings every broker within 5%.These action labels are separate from the API’s cluster balance status. Statuses without a recommendation, such as Monitoring, No rebalance needed, Balanced, at its limit and Balancing paused, show the Balance Summary, current cluster metrics and safety advisories instead of the move table and Apply.
Reading the Balance Summary
When Pilot does not recommend a plan, proposal review shows a Balance Summary card. Its header states the rule with the cluster’s own numbers:
Pilot moves partitions when a broker is more than 10% from its fair share, and evens it to within 5%.The card has one row per load, most urgent first. A row names the broker farthest from its fair share, for example Broker 4 · 28% below, or reads Even when that rounds to 0%.
The bar is centred on the fair share and reaches its ends at 10% below and above it. Two marks show 5%. Its colour rates the farthest broker:
| Colour | Meaning |
|---|---|
| Green | Within 10% of its fair share |
| Red | More than 10% from its fair share |
| Amber | At its limit |
| Grey | Not balanced: the load has no traffic, or, with PILOT_BALANCE_FLOOR set, is below the minimum difference |
A red row adds one line, for example More than 10% below its fair share. An amber row adds the cause, for example At its limit: messages vs bytes.
When a topic is concentrated, a Topics row follows the loads, for example 12 concentrated. Selecting it shows the most concentrated topic, such as orders: 60 of 100 leaders on broker 1, at most 21, and any topic at its limit, such as audit on broker 2: replica placement.
The badge sums up the card:
| Badge | Meaning |
|---|---|
| Balanced | Every broker is within 10% of its fair share of every load |
| Needs rebalancing | The proposal contains a balancing plan; the status says whether it can be applied yet |
| More than 10% from fair share | A broker is more than 10% from its fair share, has no plan yet and is not at its limit |
| Balanced, at its limit | Every broker more than 10% from its fair share is at its limit |
| Balancing paused | A broker is down and not in maintenance. See Unavailable Brokers |
How is this calculated? in the header explains the terms with the cluster’s own settings.
Grey Rows
A rate or disk row reads No traffic (No data for disk) when the cluster carries essentially none of it, and Pilot does not balance it. Leader and follower rows always count. See No Traffic.
With PILOT_BALANCE_FLOOR set, a row reads Below minimum difference when no broker is as far from its equal share as the minimum difference, however uneven the load is in percentage terms. Its note states the largest difference against the minimum.
Farthest Broker
Selecting a row shows the numbers behind its distance:
Now Broker 4 at 35, fair share 50
After the plan Broker 4 at 48, fair share 50
One partition 1 (2% of its share)Broker 4 leads 15 fewer partitions than its fair share. One partition is forgiven, so its distance is 14, or 28% of its share. A load at its limit also shows its cause in one sentence and What would help, for example Add partitions to the topics holding the large partitions so their load can be split.
Red Rows Without a Plan
A red row does not mean moves are coming. Pilot recommends a plan only when the same broker stays more than 10% from its fair share of the same load, on the same side, across four readings, and the plan is worth applying. When no plan is, the status message names the broker:
Broker 4 is 28% below its fair share of leaders. No plan improves balance enough to be worth applying.When the broker does not stay that far off, the message reads No broker stayed more than 10% from its fair share.
With PILOT_BALANCE_FLOOR set, loads below the minimum difference that would otherwise be more than 10% from their fair share get one extra sentence, for example Two metrics are uneven but below the minimum difference (PILOT_BALANCE_FLOOR). Their numbers are on the grey rows. Loads with no traffic are never named.
Observation Progress
While a plan with changes waits for confirmation (decision.reason is benefit_unconfirmed), the status shows progress such as 2 of 3 fresh measurements collected, and the message names the broker:
Broker 4 is 28% below its fair share of leaders. Pilot is checking that this lasts.See Measurement Readiness and Sustained Benefit.
Expected Effect
A plan shown for review has an Expected effect on each load table: per load, the farthest broker now and where the plan leaves it, for example Every broker within 5%. A load at its limit shows its cause under its name, and a Loads at their limit card adds what would help. Grey loads read as in the balance summary. A Topics row shows the concentrated topics now and after the plan.
Pipeline Overview
Proposal generation runs through a multi-phase pipeline in three stages:
- Preparation - collects metadata, checks cluster safety (in-flight reassignments, ISR shrinks), filters excluded or cooling-down partitions, and fixes critical issues like under-replicated partitions and rack violations. If every broker is within 10% of its fair share of every load with traffic and no topic is concentrated, the pipeline short-circuits here.
- Optimization - searches for partition moves that bring brokers to their fair share, trying leadership changes before replica moves, with focused, broad, and swap candidate strategies. Each candidate is scored by how much it improves balance versus how expensive the move is. A second pass reduces proposal cost by pruning unnecessary moves and replacing expensive inter-broker transfers with cheaper leader-only shifts.
- Finalization - moves leadership off concentrated topics, validates correctness, removes changes that do not help, and rebuilds projected metrics from the final placements. Loads at their limit produce advisories that say what would help.
See Architecture for more details.
Reducing Unnecessary Movement
Correctness repairs take priority. When a proposal contains under-replication, rack-placement or replication-factor repairs, it offers those repairs first. Optional balancing is reconsidered after the repairs complete and measurements settle. While a broker is unavailable, Pilot plans only repairs.
For ordinary balancing, the finished plan must be worth applying. All three checks must hold:
- At least one broker more than 10% from its fair share ends within 10%, or at least halfway to 5%, and at least a quarter of one partition closer than it started.
- No plan leaves a broker farther from its fair share than 5% or than where it started, whichever is larger. Only when such a plan brings a broker closer but still leaves it more than 10% from its fair share does Pilot allow other brokers up to 10% instead, and the proposal says so.
- The cluster ends more even overall: the sum of every broker’s squared distance beyond 5%, weighted by load, must fall.
A plan that lowers the leaders over a topic’s cap skips the first check, and passes the third when the sum stays the same.
A plan is not refused for the data it copies. These checks apply to both manual application and activity self-healing.
Pilot evaluates that exact final plan against the latest reading and three earlier ones. In each, the same broker must stay more than 10% from its fair share of the same load, on the same side, or the same topic over its cap on the same broker, and the plan must stay worth applying. The optimizer does not need to choose the same moves in separate calculations. See Measurement Readiness and Sustained Benefit for observation requirements and waiting states.
Finalization evaluates complete partition transitions. It keeps every change that helps and removes changes that do not, without pushing a topic over its leader cap, and keeps each topic’s replicas spread across brokers and racks. Reordering followers without changing replica membership or leadership produces no move. When a broker would end farther from its fair share than the second check allows, Pilot rolls back the optional changes touching it; when no rollback helps, it proposes no optional balancing.
An explicit proposal partition cap can require more trimming. If those cuts violate a topic’s distribution bounds, Pilot rebuilds that topic’s remaining optional plan, retaining changes that bring brokers closer to their fair share on their own, preserve distribution bounds and fit the cap. Changes the cap excludes stay excluded.
These rules favor fewer useful changes; they do not guarantee a mathematically minimal move set. An empty proposal can mean that no plan is worth applying, even when a broker remains more than 10% from its fair share.
Recent Moves
A partition Pilot just moved is not moved again by optional balancing for 30 minutes, so the next plan balances the rest of the cluster. Partitions in a running reassignment are left alone too. When only such partitions could bring a broker closer, the load shows at its limit with the cause recent moves, and Pilot checks again once they are free. A small count fix that needs wider changes waits the same way while the last plan’s moves are in their cooldown.
The cooldown follows every completed move, including applied proposals and self-healing, correctness repairs, topic and partition redistribution, bulk reassignment, audit-log reverts, and moves made through MCP tools or the assistant. It never holds back a correctness repair. Moves that drain a broker entering maintenance start no cooldown, so the first plan after maintenance can move those partitions back. They still count for settling.
The cooldown uses Pilot’s in-memory record of its moves, which a restart clears.
Even Balance
Pilot moves partitions when a broker stays more than 10% from its fair share of a load, and then evens that load until every broker is within 5% of its fair share, give or take one partition.
Both numbers come from PILOT_BALANCE_THRESHOLD (default 5): Pilot aims at the threshold and acts beyond twice the threshold. The rule is the same at every cluster size.
| Term | Meaning |
|---|---|
| Load | Leaders, followers, disk, producer messages, producer bytes, consumer messages and consumer bytes. Each is judged on its own, internal topics included |
| Fair share | The load’s total divided by the active brokers. For followers and disk, the most even split the rack layout allows. A partition bigger than a share becomes its holder’s whole share, and the other brokers split the rest |
| One partition | 1 for leaders and followers. For the other loads, the typical partition: half the load sits in partitions at least this big |
| Distance | How far a broker’s load is from its fair share, less one partition, in percent of the share. It is not a utilization |
- When Pilot acts - the same broker stays more than 10% from its fair share of the same load, on the same side, in four readings: the latest and three earlier ones, at least a minute apart.
- What a plan does - it evens each such load until every broker is within 5%. Leadership changes come first, including exchanges that keep each broker’s leader count; replica moves follow. Among moves that help about equally, the one copying less data wins.
- Small count fixes - when the only loads that start the plan are leader or follower counts beyond 10% by a partition or two, the plan carries just the changes that fix those partitions.
- No broker ends worse - no plan leaves a broker farther from its fair share than 5% or than where it started, whichever is larger. See Reducing Unnecessary Movement.
- Left alone on purpose - a broker within 10% of its fair share never starts a plan.
Balancing reads message and byte rates and partition sizes averaged over about 30 minutes, so short spikes and retention cuts every few minutes do not move partitions. The Brokers and Partitions views, monitoring and what-if simulations show the measured values, so their figures can differ from the current proposal.
| Brokers | Result |
|---|---|
| 4 brokers at 100, 100, 100 and 60 MB/s of producer bytes. The fair share is 90 MB/s and one partition 2.5 MB/s, so broker 4 is 31% below its fair share | Evened, here by leadership changes alone: every broker ends between 83 and 97 MB/s |
| 6 brokers, three at 108 and three at 92 MB/s | Left alone: every broker is within 10% of its fair share |
| 47 brokers, one new and empty | Filled to at least 90% of its fair share in one plan |
| 3 brokers carrying 300 B/s of consumer bytes in total | Left alone: no traffic |
Broker hardware and broker profiles do not change a fair share. Evening disk copies replicas, so a large cluster can receive a plan that copies hundreds of GB. See Troubleshooting.
A lower threshold evens brokers more closely and moves more partitions; a higher one tolerates more. With PILOT_BALANCE_THRESHOLD=10, Pilot acts beyond 20% and evens to within 10%. To keep a quiet dev or test cluster from moving partitions for small differences, set an optional minimum difference.
No Traffic
Metrics with essentially no load (on average below 1 msg/s, 1 KiB/s or 1 MiB per broker) show as no traffic and are not balanced. A percentage of near-zero load is noise: brokers at 10, 20 and 300 B/s are very uneven in percentage terms, but carry nothing worth moving.
| Metric | Not balanced below, on average per broker |
|---|---|
| Producer and consumer messages | 1 msg/s |
| Producer and consumer bytes | 1 KiB/s |
| Disk | 1 MiB |
- This is not a setting. Real small clusters sit far above these levels: a dev cluster averaging 127 msg/s and 25 KB/s per broker is balanced normally.
- Leader and follower counts always count.
- A load with no traffic reads a distance of 0 for every broker, so it never starts a plan, is never at its limit and does not count for drift detection. Its row reads No traffic (No data for disk), and the proposal reason does not name it.
- Each calculation decides from the placement it starts from. Each benefit observation decides from its own measurements.
Compacted Topics
The log cleaner keeps a compacted partition in a sawtooth: the log grows, the cleaner cuts it, and it grows again. A busy __consumer_offsets partition can swing between a few MB and about 100 MB every few hours. Balancing on the measured size would let every cut shift disk imbalance by several points and trigger moves that the next cut undoes.
For topics whose cleanup.policy contains compact, internal topics included, proposals weigh each partition at its typical size over its log-cleaner cycle: a 12-hour average of its measured size, never further from the current size than the cleaner’s recent cut.
- Until Pilot has seen the cleaner cut a partition, one segment (
segment.bytes) stands in for the cut: after a start, the average begins at the first size Pilot reads and stays within one segment of the measured size, so a log that only grows counts up to one segment below it. A partition created while Pilot runs counts at its measured size until Pilot has seen a repeating cut: one the cleaner makes again, or one the log grows back within 12 hours. - A partition whose log takes more than 12 hours to grow a cut back counts at the middle of the cut, since a 12-hour average would follow so slow a cycle down and back up. It stays there while the log keeps growing back, for cycles of up to about 8 days, and only for partitions that existed when Pilot started. If the log stops growing back, half the cut stays in the estimate for about a day and a half, then fades by half every 33 hours.
- A one-off shrink, such as a retention cut,
DeleteRecordsor a first compaction after a bulk load, counts at its new size straight away, with two exceptions. The first drop of up to two segments on a log of one segment or more is taken for a cleaner cut: the removed bytes leave the estimate over about a day, or like a slow cycle’s if the log was growing slowly. A shrink of up to twice the cut within 12 hours of a slow partition’s cut counts as part of that cut until the next cut. - The estimate lives in memory. After a restart it starts again from the first sizes Pilot reads, so until its first cut a busy partition can count up to half its swing away from its typical size. From that cut, the estimate moves to the middle of the cut by at most 1.4% of the cut per 10 minutes and gets there in about 6 hours.
- Topics without
compactin their policy count at their measured size averaged over about 30 minutes, so a log that only grows counts about 30 minutes of growth short.
Proposal disk figures (broker disk totals, fair shares and distances, disk utilization, and a plan’s bytes to move and duration) use these sizes, so they can differ from the Brokers and Partitions views. Those views, monitoring, what-if and /api/v1/cluster/logdirs keep the measured size.
At Its Limit
A broker more than 10% from its fair share that no plan worth applying brings closer is at its limit. When every such broker is at its limit, the status reads Balanced, at its limit, and the message names the loads:
Pilot can't even out producer messages any further.The Balance Summary names the broker and the cause. structuralLimits[].reason says both in one sentence:
Producer messages on broker 1 are 55% above their fair share and at their limit: messages and bytes can't both be even.cause | Label | Sentence |
|---|---|---|
partition_size | large partition | One partition is a large part of this broker’s share. |
rack_layout | rack layout | The rack layout limits how evenly this load can spread. |
messages_vs_bytes | messages vs bytes | Messages and bytes can’t both be even. |
messages_vs_bytes on leaders or followers | messages vs bytes | Evening them would push producer messages on broker 4 more than 10% from their fair share. |
no_move_helps | no move helps | Every move that would help unbalances another broker. |
recent_moves | recent moves | Recent moves hold the partitions that would help. |
A partition counts as large at a quarter of the broker’s share or more. Recent moves are the 30-minute cooldown and in-flight reassignments.
Pilot keeps balancing the other loads and checks the limit again in every calculation. See Troubleshooting for what helps.
Concentrated Topics
A broker that leads most of a topic takes all of the topic’s clients with it when it fails or restarts. A topic is concentrated when one broker leads more of its partitions than the cap: 1.2 times the broker’s share of the topic, rounded up, plus one. The share is the topic’s partitions divided by the active brokers in the racks that hold its replicas, or by every active broker without racks. With 100 partitions on 6 brokers, the cap is 21.
Pilot moves leadership off the broker, down to the cap, to replicas on brokers under it. No data is copied, and no broker ends farther from its fair share than Even Balance allows. The repair joins the balancing plan, or forms a plan of its own in an otherwise balanced cluster. Topics within the cap are left alone.
| State | Message |
|---|---|
| Ready | Topic orders has 60 of its 100 leaders on broker 1. The plan spreads them. |
| Ready, with balancing | The balancing message, then: It also spreads 12 topics. |
| Checking | Topic orders has 60 of its 100 leaders on broker 1. Pilot is checking that this lasts. |
| At its limit | Pilot can’t spread topic orders any further. |
- Checking - like a load, a concentration must last four readings. It counts only once the brokers that can lead (available and not in maintenance) have been unchanged for 10 minutes; a restart of Pilot starts the 10 minutes again. That gives Kafka’s automatic leader rebalancing, where enabled, time to hand leadership back to a broker that returns.
- At its limit - no leadership change lowers the topic further. It does not count toward a plan, and Pilot checks again in every calculation.
cause | Label | Sentence |
|---|---|---|
replica_placement | replica placement | The topic’s replicas sit on too few brokers to spread its leaders. |
recent_moves | recent moves | Recent moves hold the partitions that would help. |
no_move_helps | no move helps | Every move that would help unbalances another broker. |
Topics are not repaired while balancing is paused, and brokers in maintenance do not count. Where the rack layout pins leadership to a few brokers, the cap is off. The proposal reports concentrated topics in topicSpread; see the OpenAPI specification .
Unavailable Brokers
While a broker is down and not in maintenance, balancing is paused. Pilot plans only repairs (under-replicated partitions, rack placement and replication factor) and does not balance around the missing broker, whose load comes back when it returns. With nothing to repair, the status reads Balancing paused:
Balancing is paused while broker 5 is down. If it stays down, drain it with maintenance mode.balanceHold.brokers lists the brokers. Drain a broker that stays down with maintenance mode, so Pilot balances the remaining brokers.
When every broker is down or in maintenance, the message reads Balancing is paused while no broker is active.
Minimum Difference
PILOT_BALANCE_FLOOR is optional and unset by default, and Pilot then balances differences of any size. Set it to keep a quiet dev or test cluster from moving partitions for small differences: a rate or disk load is then balanced only when some broker is at least the minimum difference from its equal share, the load’s total divided by the number of active brokers. Below it, the load is left alone however uneven it is in percentage terms. Leader and follower counts are not affected.
PILOT_BALANCE_FLOOR="bytes=2MB/s disk=2GiB" # ignore smaller per-broker differences
PILOT_BALANCE_FLOOR="disk=100GiB" # bytes= uses 2 MB/s
PILOT_BALANCE_FLOOR="bytes=2MB/s msgs=5k/s disk=2GiB" # explicit message rate
PILOT_BALANCE_FLOOR=off # same as unset- Settings are separated by spaces or commas. Settings left out use 2 MB/s for
bytes=and 2 GiB fordisk=. bytes=applies to producer and consumer bytes. It takes a byte rate (2MB/s,500KiB/s) or a bit rate (16Mbit/s).Mb/sis rejected as ambiguous.msgs=(ormessages=) applies to producer and consumer messages. Left out or set toauto, it follows the byte rate at the average message size of each direction, so a message-rate difference counts when the bytes it carries would. Traffic without bytes or messages uses 10k msg/s. An explicit value takes messages per second:10k/s,10000/sor250msg/s.disk=applies to disk and takes a size:2GiB,512MiBor1.5TB.- Each value must be positive.
offcannot be combined with other settings. - An invalid value stops Pilot at startup. A minimum difference in force is logged once at startup.
With PILOT_BALANCE_FLOOR="bytes=2MB/s disk=2GiB":
| Load | Largest difference | Result |
|---|---|---|
| Producer messages on a dev cluster: busiest broker 167 msg/s, equal share 127 msg/s | 40.1 msg/s, below the derived minimum of about 10k msg/s | Left alone |
| Producer bytes: two brokers at 150 MB/s, four at 100 MB/s | 33.3 MB/s, above 2 MB/s | Evened out |
With a minimum difference set:
- Each calculation decides from the placement it starts from, derived message rates included. Each benefit observation decides from its own measurements, so a plan is confirmed only while the load stays above the minimum in every observation.
- A load below the minimum reads a distance of 0 for every broker, so it never starts a plan and is never at its limit. Drift detection ignores it.
- Optional moves never push a load below the minimum to or past it. Finalization rolls back changes that would; if no rollback helps, no optional plan is recommended. Correctness repairs are exempt.
- When a load below the minimum is uneven in percentage terms, the proposal reason names it, for example
not balanced, below the minimum difference: consumer bytes 21.3% uneven, largest difference 7.24 KB/s per broker (minimum 2 MB/s).
API Fields
Each proposal carries the fair-share view:
| Field | Meaning |
|---|---|
fairShare | line and target (10 and 5 at the default), every active broker’s fair share per load (shares), and per load the farthest broker before and after the plan, the size of one partition and whether the rack layout set the shares |
structuralLimits[] | Each load at its limit, with brokerId, distance, cause and its reason sentence. For a leader or follower count with messages_vs_bytes, conflictMetric and conflictBrokerId name the rate load and broker that evening it would push more than 10% from its fair share. An entry with a cause always carries brokerId |
balanceHold | Set while balancing is paused for a down broker: cause (broker_unavailable) and brokers. See Unavailable Brokers |
decision.reason is balance_held | Set while balancing is paused for a down broker (see balanceHold) or while no broker is active. decision.state is then blocked, nextAction is wait, and message is the hold sentence. See Unavailable Brokers |
benefitEvidence.tripped[] | The brokers the plan is for: each stayed more than 10% from its fair share of a load (metric), on the same side, in every confirmation reading |
topicSpread, benefitEvidence.topicBreaches[] | Concentrated topics now and after the plan, and those the plan is for |
metrics.effective | Half the farthest broker’s distance per load, so a value above metrics.threshold means a broker is more than 10% from its fair share. It is omitted when it would equal current and projected |
metrics.current and metrics.projected report each load’s relative imbalance, the average distance of the brokers from the equal share; advisories.metrics mirrors them. They are for display and do not decide. See the OpenAPI specification for the full schemas: FairShareReport, StructuralLimit, BenefitEvidence and TopicSpreadReport.
Generate, current, by-ID and list responses carry a balance object beside each proposal. The MCP generate_proposal and get_proposal tools return the same object.
| Field | Meaning |
|---|---|
gate | The minimum difference in force. Without one, off is true, summary is off and every value is 0 |
gate.byteRate, gate.disk | Minimum difference for producer and consumer bytes (B/s) and for disk (bytes) |
gate.producerMessageRate, gate.consumerMessageRate | Minimum difference for messages (msg/s); messageRate repeats the producer value |
gate.messageRateDerived, gate.detail | true when the message rates follow the byte rate at the average message size; detail then explains them |
gate.summary | Readable minimum difference, for example 2 MB/s of traffic or 2 GiB of data per broker, or off |
current, projected | One entry per metric, before and after the proposed changes |
[].metric, [].unit | Metric key (as in metrics.current) and its unit: count, B, msg/s or B/s |
[].imbalance | Relative imbalance in percent, matching metrics.current or metrics.projected |
[].material | true when the metric counts for balancing: it has traffic and, with a minimum difference set, some broker is at least that far from its equal share |
[].idle | true when the metric has no traffic; omitted otherwise |
[].gate | The metric’s minimum difference in its unit; 0 means none |
[].largestDeviation, [].share | Largest per-broker difference from the equal share, and the equal share |
[].busiest, [].quietest | brokerId and value of the brokers with the highest and lowest value |
[].summary | One-line explanation, for example Producer bytes: 166.2% uneven, broker 1 at 400 MB/s, broker 2 at 200 KB/s, share 66.8 MB/s or Consumer bytes: no traffic. Below a minimum difference it reads, for example, Disk: 11.3% uneven, largest difference 99 MiB per broker, below the 2 GiB minimum difference |
The Prometheus gauges pilot_proposal_imbalance_material and pilot_proposal_imbalance_largest_deviation export material and largestDeviation of the current placement; pilot_proposal_variance_current and pilot_proposal_variance_projected export the relative imbalance. See Metrics.
Broker Profiles
A broker profile describes each broker’s hardware. It does not change balancing: capacity never makes an imbalance count more or less, and fair shares do not depend on broker hardware. A profile:
- Reports each broker’s utilization in
brokerActivityScores[].capacityUtilization(network) anddiskUtilization, in percent after the proposed changes. - Switches post-move settling to evidence when it sets
network=.
Capacity is meant for overload protection, such as keeping a broker below its limits or giving larger brokers a larger share. This is planned and not active yet, so give every broker room for its fair share.
| Setting | Meaning |
|---|---|
network= | Sustained throughput per direction: the lowest of the NIC speed, the instance’s baseline (not burst) network bandwidth and the log volume’s write throughput |
disk= | Kafka log storage per broker; for JBOD, the sum of the log directories |
A profile can set either or both. Network utilization is the busier direction divided by network=: a broker receives the produce traffic of every replica it hosts, and sends the consumer traffic plus one copy of the produce traffic per follower for each partition it leads. Disk utilization is stored bytes divided by disk=.
Homogeneous Fleet
A rule without a selector applies to every broker:
PILOT_BROKER_PROFILE="network=10Gbit/s disk=2TiB"Mixed Fleet
A rule with a selector and a colon overrides the default for matching brokers. PILOT_BROKER_PROFILE separates rules with ;. PILOT_BROKER_PROFILE_FILE takes one rule per line, and # starts a comment:
network=10Gbit/s disk=2TiB # default for every broker
rack=eu-west-1c: network=25Gbit/s # newer instance type in one zone
host=kafka-big-*: network=25Gbit/s # advertised host from cluster metadata
broker=7-9,12: disk=8TiB # larger log volumesrack=andhost=take glob patterns (*,?,[...]) without spaces.host=matches the advertised host without the port.broker=takes IDs and ranges.- A selector ends at the last colon of its rule, because settings never contain one. Rack names and IPv6 hosts with colons work as written, for example
rack=eu:west:1: disk=4TiB. - Each setting resolves on its own, from the matching rules that set it. A
broker=rule beats ahost=rule, which beats arack=rule, which beats the default. Rules of the same kind are equally specific whatever their pattern, and the one written last wins. Broker 8 ineu-west-1cresolves to 25 Gbit/s and 8 TiB. - A selector can appear in several rules; each rule keeps its own position. In
rack=eu-*: disk=4TiB; rack=eu-west-1c: network=25Gbit/s; rack=eu-*: network=40Gbit/s, a broker ineu-west-1cresolves to 40 Gbit/s. A selector can set each setting only once. - Every setting a selector uses needs a default, so every broker resolves.
Units and Validation
- Units are required and follow the number without a space (
10Gbit/s, not10 Gbit/s). Network takes bit rates (10Gbit/s,25Gbps,100Mbit) or byte rates (1.25GB/s,800MiB/s). Disk takes sizes (2TiB,1.5TB,500GiB). GbandGb/sare rejected as ambiguous, as is a size fornetwork=or a rate fordisk=.- Set only one of the two variables. Blank values count as unset.
- An invalid profile stops Pilot at startup and names the failing rule or line. A valid profile is summarized once in the startup log.
- Selectors that match no broker, and brokers above 100% of their configured capacity (usually a unit mistake), are logged as warnings. The warnings repeat only when the resolved capacities or the set of affected brokers change.
Settling After Moves
After Pilot’s own moves end, optional balancing waits until measurements reflect the new placement. Repair-only proposals do not wait.
Without a network profile (no profile, or disk= only), Pilot waits an adaptive period after the last reassignment: normally 5 to 30 minutes, according to the share of partitions and leaders moved. The UI shows the remaining wait.
With a network profile, Pilot waits for evidence instead of a timer. Settling ends when:
- No partition moves are active.
- Moved replicas are back in the in-sync replica set.
- Leadership of the moved partitions is on the planned brokers.
- At least two minutes have passed since the last move ended. This covers the short client and measurement dip after a leader change.
- The rate, consumer and storage measurements of the moved partitions were all taken after the moves.
A move that failed, timed out or was cancelled also starts settling, because it can still change placement: a cancel rolls the replicas back. For these partitions, Pilot checks the placement the cluster now holds instead of the planned one: all replicas in sync and a leader among them.
Benefit evidence then counts only observations taken after the moves. Settling never ends sooner than two minutes after the last move; how much longer it takes depends on the cluster. The UI shows what settling is waiting for and, when it can be estimated, the remaining time. Settling considers moves that ended in the last 30 minutes: if a condition never clears, optional balancing resumes 30 minutes after the last move, and a lasting ISR shrink is still caught by the safety checks.
In both modes:
- Partitions Pilot moved stay out of optional balancing plans for 30 minutes, so the next plan balances the rest of the cluster instead of moving them again. Partitions moved off a broker entering maintenance are exempt, so a broker leaving maintenance can be refilled with them. Drain moves still count for settling. See Recent Moves.
- Settling decides when a plan can be applied, not which moves it contains.
- Activity self-healing applies at most once per 30 minutes, and waits while any partition is still in its 30-minute cooldown, so each automatic plan can move every partition. Repairs do not wait. See Self-Healing.
- The end of settling allows reevaluation; it does not approve or execute a proposal.
- Settling uses Pilot’s in-memory record of its moves, which a restart clears.
Rack-Safety Advisory
When PILOT_BALANCE_RACK_AWARE=true, a proposal can include a rackSafety advisory in the proposal payload (and the UI advisories panel). It fires at one of two severities:
high- rack-awareness cannot be trusted. Fires when at least one replica-hosting broker reports nobroker.rack, at least one partition’s replica set may be physically co-located, or the cluster has a single rack while the largest replication factor is 2 or more. In these cases a full replica set can land on co-located brokers while appearing rack-aware. This is common on a mixed-rollout cluster where some brokers have not yet been assigned abroker.rack.warning- the only finding is fewer distinct racks than the largest replication factor (2 or more racks, complete rack metadata, no co-location detected). Placement still spreads replicas at the maximally even ceil(RF / numRacks) per rack, so a replica set cannot collapse into one rack. The residual risk: a single rack failure removes up to ceil(RF / numRacks) replicas of a partition at once.
The warning tier is gated on min.insync.replicas; the high tier is not. For every partition whose replication factor exceeds numRacks, the engine computes the worst case for a single rack loss: RF - ceil(RF / numRacks) assigned replicas survive. This calculation assumes the assigned replicas are in sync before the failure; it does not verify current producer availability. Comparing survivors against each topic’s effective min.insync.replicas gives one of three outcomes:
- Suppressed - every affected topic has a known
min.insync.replicasand no partition drops below it under that assumption, so no advisory is emitted. Check current ISR separately before concluding thatacks=allproducers can keep writing. - Sharpened - at least one partition would drop below its topic’s
min.insync.replicas. The summary states the worst case concretely: the topic, how many in-sync replicas a single rack failure would leave versus itsmin.insync.replicas, how many topics are affected, and thatacks=allproducers would stall. - Generic -
min.insync.replicascannot be determined for at least one affected topic. The advisory keeps the topology-only warning describing the residual rack-loss exposure.
Example: 4 brokers in 2 racks with RF 4 leave 2 assigned replicas after a rack loss. If all replicas were in sync, min.insync.replicas=2 permits writes and the advisory is suppressed; with min.insync.replicas=3 it warns that producers would stall.
The advisory carries:
| Field | Meaning |
|---|---|
severity | high or warning (see above) |
summary | Human-readable description of the condition |
numRacks | Distinct non-empty racks the engine sees |
maxReplicationFactor | Largest replication factor among hosted partitions. When numRacks < maxReplicationFactor a full replica set cannot be spread one-per-rack |
unrackedBrokers | Replica-hosting brokers that report no rack |
collocatedPartitions / collocatedPartitionCount | Partitions whose current replica set is not rack-aware under the conservative reading (sample list plus the full total) |
remediation | What to do. For high: assign broker.rack on every broker and ensure at least RF distinct racks. For warning: add racks or lower RF if the rack-loss exposure is unacceptable, and review min.insync.replicas against the per-rack replica count |
Until a high condition is resolved, treat rack-aware placement as best-effort: replicas that appear rack-distributed may be physically co-located in the same fault domain.
Always-Ready Proposals
Pilot continuously watches for cluster changes and regenerates proposals automatically. When topology changes (broker added/removed, partition count changes) or metric drift exceeds an internal threshold, Pilot updates the proposal. Detection cadence, drift sensitivity, and the regeneration interval are fixed engine defaults and need no configuration.
Explicit manual generation with POST /api/v1/proposals/generate and {"immediate": true} requests a fresh calculation when generation is available. Generation can run while measurements are settling, but optional application waits. A request without that option retrieves the current background proposal. Repair-only and empty proposals use the normal generation interval. Metric drift is checked against current sampled inputs. Refreshing readiness or benefit evidence does not change the time the proposal was calculated.
Changing a broker’s maintenance state invalidates the current proposal and queues a fresh calculation. Benefit evidence from an earlier broker eligibility state cannot confirm the new plan.
Plan Stability
Pilot re-checks a balancing plan on fresh readings every 90 seconds and on metric drift. It keeps the plan, with the same ID and changes, while it is still worth applying and still covers every broker more than 10% from its fair share. It makes a new plan when partitions move, a broker or setting changes, a cooldown ends, the check fails, or the plan is 30 minutes old.
The page never swaps the plan you are reading. When Pilot moves on to another proposal, it says:
- A newer plan is available. - the new proposal has changes
- This plan is no longer current. - the new proposal has no changes
- The plan changed. Review the new plan. - shown in an open review
Configuration
| Variable | Default | Description |
|---|---|---|
PILOT_BALANCE_THRESHOLD | 5.0 | Applies per broker and load: Pilot acts when a broker stays more than twice this percentage from its fair share (10% at the default), and evens that load until every broker is within this percentage. See Even Balance |
PILOT_BALANCE_RACK_AWARE | true | Enforce rack-aware placement in proposals |
PILOT_BALANCE_MIN_RF | 0 | Minimum replication factor (0 = disabled) |
PILOT_BALANCE_WORKERS | 0 | Candidate-generation goroutines (0 = auto, uses GOMAXPROCS). Does not affect proposal output |
PILOT_BALANCE_FLOOR | "" (off) | Optional minimum difference for rate and disk balancing, for example bytes=2MB/s disk=2GiB to keep quiet test clusters from moving partitions for small differences. Unset or off balances differences of any size. See Minimum Difference |
PILOT_BROKER_PROFILE | "" | Broker hardware for utilization reporting and evidence-based settling; does not change balancing. See Broker Profiles |
PILOT_BROKER_PROFILE_FILE | "" | Broker profile file, one rule per line. Mutually exclusive with PILOT_BROKER_PROFILE |
PILOT_MOVE_MAX_PER_BROKER | 20 | Max concurrent moves per broker |
Scoring Model
Proposals are scored using a lexicographic scoring model:
- Weighted squared excess - per load, how far each broker’s distance exceeds 5%, squared and weighted by the load’s importance
- Worst excess - the load farthest from even
- Movement cost - total cost of proposed moves (leader-only < inter-broker < cross-rack)
The weights only rank trade-offs between loads: 3 for leaders and followers, 2.5 for disk, 2 for byte rates and 1.5 for message rates. A built-in cost weight balances movement cost against balance improvement so the engine prefers smaller, cheaper move sets when a comparable result is achievable.
API Endpoints
| Method | Path | Description | License |
|---|---|---|---|
POST | /api/v1/proposals/generate | Retrieve the current proposal, or recalculate with {"immediate": true} | No |
GET | /api/v1/proposals/current | Get the current candidate and authoritative readiness decision | No |
GET | /api/v1/proposals | List stored proposals with readiness decisions | No |
GET | /api/v1/proposals/{proposalId} | Get proposal details | No |
POST | /api/v1/proposals/{proposalId}/apply | Execute a proposal | Yes |
DELETE | /api/v1/proposals/{proposalId} | Delete a proposal | No |
See the API reference for request and response schemas.
Apply-Time Safety Gating
Applying a proposal runs a live cluster-health pre-flight at submit time, independent of the safety snapshot recorded when the proposal was generated. This catches a controller-health change or sustained ISR shrink that happened after generation but did not move the topology fingerprint.
- Hard-health blockers return 409 and are not overridable. Controller unhealthy, a majority of brokers unavailable, or a sustained ISR shrink block the apply outright.
force=truedoes not bypass them. - Ordinary balancing requires ready measurements and a confirmed plan checked within the last two minutes (
checkedAt). Manual application and activity self-healing share these requirements and recheck the exact plan against current measurements before submission. Repair-only proposals are exempt from these balance requirements; health and execution safety checks still apply. - A changed balance plan needs recalculation. If apply-time filtering removes moves, Pilot refuses the remaining optional subset because the original benefit estimate no longer describes it. Repair-only subsets can proceed.
- Softer reasons can be overridden with
force=true. Reasons such as a stale topology fingerprint or in-flight reassignments are reported as 409 with guidance to re-submit with?force=true. A forced apply is recorded in the audit log. - Applies are blocked during a rolling restart. A reassignment cannot run while a rolling restart is in progress.
The current, generate and by-ID responses expose an authoritative decision; list entries include the same decision for each stored proposal. Overview, proposal review and the assistant use decision.state, title, message, applicationAllowed and nextAction. The compatibility fields metadata.applicationAllowed, applicationBlockReason and applicationBlockMessage mirror it. data_not_ready means measurement history, freshness or settling requirements are not met; benefit_unconfirmed means the proposed plan has not demonstrated sustained benefit. The apply endpoint still runs live pre-flight checks. The explicit API force=true option can override soft balance-admission requirements with an audit record, but cannot authorize an altered optional subset or bypass hard-health blockers.
A candidate can contain possible moves while its decision is Settling after moves or Monitoring. These moves are diagnostic candidates, not a current recommendation. The UI shows the status and current metrics until the decision permits review. Storage status ready and optimizer status needs_rebalancing do not authorize application.
decision.retryAt supplies the earliest time to evaluate again: a refresh deadline, the end of the adaptive settling wait or, with a network broker profile, the estimated time of the first post-move measurements. It is absent while moved replicas or leadership are still settling. Reaching it allows another evaluation, not automatic approval. A ready decision includes validUntil; clients must refresh after it expires. A changed candidate invalidates an open review so it cannot submit an older plan.
Reading status does not generate another plan, advance benefit observations or refresh the calculation timestamp. The MCP generate_proposal tool returns the same server-selected current candidate and decision. Use the REST generation endpoint with immediate: true to request explicit recalculation.
Live health checks also protect topic and partition redistribution, bulk reassignment, broker maintenance, and audit-log revert. See Partition Reassignments for those shared checks.
Measurement Readiness and Sustained Benefit
After startup, Pilot normally collects about five minutes of measurement history before evaluating a proposal. Overview and proposal review show elapsed history, a progress bar and the estimated time remaining. The countdown describes the history window; fresh and sufficiently complete measurements must also arrive before balancing can be evaluated.
When the window has elapsed, Pilot reports whether it is waiting for fresh samples, waiting for complete partition measurements or preparing the proposal. It does not invent an ETA for a missing source. The API exposes these stages in decision.collection, separately from the benefit observations in decision.progress. A complete history window is not permission to apply a proposal.
Optional balancing uses two separate checks:
- Measurements are ready. The averaging window is fully covered, metadata is current, no partition moves are active and Pilot’s recent moves have settled.
- The exact plan remains worthwhile. In the latest measurements and three recent observations, the same broker must stay more than 10% from its fair share of the same load, on the same side, or the same topic over its cap on the same broker, and the complete final move list must stay worth applying.
Pilot reuses its bounded in-memory measurement history, retaining up to seven observations captured at least one minute apart and considering evidence no older than six minutes. Required source timestamps must advance between selected observations. Refreshing the page or rereading cached measurements does not add evidence, and slower storage measurements can delay progress. Existing compatible observations can confirm a newly generated plan without waiting for three more optimizer runs.
Observations must share the current broker eligibility, partition membership, replication factors and balance configuration. The plan’s starting placements and destination eligibility are checked again before application. Correctness repairs bypass the optional benefit requirement; safety and license checks still apply.
The proposal’s benefitEvidence.status explains the result:
| Status | Meaning |
|---|---|
confirmed | The plan stays worth applying across the required observations and current measurements, for the brokers listed in tripped and the topics in topicBreaches |
collecting | More fresh, compatible observations are needed |
unstable | No broker stayed more than 10% from its fair share and no topic over its cap (reasonCode is trigger_not_persistent), the plan stopped being worth applying, or a placement or eligibility change requires recalculation |
not_required | The proposal is empty or contains only correctness repairs |
observedSamples and requiredSamples describe progress. tripped lists the brokers the plan is for, with load and side. reason explains a pending result, and evaluatedAt records when it was checked. See the API reference for the complete schema. Confirmed benefit alone does not authorize application; use decision.applicationAllowed for the combined decision.
After Pilot’s own moves, readiness also waits for settling. With a network broker profile, only observations taken after the moves count toward the three. The end of settling allows reevaluation; it does not automatically approve or execute a proposal. If measurements are ready but benefit evidence is incomplete or inconsistent, the status becomes Monitoring.
Proposal Confidence
The API retains the confidence score and level as measurement diagnostics for compatibility. They are not a probability that a plan will improve the cluster, do not compare the optimizer’s chosen targets across calculations, and do not require a High level for application. The UI presents measurement readiness and the action decision instead of a confidence percentage. API clients should use decision.applicationAllowed and its explanation.
Proposal Storage
Pilot stores up to 100 proposals in memory with statuses: pending, ready, applied, cancelled, failed. Proposals are not persisted to disk - they are regenerated as needed.