Skip to Content
FeaturesMaintenance Mode

Maintenance Mode

Maintenance mode allows you to safely take a broker offline for upgrades, hardware maintenance, or decommissioning. Pilot moves all partitions off the broker before it enters maintenance, and the next proposal after maintenance ends moves partitions back onto it.

How It Works

  1. Enter maintenance - Pilot generates a proposal to move all partitions (leaders and replicas) off the target broker
  2. Execute moves - partitions are reassigned to other brokers with throttle management
  3. Broker is empty - the broker has no partition assignments and can be safely stopped
  4. Exit maintenance - the broker is marked online again. The next proposal includes it as a target and moves partitions back to bring it to its fair share. See Refill After Maintenance

API Endpoints

MethodPathDescriptionLicense
POST/api/v1/brokers/{brokerId}/maintenanceEnter maintenance modeYes
DELETE/api/v1/brokers/{brokerId}/maintenanceExit maintenance modeYes

See the API reference for request and response schemas, including the force parameter and the rackCollapses response field.

Enter Maintenance

curl -X POST http://localhost:8080/api/v1/brokers/3/maintenance \ -H "Content-Type: application/json" -d '{}'

This triggers a proposal to evacuate all partitions from broker 3, then executes it with managed throttling.

Force and Replication Factor

A drain preserves each partition’s replication factor by moving its replicas to other brokers. If the cluster cannot host a full replica set for one or more partitions without the broker being drained (for example, not enough remaining brokers, or rack constraints leave no valid placement), the request returns 409 Conflict and names the affected partitions.

Re-submit with "force": true in the request body to perform an audited reduced-RF drain. The forced drain proceeds for the affected partitions at a lower replication factor:

curl -X POST http://localhost:8080/api/v1/brokers/3/maintenance \ -H "Content-Type: application/json" -d '{"force": true}'

force performs a reduced-RF drain only where RF cannot be preserved. It is recorded in the audit log.

Rack-Collapse Warning

A forced drain that preserves the replication factor can still reduce fault-domain spread: to keep RF, Pilot relaxes rack strictness, so two or more replicas of a partition may land in the same rack. When this happens the response reports rackCollapses - the list of affected partitions with their new replica sets and racks. RF is preserved; only rack diversity is sacrificed, and the tradeoff is surfaced explicitly rather than applied silently.

Safety Gating

Maintenance drains run the same live cluster-health pre-flight as other reassignments. Hard-health blockers (controller unhealthy, majority of brokers unavailable, sustained ISR shrink) return 409 Conflict and are not overridable by force. Drains are also blocked while a rolling restart is in progress. See Apply-Time Safety Gating.

Automatic Rebalance After a Drain

When the drain’s moves finish, Pilot generates an activity rebalance of the remaining brokers, excluding the broker in maintenance. With PILOT_HEAL_DRY_RUN=false it applies the plan automatically if it passes the same checks as activity self-healing: measurement readiness and sustained benefit, proposal safety, the activity time window, the 30-minute cooldown shared with activity self-healing, a valid license and PILOT_HEAL_MAX_PARTITIONS_PER_RUN. Otherwise the plan is left for review, and activity self-healing, when enabled, reevaluates it on its next run. This runs for drains started from the UI or the REST API. A drain through MCP tools or the assistant leaves the rebalance to the next proposal or to activity self-healing.

A submitted rebalance is recorded in the audit log as selfheal.apply with trigger post-maintenance, attributed to pilot, with the drained broker in brokerId. A rebalance that fails after its submission was recorded adds an event with status 500. A dry run and a rebalance refused before submission record nothing; the latter is logged.

Exit Maintenance

curl -X DELETE http://localhost:8080/api/v1/brokers/3/maintenance

Pilot marks the broker as online. Exiting maintenance moves no partitions itself: it invalidates the current proposal and queues a fresh calculation that includes the broker as a target. That proposal moves partitions back onto the broker. Apply it like any other proposal, or let activity self-healing apply it.

Refill After Maintenance

The drain’s moves start no 30-minute cooldown, so the first plan after maintenance can move drained partitions back, even when maintenance ends within 30 minutes of the drain. This holds for drains started from the UI, the API, MCP tools and the assistant. Every other move, such as a proposal applied during maintenance, keeps its cooldown. Drain moves still count for settling, so applying the refill can wait until measurements reflect the drained placement.

Pilot keeps its record of recent moves in memory. Unlike maintenance state, it does not survive a restart.

The refill follows the fair-share rule. A returning broker is far below its fair share of every load, so Pilot evens each load until every broker is within 5% of its fair share. The first plan brings the returning broker to at least 90% of its fair share of most loads. Leaders can take a second, smaller plan. A load that cannot get there is shown at its limit with its cause.

URP-Aware Proposals

Maintenance proposals account for under-replicated partitions (URPs). If moving partitions off a broker would create URPs (e.g., RF=1 topics), the proposal warns about it. When the drain cannot preserve a partition’s replication factor it is reported as described under Force and Replication Factor rather than evacuated silently.

Monitoring

The number of brokers in maintenance is tracked:

pilot_cluster_brokers_maintenance

Maintenance state is persisted to the __pilot_broker_state Kafka topic, so it survives Pilot restarts.

Last updated on