---
title: "Maintenance Mode - Pilot Docs"
description: "Broker maintenance mode for safe rolling updates"
url: "https://docs.calinora.io/features/maintenance-mode/"
---

# Maintenance Mode

Maintenance mode allows you to safely take a broker offline for upgrades, hardware maintenance, or decommissioning. Pilot moves all partitions off the broker before it enters maintenance, and the next proposal after maintenance ends moves partitions back onto it.

## How It Works

1. **Enter maintenance** - Pilot generates a proposal to move all partitions (leaders and replicas) off the target broker
2. **Execute moves** - partitions are reassigned to other brokers with throttle management
3. **Broker is empty** - the broker has no partition assignments and can be safely stopped
4. **Exit maintenance** - the broker is marked online again. The next proposal includes it as a target and moves partitions back to bring it to its fair share. See [Refill After Maintenance](https://docs.calinora.io/features/maintenance-mode/#refill-after-maintenance)

## API Endpoints

| Method | Path | Description | License |
| - | - | - | - |
| `POST` | `/api/v1/brokers/{brokerId}/maintenance` | Enter maintenance mode | Yes |
| `DELETE` | `/api/v1/brokers/{brokerId}/maintenance` | Exit maintenance mode | Yes |

See the [API reference](https://docs.calinora.io/api/reference/) for request and response schemas, including the `force` parameter and the `rackCollapses` response field.

### Enter Maintenance

```bash
curl -X POST http://localhost:8080/api/v1/brokers/3/maintenance \
  -H "Content-Type: application/json" -d '{}'
```

This triggers a proposal to evacuate all partitions from broker 3, then executes it with managed throttling.

#### Force and Replication Factor

A drain preserves each partition’s replication factor by moving its replicas to other brokers. If the cluster cannot host a full replica set for one or more partitions without the broker being drained (for example, not enough remaining brokers, or rack constraints leave no valid placement), the request returns `409 Conflict` and names the affected partitions.

Re-submit with `"force": true` in the request body to perform an audited reduced-RF drain. The forced drain proceeds for the affected partitions at a lower replication factor:

```bash
curl -X POST http://localhost:8080/api/v1/brokers/3/maintenance \
  -H "Content-Type: application/json" -d '{"force": true}'
```

`force` performs a reduced-RF drain only where RF cannot be preserved. It is recorded in the [audit log](https://docs.calinora.io/configuration/audit-logging/).

#### Rack-Collapse Warning

A forced drain that preserves the replication factor can still reduce fault-domain spread: to keep RF, Pilot relaxes rack strictness, so two or more replicas of a partition may land in the same rack. When this happens the response reports `rackCollapses` - the list of affected partitions with their new replica sets and racks. RF is preserved; only rack diversity is sacrificed, and the tradeoff is surfaced explicitly rather than applied silently.

#### Safety Gating

Maintenance drains run the same live cluster-health pre-flight as other reassignments. Hard-health blockers (controller unhealthy, majority of brokers unavailable, sustained ISR shrink) return `409 Conflict` and are not overridable by `force`. Drains are also blocked while a rolling restart is in progress. See [Apply-Time Safety Gating](https://docs.calinora.io/features/reassignments/#apply-time-safety-gating).

#### Automatic Rebalance After a Drain

When the drain’s moves finish, Pilot generates an activity rebalance of the remaining brokers, excluding the broker in maintenance. With `PILOT_HEAL_DRY_RUN=false` it applies the plan automatically if it passes the same checks as [activity self-healing](https://docs.calinora.io/features/self-healing/#activity-based-healing): measurement readiness and sustained benefit, proposal safety, the activity time window, the 30-minute cooldown shared with activity self-healing, a valid license and `PILOT_HEAL_MAX_PARTITIONS_PER_RUN`. Otherwise the plan is left for review, and activity self-healing, when enabled, reevaluates it on its next run. This runs for drains started from the UI or the REST API. A drain through MCP tools or the assistant leaves the rebalance to the next proposal or to activity self-healing.

A submitted rebalance is recorded in the [audit log](https://docs.calinora.io/configuration/audit-logging/#automatic-applies) as `selfheal.apply` with trigger `post-maintenance`, attributed to `pilot`, with the drained broker in `brokerId`. A rebalance that fails after its submission was recorded adds an event with status `500`. A dry run and a rebalance refused before submission record nothing; the latter is logged.

### Exit Maintenance

```bash
curl -X DELETE http://localhost:8080/api/v1/brokers/3/maintenance
```

Pilot marks the broker as online. Exiting maintenance moves no partitions itself: it invalidates the current proposal and queues a fresh calculation that includes the broker as a target. That proposal moves partitions back onto the broker. Apply it like any other proposal, or let [activity self-healing](https://docs.calinora.io/features/self-healing/#activity-based-healing) apply it.

#### Refill After Maintenance

The drain’s moves start no 30-minute [cooldown](https://docs.calinora.io/features/proposals/#recent-moves), so the first plan after maintenance can move drained partitions back, even when maintenance ends within 30 minutes of the drain. This holds for drains started from the UI, the API, MCP tools and the assistant. Every other move, such as a proposal applied during maintenance, keeps its cooldown. Drain moves still count for [settling](https://docs.calinora.io/features/proposals/#settling-after-moves), so applying the refill can wait until measurements reflect the drained placement.

Pilot keeps its record of recent moves in memory. Unlike maintenance state, it does not survive a restart.

The refill follows the [fair-share rule](https://docs.calinora.io/features/proposals/#even-balance). A returning broker is far below its fair share of every load, so Pilot evens each load until every broker is within 5% of its fair share. The first plan brings the returning broker to at least 90% of its fair share of most loads. Leaders can take a second, smaller plan. A load that cannot get there is shown [at its limit](https://docs.calinora.io/features/proposals/#at-its-limit) with its cause.

## URP-Aware Proposals

Maintenance proposals account for under-replicated partitions (URPs). If moving partitions off a broker would create URPs (e.g., RF=1 topics), the proposal warns about it. When the drain cannot preserve a partition’s replication factor it is reported as described under [Force and Replication Factor](https://docs.calinora.io/features/maintenance-mode/#force-and-replication-factor) rather than evacuated silently.

## Monitoring

The number of brokers in maintenance is tracked:

```promql
pilot_cluster_brokers_maintenance
```

Maintenance state is persisted to the `__pilot_broker_state` Kafka topic, so it survives Pilot restarts.
