Skip to Content
FeaturesPartition Reassignments

Partition Reassignments

Pilot executes partition reassignments with managed throttling, progress tracking, and cancellation support. Reassignments can be triggered from proposals, the API, or self-healing.

Throttle Management

Pilot manages Kafka replication throttles to prevent reassignments from overwhelming brokers.

VariableDefaultDescription
PILOT_THROTTLE_RATE_MB50Default throttle rate (MB/s)
PILOT_THROTTLE_MAX_RATE_MB1000Highest rate the tuning API accepts (MB/s)
PILOT_THROTTLE_MANAGEDtruePilot owns replication throttles

With PILOT_THROTTLE_MANAGED=true, Pilot sets throttles for its moves and removes leftover broker throttles at startup. Set it to false when another tool manages throttles; proposals are then marked unsafe to apply (throttles_externally_managed), so self-healing skips them and a manual apply needs force=true.

Active execution cards show the Target throttle per receiving broker for that operation. The value comes from observed follower replication limits on brokers receiving new replicas for unfinished assignments. Other involved brokers are listed separately and do not affect the target summary. Source leader limits can still constrain outgoing replication; the target rate is not a cluster-wide bandwidth cap or a measured transfer rate.

Different target limits are shown as Mixed limits. Missing broker configuration is shown as Unavailable. Operations without new replica destinations, such as leadership-only changes, show Not applicable.

Concurrency Control

PILOT_MOVE_MAX_PER_BROKER=20 # Max concurrent moves per broker

Partitions are queued and submitted in batches, respecting the per-broker concurrency limit. The reassignment executor tracks active moves per broker and submits new partitions as capacity becomes available.

Progress Tracking

Open Execution to inspect active operations and their history. The reassignment monitor periodically checks the status of all tracked assignments:

  • Active - partitions currently being moved by Kafka
  • Pending - partitions queued, waiting for broker capacity
  • Completed - moves finished successfully
  • Failed - moves that encountered errors
  • Cancelled - assignments recorded as cancelled

Progress counts completed assignments out of the submitted targets. It does not measure bytes copied or provide a transfer ETA. Queued assignments remain unfinished work even when Kafka has no active transfers.

A tracked move that is still running 12 hours after it was last submitted to Kafka, and is no longer driven by an active execution, is cancelled back to its original replicas and shown as timed out. The cancel is recorded in the audit log.

Cancellation

Use Cancel remaining operation or the API to request cancellation. Kafka cancellation may remain unconfirmed, and completed assignments remain changed. Review the final replica placement and both active and historical outcomes after cancelling. Cancellation does not remove broker maintenance or exclusion settings.

Topic Redistribution Preview

The topic redistribution dialog reviews partition targets before applying them. API clients use POST /api/v1/topics/{topic}/redistribute with "preview": true for the same planning step.

A topic preview uses cached cluster metadata no older than 60 seconds. Missing, incomplete, or older metadata returns 503 Service Unavailable; retry after the next metadata refresh. Selected brokers that are unavailable or in maintenance are rejected. Applying with "preview": false rechecks availability and safety against current state, so a preview is not a guarantee that the change can still be submitted.

Apply-Time Safety Gating

Every mutation endpoint that moves partitions runs a live cluster-health pre-flight before submitting to Kafka. This applies to applying a proposal, topic and single-partition redistribution, bulk reassignment, broker maintenance, and audit-log revert.

Hard-health blockers (not overridable)

The following return 409 Conflict and cannot be bypassed with force=true:

  • Controller unhealthy
  • A majority of brokers unavailable
  • Sustained ISR shrink

Soft reasons (overridable, audited)

Reasons such as a stale proposal topology fingerprint or reassignments already in flight return 409 Conflict with guidance to re-submit with ?force=true. A forced override is recorded in the audit log.

Rolling-restart exclusion

Reassignments are blocked while a rolling restart is in progress, and broker drains are blocked for the same reason. Partition moves and rolling restarts are mutually exclusive.

Operator-supplied replica sets

When you supply an explicit replica set (single-partition or bulk reassignment), Pilot validates it against the live cluster before submitting. A set is rejected when it is empty, contains duplicate or negative broker IDs, or would fall below the replication-factor floor for the partition (the larger of the original RF and min.insync.replicas + 1). The single-partition endpoint returns 400/409; the bulk endpoint reports the entry in its results (see Bulk Reassignment). The replication factor is preserved, never silently grown. Unhealthy brokers in a proposal-driven set are dropped and substituted back up to the original RF where possible; if a safe set cannot be formed, the partition is skipped rather than submitted with a shrunk assignment.

Bulk Reassignment

POST /api/v1/topics/{topic}/partitions/bulk moves several partitions of one topic as one operation. Each entry takes the same fields as the single-partition endpoint, plus the partition number:

{ "description": "Move orders partitions off broker 3", "partitions": [ { "partition": 0, "replicas": [1, 2, 4] }, { "partition": 1, "replicas": [2, 4, 1] } ] }

Pilot validates and plans each entry on its own, then executes all planned live moves as one monitored operation with strategy BULK_MANUAL, through the same throttled path as topic and single-partition redistribution. The response returns once execution has started. Follow the operation under Execution or with GET /api/v1/reassignments/{id}, and cancel it with POST /api/v1/reassignments/{id}/cancel.

Response fieldMeaning
resultsOne result per request entry, in request order: partition, success, preview, the planned changes, and error when the entry was rejected
startedWhether an operation started. Present when the request has a live entry
bulkIdOperation ID for the reassignment endpoints. Present when started is true
messageOutcome summary. Present when the request has a live entry

The top-level success is always true on 200; check started and each result.

  • Rejected entries. An entry is reported with success: false and an error, and is never executed, when its replica set is empty, has duplicate or negative broker IDs, or would fall below the replication-factor floor, when its partition is listed more than once (every entry for that partition), or when it cannot be planned, for example because its current placement cannot be read. The remaining live entries still run. If every live entry is rejected, nothing starts and started is false.
  • Live requests. A request with at least one live entry runs the hard-health pre-flight and holds the reassignment slot until the operation ends, so another mutating reassignment gets 409 meanwhile. It is audited as reassignments.apply, including the entries that were not executed.
  • Previews. A request whose entries are all previews plans from cached metadata and starts nothing. It takes no reassignment slot, writes no audit record and sends nothing to Kafka. It still requires a license.
  • An empty partitions list returns 400.

Stale Placement Protection

Every move records the replicas it starts from. Pilot checks that source against Kafka so that a plan made for an older placement is not applied to a newer one.

  • Planning. A live single-partition or bulk redistribution refreshes the topic’s metadata and plans from the partition’s current replicas and leader. If the placement cannot be read, the single-partition endpoint returns 503 Service Unavailable and changes nothing; a bulk entry is rejected. Previews plan from cached metadata without a refresh, so a preview’s oldReplicas can be out of date.
  • Submission. Before each submission wave, the executor refreshes the affected topics once and compares each partition’s live replicas with its planned source and target. A plan is refused when the live replicas include a broker in neither, or lack a broker both keep. The partition is marked Failed with stale plan: live replicas [...] no longer match the planned source [...] and is never sent to Kafka; plan it again from the current placement. A partition already at its target, or partway through its own move, is submitted normally. If the refresh or a read fails, the plan is submitted without this check. The check covers every execution: proposals, self-healing, maintenance drains, redistribution, MCP tools and audit-log revert.
  • Completion. Pilot’s periodic metadata snapshot can trail Kafka. Before failing a move as a stale proposal, Pilot rereads the partition from refreshed metadata. A partition at its target or on its planned path keeps waiting and completes once the snapshot catches up.

API Endpoints

MethodPathDescriptionLicense
POST/api/v1/topics/{topic}/redistributeRedistribute all partitions of a topicYes
POST/api/v1/topics/{topic}/partitions/{partition}/redistributeReassign a single partitionYes
POST/api/v1/topics/{topic}/partitions/bulkReassign several partitions of a topic as one operationYes
POST/api/v1/preferred-leader-electionCluster-wide preferred leader electionYes
POST/api/v1/topics/{topic}/preferred-leader-electionTopic-specific preferred leader electionYes

See the API reference for request and response schemas, including the force parameter and 409 error bodies.

Prometheus Metrics

Key reassignment metrics:

MetricDescription
pilot_reassignment_activeCurrently active moves
pilot_reassignment_pendingQueued moves
pilot_reassignment_completed_totalCumulative completions
pilot_reassignment_duration_secondsTime per partition move
pilot_reassignment_broker_active_movesActive moves per broker

See Prometheus Metrics for the full reference.

Last updated on