Skip to Content
FeaturesSelf-Healing

Self-Healing

Pilot can automatically detect and repair cluster issues without manual intervention. Three independent healing loops run in the background, each targeting a different class of problem.

Healing Loops

Activity-Based Healing

Rebalances partitions based on throughput and activity to maintain even load distribution across brokers.

Activity healing uses the same measurement readiness, sustained-benefit and freshness checks as default manual application. The exact final plan must stay worth applying across three fresh, compatible measurements after pruning and any partition limit. Pilot rechecks that plan against current measurements before submission. See Reducing Unnecessary Movement and Recent Moves.

A skipped cycle with data_not_ready is waiting for measurement history, freshness or settling. benefit_unconfirmed means the plan’s sustained benefit has not been established. Finishing the settling period allows reevaluation; automatic application still requires the healing loop to run and pass its safety, time-window, cooldown, dry-run and license checks.

Settling after Pilot’s own moves takes an adaptive 5 to 30 minutes by default. With a network broker profile, it ends on evidence instead: at least two minutes after the last move, once moved replicas are in sync, leadership is as planned and fresh measurements of the moved partitions exist, and at most 30 minutes after the last move. See Settling After Moves. In both modes, each healing loop skips cycles for 30 minutes after its last application (skip reason cooldown), and partitions Pilot moved stay out of balancing plans for 30 minutes, except partitions moved off a broker entering maintenance. Activity healing waits while any partition is still in that cooldown (skip reason partitions_in_cooldown); repairs never wait. An activity cycle skipped for a reason that ends soon runs again, at most three times before the next regular cycle: 5 seconds after a cooldown ends (the partitions’ or the loop’s own), two minutes later when its own new plan is not confirmed yet, or one minute later while the critical or RF loop runs. A retry is a normal cycle with every check, and the regular interval does not move. The partition cooldown is kept in memory, so a Pilot restart clears it. See Recent Moves.

PILOT_HEAL_ENABLED=false # Enable activity rebalancing PILOT_HEAL_INTERVAL=30m # Check interval

Critical Fixes

Repairs under-replicated partitions (URPs) and rack-awareness violations. This is the most urgent healing loop.

PILOT_HEAL_CRITICAL_ENABLED=false # Enable critical fixes PILOT_HEAL_CRITICAL_INTERVAL=5m # Check interval

Replication Factor Increases

Automatically increases replication factor for partitions below the configured target.

PILOT_HEAL_RF_ENABLED=false # Enable RF increases PILOT_HEAL_RF_INTERVAL=15m # Check interval

See Replication Factor for details on the target RF feature.

Common Settings

These settings apply to all three healing loops:

VariableDefaultDescription
PILOT_HEAL_DRY_RUNtrueSimulate only - log what would be done without applying changes
PILOT_HEAL_MAX_PARTITIONS_PER_RUN0Maximum partitions moved per healing cycle (0 = unlimited, set > 0 to cap)
PILOT_HEAL_WINDOW_START-1Allowed start hour (0-23, -1 = always)
PILOT_HEAL_WINDOW_END-1Allowed end hour (0-23, -1 = always)

For activity healing, the per-run partition limit is applied during proposal generation so the smaller plan’s benefit is evaluated before application. A limit may leave no worthwhile balance plan. Correctness repairs can be applied in batches; they do not need to improve balance. If other apply-time filtering changes an optional balance plan, healing skips that plan until recalculation.

Time Windows

Restrict healing to specific hours of the day:

# Business hours only PILOT_HEAL_WINDOW_START=8 PILOT_HEAL_WINDOW_END=17 # Overnight maintenance window PILOT_HEAL_WINDOW_START=22 PILOT_HEAL_WINDOW_END=5 # Always (default) PILOT_HEAL_WINDOW_START=-1 PILOT_HEAL_WINDOW_END=-1

Overnight windows are supported - setting START=22 and END=5 means 22:00-05:59.

  1. Enable with dry-run - set PILOT_HEAL_DRY_RUN=true and enable the desired loops
  2. Monitor logs - review what Pilot would have done
  3. Enable critical fixes first - PILOT_HEAL_CRITICAL_ENABLED=true with PILOT_HEAL_DRY_RUN=false
  4. Add activity healing - after validating critical fixes work correctly
  5. Consider time windows - restrict healing to low-traffic periods if needed

Priority

When multiple healing loops want to act simultaneously, Pilot uses a priority system:

  1. Critical fixes - highest priority (URP repair, rack violations)
  2. RF increases - medium priority
  3. Activity rebalancing - lowest priority

A lower-priority loop is skipped if a higher-priority loop is actively running.

Generated proposals also separate repairs from optional balancing. If repairs are needed, Pilot proposes those first and reassesses balance after they complete and measurements settle.

License Requirement

Self-healing requires a valid license to apply changes. Without one, each cycle is skipped with the license_invalid skip reason. In dry-run mode with a valid license, cycles are skipped with dry_run after logging what would be applied.

Audit Log

Every submitted automatic apply is recorded in the audit log as selfheal.apply, attributed to pilot with method INTERNAL. The event is written when the moves are submitted for execution, before any move is sent to Kafka, with the proposal ID, the trigger (activity, critical, rf, or post-maintenance for the rebalance after a drain) and one row per partition with old and new replicas and leaders. An apply that fails after its submission was recorded adds an event with status 500. An apply refused before submission, such as one whose moves are all filtered out as leaderless during an outage, records nothing; it is logged and counted as pilot_selfhealing_runs_total{result="error"}. When the execution window closes on a running heal, the cancellation is recorded as selfheal.cancel. Dry runs and skipped cycles record nothing. These events cannot be reverted from the audit log.

Prometheus Metrics

MetricDescription
pilot_selfhealing_runs_totalTotal executions by type and result
pilot_selfhealing_skipped_totalSkipped runs with reasons
pilot_selfhealing_last_success_timestamp_secondsLast successful run
pilot_selfhealing_reassignments_applied_totalPartitions submitted by successful applies

Alert when critical healing stops running while URPs exist:

(time() - pilot_selfhealing_last_success_timestamp_seconds{type="critical"} > 1800) and pilot_cluster_under_replicated_partitions > 0

See Prometheus Metrics for the full reference.

Last updated on