---
title: "Self-Healing - Pilot Docs"
description: "Automatic partition repair and rebalancing"
url: "https://docs.calinora.io/features/self-healing/"
---

# Self-Healing

Pilot can automatically detect and repair cluster issues without manual intervention. Three independent healing loops run in the background, each targeting a different class of problem.

## Healing Loops

### Activity-Based Healing

Rebalances partitions based on throughput and activity to maintain even load distribution across brokers.

Activity healing uses the same measurement readiness, sustained-benefit and freshness checks as default manual application. The exact final plan must stay worth applying across three fresh, compatible measurements after pruning and any partition limit. Pilot rechecks that plan against current measurements before submission. See [Reducing Unnecessary Movement](https://docs.calinora.io/features/proposals/#reducing-unnecessary-movement) and [Recent Moves](https://docs.calinora.io/features/proposals/#recent-moves).

A skipped cycle with `data_not_ready` is waiting for measurement history, freshness or settling. `benefit_unconfirmed` means the plan’s sustained benefit has not been established. Finishing the settling period allows reevaluation; automatic application still requires the healing loop to run and pass its safety, time-window, cooldown, dry-run and license checks.

Settling after Pilot’s own moves takes an adaptive 5 to 30 minutes by default. With a network [broker profile](https://docs.calinora.io/features/proposals/#broker-profiles), it ends on evidence instead: at least two minutes after the last move, once moved replicas are in sync, leadership is as planned and fresh measurements of the moved partitions exist, and at most 30 minutes after the last move. See [Settling After Moves](https://docs.calinora.io/features/proposals/#settling-after-moves). In both modes, each healing loop skips cycles for 30 minutes after its last application (skip reason `cooldown`), and partitions Pilot moved stay out of balancing plans for 30 minutes, except partitions moved off a broker entering [maintenance](https://docs.calinora.io/features/maintenance-mode/#refill-after-maintenance). Activity healing waits while any partition is still in that cooldown (skip reason `partitions_in_cooldown`); repairs never wait. An activity cycle skipped for a reason that ends soon runs again, at most three times before the next regular cycle: 5 seconds after a cooldown ends (the partitions’ or the loop’s own), two minutes later when its own new plan is not confirmed yet, or one minute later while the critical or RF loop runs. A retry is a normal cycle with every check, and the regular interval does not move. The partition cooldown is kept in memory, so a Pilot restart clears it. See [Recent Moves](https://docs.calinora.io/features/proposals/#recent-moves).

```bash
PILOT_HEAL_ENABLED=false     # Enable activity rebalancing
PILOT_HEAL_INTERVAL=30m      # Check interval
```

### Critical Fixes

Repairs under-replicated partitions (URPs) and rack-awareness violations. This is the most urgent healing loop.

```bash
PILOT_HEAL_CRITICAL_ENABLED=false  # Enable critical fixes
PILOT_HEAL_CRITICAL_INTERVAL=5m    # Check interval
```

### Replication Factor Increases

Automatically increases replication factor for partitions below the configured target.

```bash
PILOT_HEAL_RF_ENABLED=false   # Enable RF increases
PILOT_HEAL_RF_INTERVAL=15m    # Check interval
```

See [Replication Factor](https://docs.calinora.io/features/replication-factor/) for details on the target RF feature.

## Common Settings

These settings apply to all three healing loops:

| Variable | Default | Description |
| - | - | - |
| `PILOT_HEAL_DRY_RUN` | `true` | Simulate only - log what would be done without applying changes |
| `PILOT_HEAL_MAX_PARTITIONS_PER_RUN` | `0` | Maximum partitions moved per healing cycle (`0` = unlimited, set > 0 to cap) |
| `PILOT_HEAL_WINDOW_START` | `-1` | Allowed start hour (0-23, -1 = always) |
| `PILOT_HEAL_WINDOW_END` | `-1` | Allowed end hour (0-23, -1 = always) |

For activity healing, the per-run partition limit is applied during proposal generation so the smaller plan’s benefit is evaluated before application. A limit may leave no worthwhile balance plan. Correctness repairs can be applied in batches; they do not need to improve balance. If other apply-time filtering changes an optional balance plan, healing skips that plan until recalculation.

## Time Windows

Restrict healing to specific hours of the day:

```bash
# Business hours only
PILOT_HEAL_WINDOW_START=8
PILOT_HEAL_WINDOW_END=17

# Overnight maintenance window
PILOT_HEAL_WINDOW_START=22
PILOT_HEAL_WINDOW_END=5

# Always (default)
PILOT_HEAL_WINDOW_START=-1
PILOT_HEAL_WINDOW_END=-1
```

Overnight windows are supported - setting `START=22` and `END=5` means 22:00-05:59.

## Recommended Rollout

1. **Enable with dry-run** - set `PILOT_HEAL_DRY_RUN=true` and enable the desired loops
2. **Monitor logs** - review what Pilot would have done
3. **Enable critical fixes first** - `PILOT_HEAL_CRITICAL_ENABLED=true` with `PILOT_HEAL_DRY_RUN=false`
4. **Add activity healing** - after validating critical fixes work correctly
5. **Consider time windows** - restrict healing to low-traffic periods if needed

## Priority

When multiple healing loops want to act simultaneously, Pilot uses a priority system:

1. **Critical fixes** - highest priority (URP repair, rack violations)
2. **RF increases** - medium priority
3. **Activity rebalancing** - lowest priority

A lower-priority loop is skipped if a higher-priority loop is actively running.

Generated proposals also separate repairs from optional balancing. If repairs are needed, Pilot proposes those first and reassesses balance after they complete and measurements settle.

## License Requirement

Self-healing requires a valid license to apply changes. Without one, each cycle is skipped with the `license_invalid` skip reason. In dry-run mode with a valid license, cycles are skipped with `dry_run` after logging what would be applied.

## Audit Log

Every submitted automatic apply is recorded in the [audit log](https://docs.calinora.io/configuration/audit-logging/#automatic-applies) as `selfheal.apply`, attributed to `pilot` with method `INTERNAL`. The event is written when the moves are submitted for execution, before any move is sent to Kafka, with the proposal ID, the trigger (`activity`, `critical`, `rf`, or `post-maintenance` for the [rebalance after a drain](https://docs.calinora.io/features/maintenance-mode/#automatic-rebalance-after-a-drain)) and one row per partition with old and new replicas and leaders. An apply that fails after its submission was recorded adds an event with status `500`. An apply refused before submission, such as one whose moves are all filtered out as leaderless during an outage, records nothing; it is logged and counted as `pilot_selfhealing_runs_total{result="error"}`. When the execution window closes on a running heal, the cancellation is recorded as `selfheal.cancel`. Dry runs and skipped cycles record nothing. These events cannot be reverted from the audit log.

## Prometheus Metrics

| Metric | Description |
| - | - |
| `pilot_selfhealing_runs_total` | Total executions by type and result |
| `pilot_selfhealing_skipped_total` | Skipped runs with reasons |
| `pilot_selfhealing_last_success_timestamp_seconds` | Last successful run |
| `pilot_selfhealing_reassignments_applied_total` | Partitions submitted by successful applies |

Alert when critical healing stops running while URPs exist:

```promql
(time() - pilot_selfhealing_last_success_timestamp_seconds{type="critical"} > 1800)
and pilot_cluster_under_replicated_partitions > 0
```

See [Prometheus Metrics](https://docs.calinora.io/reference/metrics/#self-healing) for the full reference.
