---
title: "Troubleshooting - Pilot Docs"
description: "Common issues and solutions"
url: "https://docs.calinora.io/guides/troubleshooting/"
---

# Troubleshooting

## Connection Issues

### Pilot cannot connect to Kafka

**Symptoms:** Readiness probe fails, no cluster data in UI, connection errors in logs.

**Solutions:**

1. Verify `KAFKA_BOOTSTRAP_SERVERS` matches the broker’s **advertised listener** addresses
2. Ensure Pilot can resolve broker hostnames (Docker networking, DNS)
3. Check the security protocol matches your broker configuration:
   - `PLAINTEXT` - no auth, no encryption
   - `SSL` - TLS encryption
   - `SASL_PLAINTEXT` - SASL auth, no encryption
   - `SASL_SSL` - SASL auth + TLS encryption
4. For SASL, verify credentials: `KAFKA_SASL_USERNAME`, `KAFKA_SASL_PASSWORD`, and `KAFKA_SASL_MECHANISM`
5. For TLS, verify certificate paths exist and are readable (check volume mounts in Docker)

### TLS certificate errors

**Symptoms:** `x509: certificate signed by unknown authority` or `certificate verify failed`.

**Solutions:**

1. Provide the CA certificate: `KAFKA_SSL_CA_CERT_FILE=/path/to/ca.pem`
2. If using self-signed certificates for testing: `KAFKA_SSL_INSECURE_SKIP_VERIFY=true` (not for production)
3. If broker certificates are CA-valid but don’t match the hostnames Pilot connects to: `KAFKA_SSL_VERIFY_HOSTNAME=false` skips only the hostname match. The certificate chain is still verified, so this does not fix `unknown authority` errors - provide the CA instead
4. With `KAFKA_SASL_MECHANISM=OAUTHBEARER`, the same error can come from the token endpoint rather than a broker. The token endpoint trusts the system roots plus `KAFKA_SSL_CA_CERT_FILE` / `KAFKA_SSL_CA_CERT_PEM`, so add the endpoint’s CA to that file (a PEM bundle may hold several CAs). Alternatively, point the process-wide `SSL_CERT_FILE` at a full CA bundle that includes the corporate CA; note that this replaces the system roots for every HTTPS client in the process, not just the token endpoint

## Proposal Issues

### No proposal generated

**Symptoms:** Proposal generation returns empty or “no action needed”.

**Possible causes:**

1. **Cluster is balanced** - every broker is within 10% of its fair share of every load, twice `PILOT_BALANCE_THRESHOLD` (default: 5). Lower the threshold to balance smaller differences.
2. **All partitions excluded** - check `PILOT_EXCLUDE_TOPICS` isn’t too broad
3. **Measurement history** - after startup, Pilot collects about five minutes of measurement history before evaluating a proposal; the status reads **Collecting measurements** with the remaining time. See [Measurement Readiness](https://docs.calinora.io/features/proposals/#measurement-readiness-and-sustained-benefit).
4. **Recent moves** - partitions Pilot moved in the last 30 minutes are left alone, so a plan can be empty with the load at its limit (cause **recent moves**). See [Recent Moves](https://docs.calinora.io/features/proposals/#recent-moves).
5. **No traffic** - a rate or disk load whose cluster averages below 1 msg/s, 1 KiB/s or 1 MiB per broker reads **No traffic** and is not balanced, however uneven it is in percentage terms. See [No plan although imbalance is shown](https://docs.calinora.io/guides/troubleshooting/#no-plan-although-imbalance-is-shown).
6. **Balancing paused** - while a broker is down and not in maintenance, Pilot plans only repairs. Drain a broker that stays down with [maintenance mode](https://docs.calinora.io/features/maintenance-mode/). See [Unavailable Brokers](https://docs.calinora.io/features/proposals/#unavailable-brokers).

### Proposal has too many moves

**Solutions:**

1. Increase `PILOT_BALANCE_THRESHOLD` to tolerate larger differences: with `10`, Pilot acts beyond 20% from a broker’s fair share and evens to within 10%
2. Reduce `PILOT_MOVE_MAX_PER_BROKER` to limit concurrency
3. Set `PILOT_BALANCE_WORKERS` to a positive value if you want to cap CPU usage during candidate generation (does not change proposal output)
4. On a dev or test cluster, see [Frequent plans on a quiet dev or test cluster](https://docs.calinora.io/guides/troubleshooting/#frequent-plans-on-a-quiet-dev-or-test-cluster). For large disk copies, see [Large copy recommended for uneven disk](https://docs.calinora.io/guides/troubleshooting/#large-copy-recommended-for-uneven-disk)

### Frequent plans on a quiet dev or test cluster

**Symptoms:** A cluster with little traffic keeps receiving plans that move partitions for differences of a few KB/s or MiB per broker.

**Cause:** Pilot balances differences of any size by default: a broker more than 10% from its fair share counts whatever the amount, so a small cluster is evened out like a large one. Only loads with [no traffic](https://docs.calinora.io/features/proposals/#no-traffic) are left alone.

**Solution:** Set a [minimum difference](https://docs.calinora.io/features/proposals/#minimum-difference) to keep the cluster quiet:

```bash
PILOT_BALANCE_FLOOR="bytes=2MB/s disk=2GiB"
```

A rate or disk load is then balanced only when some broker is at least that far from its equal share. Message rates follow the byte rate at the cluster’s average message size. Leader and follower counts are still evened out; raise `PILOT_BALANCE_THRESHOLD` to tolerate larger count differences.

### No plan although imbalance is shown

**Symptoms:** The Balance Summary shows a red or amber row, a broker more than 10% from its fair share, but the status is **No rebalance needed**, **Monitoring**, **Balanced, at its limit** or **Balancing paused** and no moves are proposed.

**Possible causes:**

1. **Grey rows** - a row reading **No traffic** belongs to a load with essentially no traffic, which is not balanced. With `PILOT_BALANCE_FLOOR` set, a row reading **Below minimum difference** is left alone too; the status message then counts these loads, for example “Two metrics are uneven but below the minimum difference (PILOT_BALANCE_FLOOR)”.
2. **Not confirmed yet** - Pilot recommends a plan only when the same broker stays more than 10% from its fair share of the same load in four readings, and reads rates and partition sizes averaged over about 30 minutes. When the broker came back, the message reads **No broker stayed more than 10% from its fair share.** A plan being confirmed shows progress such as **2 of 3 fresh measurements collected**.
3. **No plan worth applying** - the message ends with **No plan improves balance enough to be worth applying.** No plan brings the broker within 10%, or halfway to 5%, and at least a quarter of a partition closer, while leaving the cluster more even overall. **A broker’s producer messages come less than a quarter partition closer** means the plan would move the broker by less than a quarter of one partition; the load named varies. **The plan would leave the cluster less even overall** means the plan helps one broker but leaves the cluster as a whole less even. Pilot keeps checking.
4. **At its limit** - an amber row names the cause. See [Balanced, at its limit](https://docs.calinora.io/guides/troubleshooting/#balanced-at-its-limit).
5. **Balancing paused** - the message reads **Balancing is paused while broker 5 is down. If it stays down, drain it with maintenance mode.** While a broker is down and not in maintenance, Pilot plans only repairs. See [Unavailable Brokers](https://docs.calinora.io/features/proposals/#unavailable-brokers).

**Solutions:**

1. Select a row to see the farthest broker’s load, its fair share and the size of one partition.
2. If `PILOT_BALANCE_FLOOR` is set and you want smaller differences balanced, lower it, for example `PILOT_BALANCE_FLOOR="bytes=500KB/s"`, or unset it. See [Minimum Difference](https://docs.calinora.io/features/proposals/#minimum-difference).
3. Drain a broker that stays down with [maintenance mode](https://docs.calinora.io/features/maintenance-mode/), so Pilot balances the remaining brokers.

### Balanced, at its limit

**Symptoms:** The status reads **Balanced, at its limit**, and an amber row names a broker more than 10% from its fair share and a cause, for example **At its limit: messages vs bytes**.

**Cause:** No plan worth applying brings the broker closer. Pilot keeps balancing the other loads and checks again in every calculation.

**Solutions** by cause, as the row labels it:

| Cause | What would help |
| - | - |
| large partition | Add partitions to the topics holding the large partitions so their load can be split |
| rack layout | Add brokers to the smaller racks, or use a replication factor the racks can spread evenly |
| messages vs bytes, on producer or consumer messages | Accept the residual imbalance, or split the topics whose messages are much smaller than the average |
| messages vs bytes, on producer or consumer bytes | Accept the residual imbalance, or split the topics whose messages are much larger than the average |
| messages vs bytes, on leaders or followers: “Evening them would push producer messages on broker 4 more than 10% from their fair share.” | Accept the residual imbalance |
| no move helps | Add brokers or partitions, or accept the residual imbalance |
| recent moves | Wait for the cooldown and the in-flight reassignments to pass; Pilot checks again |

The UI shows this text as **What would help**, and the API returns it in `structuralLimits[].remediation`. See [At Its Limit](https://docs.calinora.io/features/proposals/#at-its-limit).

### Large copy recommended for uneven disk

**Symptoms:** A proposal copies a large amount of data (`movement.estimatedBytes` in the API) to even out disk, although traffic is low.

**Cause:** Disk is balanced like every other load. When a broker stays more than 10% from its fair share of disk, Pilot evens that load until every broker is within 5%, which copies replicas: on a large cluster, tens or hundreds of GiB per broker. A plan is not refused for the data it copies, and filling a new, empty broker copies most of its fair share of disk. See [Even Balance](https://docs.calinora.io/features/proposals/#even-balance).

**Solutions:**

1. Review the replica moves before applying. The copy runs throttled (`PILOT_THROTTLE_RATE_MB`) and can be cancelled.
2. Split the work into smaller plans with `PILOT_PROPOSAL_MAX_PARTITIONS`, or `PILOT_HEAL_MAX_PARTITIONS_PER_RUN` for self-healing.
3. Leave smaller disk differences alone with a minimum difference, for example `PILOT_BALANCE_FLOOR="disk=100GiB"`. This also ignores traffic differences below 2 MB/s per broker, the default for `bytes=`. See [Minimum Difference](https://docs.calinora.io/features/proposals/#minimum-difference).

## Reassignment Issues

### Reassignment stuck

**Symptoms:** `pilot_reassignment_active > 0` for extended period.

**Possible causes:**

1. **Throttle too low** - increase `PILOT_THROTTLE_RATE_MB` (default: 50 MB/s)
2. **Broker unavailable** - check if the destination broker is reachable
3. **ISR shrink** - the safety signal phase may block new moves when ISR is shrinking
4. **Large partitions** - very large partitions take longer to move

**Solutions:**

1. Check reassignment status: `curl http://localhost:8080/api/v1/reassignments`
2. Increase throttle: update via API or set `PILOT_THROTTLE_RATE_MB` higher
3. Cancel a stuck reassignment if needed: `POST /api/v1/reassignments/{id}/cancel`

## Self-Healing Issues

### Self-healing not running

**Symptoms:** `pilot_selfhealing_skipped_total` increasing.

**Check skip reasons:**

```promql
pilot_selfhealing_skipped_total
```

| Skip Reason | Solution |
| - | - |
| `outside_window` | Current time is outside `PILOT_HEAL_WINDOW_START` / `PILOT_HEAL_WINDOW_END` |
| `cooldown` | A loop applied within the last 30 minutes; wait |
| `partitions_in_cooldown` | A partition Pilot moved is still in its 30-minute cooldown; wait, so the next plan can move every partition |
| `dry_run` | Set `PILOT_HEAL_DRY_RUN=false` to apply |
| `active_reassignments` | Wait for current reassignments to complete |
| `license_invalid` | Provide a valid license |
| `data_not_ready` | Measurements are collecting or settling after moves |
| `benefit_unconfirmed` | The plan has not yet shown sustained benefit on three fresh measurements |
| `no_proposal` | Cluster is balanced or insufficient data |
| `higher_priority_running` | Wait for higher-priority healing loop to finish |

### Healing in dry-run mode

If `PILOT_HEAL_DRY_RUN=true` (default), healing generates proposals but does not apply them. Set to `false` to enable actual healing.

## UI Issues

### Dashboard not loading

**Possible causes:**

1. **UI not embedded** - binary may have been built without `uiassets` tag. Rebuild with `make build` or `make build-full`
2. **Wrong port** - verify `PORT` environment variable (default: 8080)
3. **Proxy issues** - if behind a reverse proxy, check proxy configuration. See [Ingress](https://docs.calinora.io/deployment/kubernetes/#ingress)

### SSE streaming not working (chat)

If AI chat responses don’t stream, check reverse proxy configuration:

- Disable response buffering (`proxy_buffering off` in Nginx)
- Set long read timeouts (3600s+)
- Disable caching for SSE endpoints

## License Issues

### 403 Forbidden on mutation endpoints

**Symptoms:** API returns `403` with `{"data": null, "error": {"code": "LICENSE_INVALID", "message": "License is invalid"}}` (or `LICENSE_EXPIRED`).

**Solutions:**

1. Check license status: `curl http://localhost:8080/api/v1/license | jq '.'`
2. If expired, renew or configure auto-fetch with `LICENSE_FETCH_SUBSCRIPTION_ID` and `LICENSE_FETCH_TOKEN`
3. Read-only operations (GET, proposal generation, what-if) work without a license

## Logs

Set `LOG_LEVEL=DEBUG` for verbose output:

```bash
docker run -e LOG_LEVEL=DEBUG ...
```

Log levels: `DEBUG`, `INFO`, `WARN`.
