Troubleshooting
Connection Issues
Pilot cannot connect to Kafka
Symptoms: Readiness probe fails, no cluster data in UI, connection errors in logs.
Solutions:
- Verify
KAFKA_BOOTSTRAP_SERVERSmatches the broker’s advertised listener addresses - Ensure Pilot can resolve broker hostnames (Docker networking, DNS)
- Check the security protocol matches your broker configuration:
PLAINTEXT- no auth, no encryptionSSL- TLS encryptionSASL_PLAINTEXT- SASL auth, no encryptionSASL_SSL- SASL auth + TLS encryption
- For SASL, verify credentials:
KAFKA_SASL_USERNAME,KAFKA_SASL_PASSWORD, andKAFKA_SASL_MECHANISM - For TLS, verify certificate paths exist and are readable (check volume mounts in Docker)
TLS certificate errors
Symptoms: x509: certificate signed by unknown authority or certificate verify failed.
Solutions:
- Provide the CA certificate:
KAFKA_SSL_CA_CERT_FILE=/path/to/ca.pem - If using self-signed certificates for testing:
KAFKA_SSL_INSECURE_SKIP_VERIFY=true(not for production) - If broker certificates are CA-valid but don’t match the hostnames Pilot connects to:
KAFKA_SSL_VERIFY_HOSTNAME=falseskips only the hostname match. The certificate chain is still verified, so this does not fixunknown authorityerrors - provide the CA instead - With
KAFKA_SASL_MECHANISM=OAUTHBEARER, the same error can come from the token endpoint rather than a broker. The token endpoint trusts the system roots plusKAFKA_SSL_CA_CERT_FILE/KAFKA_SSL_CA_CERT_PEM, so add the endpoint’s CA to that file (a PEM bundle may hold several CAs). Alternatively, point the process-wideSSL_CERT_FILEat a full CA bundle that includes the corporate CA; note that this replaces the system roots for every HTTPS client in the process, not just the token endpoint
Proposal Issues
No proposal generated
Symptoms: Proposal generation returns empty or “no action needed”.
Possible causes:
- Cluster is balanced - every broker is within 10% of its fair share of every load, twice
PILOT_BALANCE_THRESHOLD(default: 5). Lower the threshold to balance smaller differences. - All partitions excluded - check
PILOT_EXCLUDE_TOPICSisn’t too broad - Measurement history - after startup, Pilot collects about five minutes of measurement history before evaluating a proposal; the status reads Collecting measurements with the remaining time. See Measurement Readiness.
- Recent moves - partitions Pilot moved in the last 30 minutes are left alone, so a plan can be empty with the load at its limit (cause recent moves). See Recent Moves.
- No traffic - a rate or disk load whose cluster averages below 1 msg/s, 1 KiB/s or 1 MiB per broker reads No traffic and is not balanced, however uneven it is in percentage terms. See No plan although imbalance is shown.
- Balancing paused - while a broker is down and not in maintenance, Pilot plans only repairs. Drain a broker that stays down with maintenance mode. See Unavailable Brokers.
Proposal has too many moves
Solutions:
- Increase
PILOT_BALANCE_THRESHOLDto tolerate larger differences: with10, Pilot acts beyond 20% from a broker’s fair share and evens to within 10% - Reduce
PILOT_MOVE_MAX_PER_BROKERto limit concurrency - Set
PILOT_BALANCE_WORKERSto a positive value if you want to cap CPU usage during candidate generation (does not change proposal output) - On a dev or test cluster, see Frequent plans on a quiet dev or test cluster. For large disk copies, see Large copy recommended for uneven disk
Frequent plans on a quiet dev or test cluster
Symptoms: A cluster with little traffic keeps receiving plans that move partitions for differences of a few KB/s or MiB per broker.
Cause: Pilot balances differences of any size by default: a broker more than 10% from its fair share counts whatever the amount, so a small cluster is evened out like a large one. Only loads with no traffic are left alone.
Solution: Set a minimum difference to keep the cluster quiet:
PILOT_BALANCE_FLOOR="bytes=2MB/s disk=2GiB"A rate or disk load is then balanced only when some broker is at least that far from its equal share. Message rates follow the byte rate at the cluster’s average message size. Leader and follower counts are still evened out; raise PILOT_BALANCE_THRESHOLD to tolerate larger count differences.
No plan although imbalance is shown
Symptoms: The Balance Summary shows a red or amber row, a broker more than 10% from its fair share, but the status is No rebalance needed, Monitoring, Balanced, at its limit or Balancing paused and no moves are proposed.
Possible causes:
- Grey rows - a row reading No traffic belongs to a load with essentially no traffic, which is not balanced. With
PILOT_BALANCE_FLOORset, a row reading Below minimum difference is left alone too; the status message then counts these loads, for example “Two metrics are uneven but below the minimum difference (PILOT_BALANCE_FLOOR)”. - Not confirmed yet - Pilot recommends a plan only when the same broker stays more than 10% from its fair share of the same load in four readings, and reads rates and partition sizes averaged over about 30 minutes. When the broker came back, the message reads No broker stayed more than 10% from its fair share. A plan being confirmed shows progress such as 2 of 3 fresh measurements collected.
- No plan worth applying - the message ends with No plan improves balance enough to be worth applying. No plan brings the broker within 10%, or halfway to 5%, and at least a quarter of a partition closer, while leaving the cluster more even overall. A broker’s producer messages come less than a quarter partition closer means the plan would move the broker by less than a quarter of one partition; the load named varies. The plan would leave the cluster less even overall means the plan helps one broker but leaves the cluster as a whole less even. Pilot keeps checking.
- At its limit - an amber row names the cause. See Balanced, at its limit.
- Balancing paused - the message reads Balancing is paused while broker 5 is down. If it stays down, drain it with maintenance mode. While a broker is down and not in maintenance, Pilot plans only repairs. See Unavailable Brokers.
Solutions:
- Select a row to see the farthest broker’s load, its fair share and the size of one partition.
- If
PILOT_BALANCE_FLOORis set and you want smaller differences balanced, lower it, for examplePILOT_BALANCE_FLOOR="bytes=500KB/s", or unset it. See Minimum Difference. - Drain a broker that stays down with maintenance mode, so Pilot balances the remaining brokers.
Balanced, at its limit
Symptoms: The status reads Balanced, at its limit, and an amber row names a broker more than 10% from its fair share and a cause, for example At its limit: messages vs bytes.
Cause: No plan worth applying brings the broker closer. Pilot keeps balancing the other loads and checks again in every calculation.
Solutions by cause, as the row labels it:
| Cause | What would help |
|---|---|
| large partition | Add partitions to the topics holding the large partitions so their load can be split |
| rack layout | Add brokers to the smaller racks, or use a replication factor the racks can spread evenly |
| messages vs bytes, on producer or consumer messages | Accept the residual imbalance, or split the topics whose messages are much smaller than the average |
| messages vs bytes, on producer or consumer bytes | Accept the residual imbalance, or split the topics whose messages are much larger than the average |
| messages vs bytes, on leaders or followers: “Evening them would push producer messages on broker 4 more than 10% from their fair share.” | Accept the residual imbalance |
| no move helps | Add brokers or partitions, or accept the residual imbalance |
| recent moves | Wait for the cooldown and the in-flight reassignments to pass; Pilot checks again |
The UI shows this text as What would help, and the API returns it in structuralLimits[].remediation. See At Its Limit.
Large copy recommended for uneven disk
Symptoms: A proposal copies a large amount of data (movement.estimatedBytes in the API) to even out disk, although traffic is low.
Cause: Disk is balanced like every other load. When a broker stays more than 10% from its fair share of disk, Pilot evens that load until every broker is within 5%, which copies replicas: on a large cluster, tens or hundreds of GiB per broker. A plan is not refused for the data it copies, and filling a new, empty broker copies most of its fair share of disk. See Even Balance.
Solutions:
- Review the replica moves before applying. The copy runs throttled (
PILOT_THROTTLE_RATE_MB) and can be cancelled. - Split the work into smaller plans with
PILOT_PROPOSAL_MAX_PARTITIONS, orPILOT_HEAL_MAX_PARTITIONS_PER_RUNfor self-healing. - Leave smaller disk differences alone with a minimum difference, for example
PILOT_BALANCE_FLOOR="disk=100GiB". This also ignores traffic differences below 2 MB/s per broker, the default forbytes=. See Minimum Difference.
Reassignment Issues
Reassignment stuck
Symptoms: pilot_reassignment_active > 0 for extended period.
Possible causes:
- Throttle too low - increase
PILOT_THROTTLE_RATE_MB(default: 50 MB/s) - Broker unavailable - check if the destination broker is reachable
- ISR shrink - the safety signal phase may block new moves when ISR is shrinking
- Large partitions - very large partitions take longer to move
Solutions:
- Check reassignment status:
curl http://localhost:8080/api/v1/reassignments - Increase throttle: update via API or set
PILOT_THROTTLE_RATE_MBhigher - Cancel a stuck reassignment if needed:
POST /api/v1/reassignments/{id}/cancel
Self-Healing Issues
Self-healing not running
Symptoms: pilot_selfhealing_skipped_total increasing.
Check skip reasons:
pilot_selfhealing_skipped_total| Skip Reason | Solution |
|---|---|
outside_window | Current time is outside PILOT_HEAL_WINDOW_START / PILOT_HEAL_WINDOW_END |
cooldown | A loop applied within the last 30 minutes; wait |
partitions_in_cooldown | A partition Pilot moved is still in its 30-minute cooldown; wait, so the next plan can move every partition |
dry_run | Set PILOT_HEAL_DRY_RUN=false to apply |
active_reassignments | Wait for current reassignments to complete |
license_invalid | Provide a valid license |
data_not_ready | Measurements are collecting or settling after moves |
benefit_unconfirmed | The plan has not yet shown sustained benefit on three fresh measurements |
no_proposal | Cluster is balanced or insufficient data |
higher_priority_running | Wait for higher-priority healing loop to finish |
Healing in dry-run mode
If PILOT_HEAL_DRY_RUN=true (default), healing generates proposals but does not apply them. Set to false to enable actual healing.
UI Issues
Dashboard not loading
Possible causes:
- UI not embedded - binary may have been built without
uiassetstag. Rebuild withmake buildormake build-full - Wrong port - verify
PORTenvironment variable (default: 8080) - Proxy issues - if behind a reverse proxy, check proxy configuration. See Ingress
SSE streaming not working (chat)
If AI chat responses don’t stream, check reverse proxy configuration:
- Disable response buffering (
proxy_buffering offin Nginx) - Set long read timeouts (3600s+)
- Disable caching for SSE endpoints
License Issues
403 Forbidden on mutation endpoints
Symptoms: API returns 403 with {"data": null, "error": {"code": "LICENSE_INVALID", "message": "License is invalid"}} (or LICENSE_EXPIRED).
Solutions:
- Check license status:
curl http://localhost:8080/api/v1/license | jq '.' - If expired, renew or configure auto-fetch with
LICENSE_FETCH_SUBSCRIPTION_IDandLICENSE_FETCH_TOKEN - Read-only operations (GET, proposal generation, what-if) work without a license
Logs
Set LOG_LEVEL=DEBUG for verbose output:
docker run -e LOG_LEVEL=DEBUG ...Log levels: DEBUG, INFO, WARN.