---
title: "Production Checklist - Pilot Docs"
description: "Checklist for running Pilot in production"
url: "https://docs.calinora.io/guides/production-checklist/"
---

# Production Checklist

Use this checklist to prepare Pilot for production deployment.

## Security

- [ ] **Kafka authentication** - configure SASL or TLS for Kafka connections ([SASL](https://docs.calinora.io/configuration/sasl/), [TLS](https://docs.calinora.io/configuration/tls/))
- [ ] **UI authentication** - enable OAuth2/OIDC to restrict dashboard access ([Authentication](https://docs.calinora.io/configuration/authentication/))
- [ ] **Cookie secret** - set `AUTH_COOKIE_SECRET` to a strong random value (do not rely on auto-generation)
- [ ] **Token verification** - set `AUTH_<PROVIDER>_ISSUER` for any provider whose bearer tokens are JWTs; without it no signature can be checked ([Bearer Token Validation](https://docs.calinora.io/configuration/authentication/#bearer-token-validation))
- [ ] **Audience binding** - confirm the audience your IdP stamps matches `AUTH_<PROVIDER>_ALLOWED_AUDIENCES`, which defaults to the client id ([Audience Binding](https://docs.calinora.io/configuration/authentication/#audience-binding))
- [ ] **HTTPS** - enable server TLS or use a reverse proxy with TLS termination ([TLS](https://docs.calinora.io/configuration/tls/), [Ingress](https://docs.calinora.io/deployment/kubernetes/#ingress))
- [ ] **License** - configure a valid license for management operations ([Licensing](https://docs.calinora.io/reference/licensing/))

## Kafka Configuration

- [ ] **Bootstrap servers** - set `KAFKA_BOOTSTRAP_SERVERS` to at least 3 brokers for redundancy
- [ ] **Exclude internal topics** - set `PILOT_EXCLUDE_TOPICS="__consumer_offsets,__transaction_state"` to avoid moving internal Kafka topics
- [ ] **Rack awareness** - ensure `PILOT_BALANCE_RACK_AWARE=true` (default) and that your brokers have rack assignments configured

## Balance & Throttling

- [ ] **Balance threshold** - review `PILOT_BALANCE_THRESHOLD` (default: 5): Pilot acts when a broker is more than twice the value from its fair share and evens to within the value (see [Even Balance](https://docs.calinora.io/features/proposals/#even-balance))
- [ ] **Minimum difference** - optional; leave `PILOT_BALANCE_FLOOR` unset to even out brokers whatever the size of the difference, or set it (for example `bytes=2MB/s disk=2GiB`) if smaller per-broker differences are not worth moving data for on your cluster (see [Minimum Difference](https://docs.calinora.io/features/proposals/#minimum-difference))
- [ ] **Broker profile** - optionally set `PILOT_BROKER_PROFILE` to each broker’s sustained network throughput and log storage to see broker utilization and to settle on evidence after moves; it does not change balancing (see [Broker Profiles](https://docs.calinora.io/features/proposals/#broker-profiles))
- [ ] **Throttle rate** - set `PILOT_THROTTLE_RATE_MB` based on your network capacity (default: 50 MB/s, increase for faster rebalancing)
- [ ] **Max throttle rate** - set `PILOT_THROTTLE_MAX_RATE_MB` as an upper bound (default: 1000 MB/s)
- [ ] **Max moves per broker** - set `PILOT_MOVE_MAX_PER_BROKER` based on broker capacity (default: 20)

## Self-Healing

- [ ] **Start with dry-run** - enable self-healing with `PILOT_HEAL_DRY_RUN=true` first
- [ ] **Enable critical fixes** - set `PILOT_HEAL_CRITICAL_ENABLED=true` after validating dry-run behavior
- [ ] **Consider time windows** - restrict healing to maintenance windows if needed
- [ ] **Consider a partition cap** - healing cycles move unlimited partitions by default; set `PILOT_HEAL_MAX_PARTITIONS_PER_RUN` to a value > 0 to cap moves per cycle on large clusters

## Monitoring

- [ ] **Prometheus scraping** - configure Prometheus to scrape `/metrics` ([Monitoring Setup](https://docs.calinora.io/guides/monitoring-setup/))
- [ ] **Alerts** - set up alerts for under-replicated partitions, offline partitions, license expiry, and sampler staleness
- [ ] **Health checks** - configure liveness (`/api/v1/health`) and readiness (`/api/v1/ready`) probes

## Deployment

- [ ] **Single replica** - run exactly one Pilot instance per cluster (multiple replicas cause conflicts)
- [ ] **Resource limits** - set appropriate CPU and memory limits based on cluster size
- [ ] **Restart policy** - configure automatic restart on failure (Docker restart policy or Kubernetes liveness probe)
- [ ] **Network access** - ensure Pilot can reach all Kafka broker advertised listeners
- [ ] **License fetch** - if using auto-fetch, ensure outbound HTTPS access to `license.calinora.io`

## Audit & Compliance

- [ ] **Audit logging** - enable authentication so audit events record who made each change; without it, events have an empty user ([Audit Logging](https://docs.calinora.io/configuration/audit-logging/))
- [ ] **Audit replication** - check that `__pilot_audit_log` has more than one replica; Pilot warns at startup about an existing single-replica audit topic but does not change it
- [ ] **Audit delivery** - alert on `pilot_audit_events_failure_total{error_type="dropped"}`, which counts audit events lost for good ([Delivery](https://docs.calinora.io/configuration/audit-logging/#delivery))
- [ ] **Review audit events** - periodically review the audit log for unexpected operations

## Backup & Recovery

- [ ] **Understand state model** - Pilot stores state in internal Kafka topics (see [Architecture](https://docs.calinora.io/features/architecture/#state-persistence) for details)
- [ ] **Broker state** - ensure `__pilot_broker_state` uses compaction (created and enforced automatically by Pilot)
- [ ] **Audit log** - ensure `__pilot_audit_log` retention meets your compliance needs (default 30 days, configurable via `AUDIT_RETENTION_MS`)
- [ ] **PAT tokens** - if PATs are enabled, `__pilot_pat_tokens` stores hashed tokens with compaction (created automatically)
- [ ] **Proposals** - proposals are stored in memory and regenerated on demand; they are not backed up
