Skip to Content
GuidesProduction Checklist

Production Checklist

Use this checklist to prepare Pilot for production deployment.

Security

  • Kafka authentication - configure SASL or TLS for Kafka connections (SASL, TLS)
  • UI authentication - enable OAuth2/OIDC to restrict dashboard access (Authentication)
  • Cookie secret - set AUTH_COOKIE_SECRET to a strong random value (do not rely on auto-generation)
  • Token verification - set AUTH_<PROVIDER>_ISSUER for any provider whose bearer tokens are JWTs; without it no signature can be checked (Bearer Token Validation)
  • Audience binding - confirm the audience your IdP stamps matches AUTH_<PROVIDER>_ALLOWED_AUDIENCES, which defaults to the client id (Audience Binding)
  • HTTPS - enable server TLS or use a reverse proxy with TLS termination (TLS, Ingress)
  • License - configure a valid license for management operations (Licensing)

Kafka Configuration

  • Bootstrap servers - set KAFKA_BOOTSTRAP_SERVERS to at least 3 brokers for redundancy
  • Exclude internal topics - set PILOT_EXCLUDE_TOPICS="__consumer_offsets,__transaction_state" to avoid moving internal Kafka topics
  • Rack awareness - ensure PILOT_BALANCE_RACK_AWARE=true (default) and that your brokers have rack assignments configured

Balance & Throttling

  • Balance threshold - review PILOT_BALANCE_THRESHOLD (default: 5): Pilot acts when a broker is more than twice the value from its fair share and evens to within the value (see Even Balance)
  • Minimum difference - optional; leave PILOT_BALANCE_FLOOR unset to even out brokers whatever the size of the difference, or set it (for example bytes=2MB/s disk=2GiB) if smaller per-broker differences are not worth moving data for on your cluster (see Minimum Difference)
  • Broker profile - optionally set PILOT_BROKER_PROFILE to each broker’s sustained network throughput and log storage to see broker utilization and to settle on evidence after moves; it does not change balancing (see Broker Profiles)
  • Throttle rate - set PILOT_THROTTLE_RATE_MB based on your network capacity (default: 50 MB/s, increase for faster rebalancing)
  • Max throttle rate - set PILOT_THROTTLE_MAX_RATE_MB as an upper bound (default: 1000 MB/s)
  • Max moves per broker - set PILOT_MOVE_MAX_PER_BROKER based on broker capacity (default: 20)

Self-Healing

  • Start with dry-run - enable self-healing with PILOT_HEAL_DRY_RUN=true first
  • Enable critical fixes - set PILOT_HEAL_CRITICAL_ENABLED=true after validating dry-run behavior
  • Consider time windows - restrict healing to maintenance windows if needed
  • Consider a partition cap - healing cycles move unlimited partitions by default; set PILOT_HEAL_MAX_PARTITIONS_PER_RUN to a value > 0 to cap moves per cycle on large clusters

Monitoring

  • Prometheus scraping - configure Prometheus to scrape /metrics (Monitoring Setup)
  • Alerts - set up alerts for under-replicated partitions, offline partitions, license expiry, and sampler staleness
  • Health checks - configure liveness (/api/v1/health) and readiness (/api/v1/ready) probes

Deployment

  • Single replica - run exactly one Pilot instance per cluster (multiple replicas cause conflicts)
  • Resource limits - set appropriate CPU and memory limits based on cluster size
  • Restart policy - configure automatic restart on failure (Docker restart policy or Kubernetes liveness probe)
  • Network access - ensure Pilot can reach all Kafka broker advertised listeners
  • License fetch - if using auto-fetch, ensure outbound HTTPS access to license.calinora.io

Audit & Compliance

  • Audit logging - enable authentication so audit events record who made each change; without it, events have an empty user (Audit Logging)
  • Audit replication - check that __pilot_audit_log has more than one replica; Pilot warns at startup about an existing single-replica audit topic but does not change it
  • Audit delivery - alert on pilot_audit_events_failure_total{error_type="dropped"}, which counts audit events lost for good (Delivery)
  • Review audit events - periodically review the audit log for unexpected operations

Backup & Recovery

  • Understand state model - Pilot stores state in internal Kafka topics (see Architecture for details)
  • Broker state - ensure __pilot_broker_state uses compaction (created and enforced automatically by Pilot)
  • Audit log - ensure __pilot_audit_log retention meets your compliance needs (default 30 days, configurable via AUDIT_RETENTION_MS)
  • PAT tokens - if PATs are enabled, __pilot_pat_tokens stores hashed tokens with compaction (created automatically)
  • Proposals - proposals are stored in memory and regenerated on demand; they are not backed up
Last updated on