Environment Variables
Pilot is configured entirely via environment variables. This page groups them by feature area with explanations and examples. For a flat alphabetical reference, see All Environment Variables.
Core
PORT=8080 # HTTP server port
KAFKA_BOOTSTRAP_SERVERS=broker1:9092,broker2:9092 # Kafka brokers
LOG_LEVEL=INFO # DEBUG, INFO or WARN
METADATA_UPDATE_INTERVAL=10s # Metadata collection frequencyMETADATA_UPDATE_INTERVAL controls how often Pilot samples partition watermarks and log directory sizes. Lower values give more responsive activity scores but increase Kafka metadata calls.
Balance Engine
PILOT_BALANCE_THRESHOLD=5.0 # Per broker: act beyond 2x this % from the fair share, then even to within this %
PILOT_BALANCE_RACK_AWARE=true # Enforce rack-aware placement
PILOT_BALANCE_MIN_RF=0 # Minimum replication factor (0 = disabled)
PILOT_BALANCE_WORKERS=0 # Candidate-generation goroutines (0 = auto)
# PILOT_BALANCE_FLOOR="bytes=2MB/s disk=2GiB" # Optional minimum difference (unset = off)- Threshold - Applies per broker and load. Pilot moves partitions only when a broker stays more than twice the threshold from its fair share of a load (10% at the default), and then evens that load until every broker is within the threshold, give or take one partition. Lower values even out brokers more closely and move more partitions. Higher values tolerate larger differences. See Even Balance.
- Rack awareness - When enabled, proposals never place two replicas of the same partition on brokers in the same rack.
- Min RF - Set to
2or3to automatically increase replication factor for under-replicated topics. See Replication Factor. - Minimum difference - Optional, unset (off) by default. Set it, for example to
bytes=2MB/s disk=2GiB, to keep a quiet dev or test cluster from moving partitions for small differences: a rate or stored-bytes metric is then balanced only when some broker is at least that far from its equal share. Leader and follower counts are not affected. Settings are separated by spaces or commas; settings left out usebytes=2MB/sanddisk=2GiB. Message rates follow the byte rate at the average message size unlessmsgs=sets them (for examplemsgs=5k/s);msgs=autois accepted.offis the same as unset. An invalid value stops Pilot at startup. See Minimum Difference. - Workers - Number of goroutines the proposal engine uses for candidate generation.
0(default) auto-detects the available CPUs (GOMAXPROCS, which respects container CPU limits). Set a positive value to cap parallelism, for example to leave CPU headroom for the API and samplers. Does not affect proposal output.
Broker Profile
PILOT_BROKER_PROFILE="network=10Gbit/s disk=2TiB" # Capacity per broker (unset = no utilization reporting)
PILOT_BROKER_PROFILE_FILE=/etc/pilot/brokers.conf # Same rules, one per line, # comments- Profile - Describes each broker’s sustained network throughput per direction (
network=) and log storage (disk=). Pilot reports each broker’s network and disk utilization from it. A profile does not change balancing: capacity never makes an imbalance count more or less, and fair shares do not depend on it. Capacity-based overload protection is planned. See Broker Profiles. - Settling - With
network=set, optional balancing after Pilot’s own moves waits for evidence (moved replicas in sync, planned leaders in place, post-move measurements) instead of the adaptive 5 to 30 minute wait. See Settling After Moves. - Mixed fleets - Add rules with a
rack=,host=orbroker=selector and a colon, for examplenetwork=10Gbit/s disk=2TiB; rack=eu-west-1c: network=25Gbit/s; broker=7-9: disk=8TiB. A selector ends at the last colon of its rule, so rack names with colons work. Per setting, a matchingbroker=rule beatshost=, which beatsrack=, which beats the default; among rules of the same kind, the one written last wins. Every setting a selector uses needs a default. - Units - Required, written without a space (
10Gbit/s). Network takes bit or byte rates (10Gbit/s,25Gbps,1.25GB/s,800MiB/s); disk takes sizes (2TiB,1.5TB,500GiB).Gbis rejected as ambiguous. - Validation - Set only one of the two variables; blank values count as unset. An invalid profile stops Pilot at startup. Balancing is the same with or without a profile.
Movement Control
PILOT_MOVE_MAX_PER_BROKER=20 # Max concurrent partition moves per broker
PILOT_PROPOSAL_MAX_PARTITIONS=0 # Max partitions a single proposal may move (0 = unlimited)- Max per broker - Limits concurrent reassignments to prevent overwhelming brokers. Applies per broker, not cluster-wide.
- Max partitions per proposal - Caps how many partitions a single generated proposal may move.
0(default) is unlimited. When a plan exceeds the cap, whole partition transitions with the lowest marginal benefit per movement cost are dropped first. The remaining optional plan must still be worth applying. Critical and replication-factor fixes are always kept, so correctness repairs can exceed this cap. The proposal metadata records that the plan was capped.
Throttle
PILOT_THROTTLE_RATE_MB=50 # Base replication throttle rate (MB/s)
PILOT_THROTTLE_MAX_RATE_MB=1000 # Maximum throttle rate (MB/s)
PILOT_THROTTLE_MANAGED=true # Pilot manages replication throttlesPilot applies PILOT_THROTTLE_RATE_MB as the replication throttle for its moves. You can change it at runtime with PUT /api/v1/reassignments/tuning, up to PILOT_THROTTLE_MAX_RATE_MB. Set PILOT_THROTTLE_MANAGED=false only if you manage throttles outside Pilot. Proposals are then marked unsafe to apply (throttles_externally_managed), a manual apply needs force=true, activity self-healing skips them, and Pilot does not clear leftover throttles at startup.
When throttle management is enabled, PILOT_THROTTLE_RATE_MB must be at least 1 MB/s. A value of 0 or negative stalls replication and can strand under-replicated partitions, so it is rejected at startup. If set, PILOT_THROTTLE_MAX_RATE_MB must not be below PILOT_THROTTLE_RATE_MB.
Self-Healing
Pilot runs three independent self-healing loops. All are disabled by default and start in dry-run mode.
Activity-Based Healing
PILOT_HEAL_ENABLED=false # Enable activity rebalancing
PILOT_HEAL_INTERVAL=30m # Check intervalCritical Fixes
PILOT_HEAL_CRITICAL_ENABLED=false # Enable URP and rack violation fixes
PILOT_HEAL_CRITICAL_INTERVAL=5m # Check intervalReplication Factor Increases
PILOT_HEAL_RF_ENABLED=false # Enable RF increase healing
PILOT_HEAL_RF_INTERVAL=15m # Check intervalCommon Settings
PILOT_HEAL_DRY_RUN=true # Simulate only (no actual moves)
PILOT_HEAL_MAX_PARTITIONS_PER_RUN=0 # Max partitions moved per cycle (0 = unlimited)
PILOT_HEAL_WINDOW_START=-1 # Start hour (0-23, -1 = always)
PILOT_HEAL_WINDOW_END=-1 # End hour (0-23, -1 = always)The time window supports overnight windows (e.g. PILOT_HEAL_WINDOW_START=22, PILOT_HEAL_WINDOW_END=5 means 22:00-05:59). Set both to -1 to disable the time restriction.
Recommended rollout: Enable with PILOT_HEAL_DRY_RUN=true first to observe what Pilot would do, then switch to false when confident.
Exclusions
PILOT_EXCLUDE_TOPICS="^__consumer_offsets$,^__transaction_state$"Comma-separated regular expressions. A topic that matches any of them is left out of proposals and self-healing. Patterns are not anchored, so orders also matches orders-dlq; use ^orders$ for a single topic. An invalid pattern is skipped with a warning. Internal Kafka topics (__consumer_offsets, __transaction_state) are commonly excluded.
Consumer Groups
PILOT_CONSUMER_GROUP_COLLECTION_ENABLED=true # Background collection
PILOT_CONSUMER_GROUP_COLLECTION_INTERVAL=15s # Collection intervalAlways-Ready Proposals
Pilot continuously monitors for cluster changes and regenerates proposals automatically. Detection cadence, drift sensitivity, and the regeneration interval are fixed engine defaults and need no configuration.
Observability
PILOT_FOLLOWER_LAG_WARNING_THRESHOLD=1000 # Warning threshold for total follower lag
PILOT_BROKER_EXPIRY=1h # Forget a permanently unavailable broker after this longPILOT_FOLLOWER_LAG_WARNING_THRESHOLD sets the offset count at which total cluster follower lag triggers a warning on the health endpoint. The default of 1000 is suitable for most clusters. Raise this value for high-throughput clusters where transient lag spikes are expected.
PILOT_BROKER_EXPIRY controls how long a broker must stay continuously unavailable, and unreferenced by any partition replica set, before Pilot forgets it. Pilot tracks the union of every broker it has ever seen so that a broker which goes offline still shows as unavailable. After a permanent scale-down that tracking would otherwise keep the removed brokers forever: TotalBrokers stays inflated and rolling restarts report the gone brokers as unavailable indefinitely. When a broker has been unavailable for at least this duration and no partition still lists it as a replica, Pilot drops it from the tracked union and tombstones its persisted state. Set to 0 to disable expiry. A broker that reappears in cluster metadata is tracked again, so expiry is reversible. Brokers still hosting replicas never expire.
AI Chat
PILOT_CHAT_ENABLED=false
PILOT_CHAT_PROVIDER=anthropic # "anthropic" or "openai"
PILOT_CHAT_OPENAI_API=auto # "auto", "responses", or "chat_completions" (openai only)
PILOT_CHAT_API_KEY=sk-... # Required if chat enabled
PILOT_CHAT_MODEL=claude-haiku-4-5-20251001
PILOT_CHAT_MAX_TOKENS=4096
PILOT_CHAT_MAX_CONTEXT_TOKENS=30000 # Estimated context budget per request (min 4096)
PILOT_CHAT_MAX_TOOL_RESULT_CHARS=8000 # Tool result size kept in context (min 1000)
PILOT_CHAT_BASE_URL= # Override API base URL (self-hosted models, Azure OpenAI)
PILOT_CHAT_APPROVAL_TIMEOUT=5m # Approval window for mutations
PILOT_CHAT_RATE_LIMIT=20 # Messages/min/user
PILOT_CHAT_MAX_CONVERSATIONS=10 # Per user
PILOT_CHAT_CONVERSATION_TTL=24h # Auto-expire conversationsWith the openai provider, PILOT_CHAT_OPENAI_API=auto uses Responses for GPT-5/GPT-6 model names and Chat Completions otherwise. See AI Chat for provider setup, including local models.
Azure OpenAI
PILOT_CHAT_PROVIDER=openai
PILOT_CHAT_BASE_URL=https://<resource>.openai.azure.com/openai # Trailing /openai required
PILOT_CHAT_API_KEY=<azure-openai-api-key>
PILOT_CHAT_MODEL=<deployment-name> # Azure deployment name; must support tool calling
PILOT_CHAT_OPENAI_API=responses # For GPT-5/GPT-6 deployment aliasesPilot appends /v1/responses or /v1/chat/completions to the base URL, so the trailing /openai segment is required. Select responses for GPT-5/GPT-6 deployment aliases; the deployment must support Responses.
Corporate CA Certificates (SSL_CERT_FILE)
Behind a TLS-intercepting proxy, set SSL_CERT_FILE to your CA bundle; the chat HTTP client honors this standard Go variable, while KAFKA_SSL_* affects only the Kafka client. The file replaces the system roots for the whole process, so use a full bundle with the corporate CA appended (cat /etc/ssl/certs/ca-certificates.crt corp-root.pem > ca-bundle.pem), not the corporate certificate alone.
services:
pilot:
volumes:
- ./certs/ca-bundle.pem:/etc/pilot/certs/ca-bundle.pem:ro
environment:
SSL_CERT_FILE: /etc/pilot/certs/ca-bundle.pemPilot Agent
PILOT_AGENT_ENABLED=false # Enable the agent gRPC server
PILOT_AGENT_GRPC_PORT=9190 # Port for agent gRPC connections
PILOT_AGENT_DATA_DIR=/data/agent-certs # Directory for auto-generated certificates
PILOT_AGENT_BOOTSTRAP_TOKEN= # Master token for agent cert enrollment (required for auto-TLS, generate with: openssl rand -hex 32)
PILOT_AGENT_TLS_SANS= # Extra DNS names/IPs for auto-generated server cert (comma-separated)
PILOT_AGENT_INSECURE=false # Allow gRPC without TLS (development only)
PILOT_DEPLOY_ALLOW_HTTP=false # Allow SSH deploy endpoints over plain HTTP (default: false, requires HTTPS)
PILOT_TRUSTED_PROXIES= # CIDRs or IPs of reverse proxies whose X-Forwarded-Proto/-Host are trusted (unset = trust all)
# Agent binary distribution
PILOT_AGENT_DOWNLOAD_URL=https://downloads.calinora.io/agent/v{version}/pilot-agent-linux-{arch} # Upstream URL (supports {version}/{arch} placeholders)
PILOT_AGENT_BINARY_DIR= # Local directory for air-gapped mode (no downloads when set)
PILOT_AGENT_BINARY_CACHE_DIR= # Cache dir for downloaded binaries (default: {data-dir}/binaries)
# Agent SSH deploy trust
PILOT_AGENT_SSH_KNOWN_FINGERPRINTS= # Comma-separated SHA256:<base64> host-key fingerprints trusted for agent SSH deploys. Merged into hops that do not supply their own list. See /features/agents#ssh-host-key-verificationWhen no explicit TLS paths are set, Pilot auto-generates an ECDSA P-256 CA and server certificate. Set PILOT_AGENT_BOOTSTRAP_TOKEN to a secure random string and provide the same value as AGENT_BOOTSTRAP_TOKEN on each agent (master token), or use single-use enrollment tokens created via the API. This variable is required when using auto-TLS. For self-provided certificates:
PILOT_AGENT_TLS_CERT=/certs/server.crt # Server certificate (disables auto-TLS)
PILOT_AGENT_TLS_KEY=/certs/server.key # Server private key
PILOT_AGENT_TLS_CA=/certs/ca.crt # CA for verifying agent client certsPILOT_TRUSTED_PROXIES lists the reverse proxies whose X-Forwarded-Proto and X-Forwarded-Host Pilot trusts, for the HTTPS check on SSH deploy endpoints and the URLs given to agents. Unset trusts every source; set it when clients can reach Pilot without the proxy.
See Pilot Agent for full deployment details.
Agent-Side (Broker Lifecycle)
Agents can restart, stop, and start the Kafka broker on their host. These variables configure how the agent manages the broker process:
AGENT_KAFKA_SERVICE_UNIT=kafka.service # Override systemd unit name
AGENT_KAFKA_START_CMD= # Override start command
AGENT_KAFKA_STOP_CMD= # Override stop command
AGENT_ALLOW_RESTART=true # Set to false to disable lifecycle commands
AGENT_DOCKER_CONTAINER= # Docker container name for lifecycle commands (containerized brokers)AGENT_DOCKER_CONTAINER specifies the Docker container name to manage via docker stop/start/restart. This is required for containerized broker deployments when using pid: "host" mode, so the agent knows which container to control.
The agent auto-detects the service manager (systemd or exec fallback) during discovery. Override these only if auto-detection does not match your setup.
Authentication (UI & API)
AUTH_ENTRAID_ENABLED=true
AUTH_ENTRAID_TENANT=your-tenant-id # derives auth, token, and issuer URLs
AUTH_ENTRAID_CLIENT_ID=your-client-id
AUTH_ENTRAID_CLIENT_SECRET=your-client-secret
AUTH_COOKIE_SECRET=random-32-char-secretAUTH_<PROVIDER>_ENABLED must be present for the rest of a provider’s block to be read.
AUTH_<PROVIDER>_ISSUER is what lets Pilot verify a bearer JWT: signing keys are discovered at <issuer>/.well-known/openid-configuration and cached for AUTH_<PROVIDER>_JWKS_CACHE_TTL (default 15m). It is the trust anchor, so it is required for local verification. AUTH_<PROVIDER>_JWKS_URL on its own only overrides where keys are fetched, it does not establish which issuer to trust. The same issuer verifies the id_token during browser login, so a stale value that no longer matches the token’s iss returns 401. Providers that issue opaque tokens need AUTH_<PROVIDER>_API_URL instead, since only the provider can validate those. GitHub always does; Google access tokens (ya29.*) are opaque too; and a generic OIDC provider does whenever it is configured for reference tokens (Auth0 without an API audience, Okta reference tokens). For those, setting AUTH_<PROVIDER>_ISSUER enables id_token verification at login but disables bearer auth for that provider, because an opaque token declares no audience to bind.
A verified bearer JWT must also be minted for this deployment. Its aud or azp must appear in AUTH_<PROVIDER>_ALLOWED_AUDIENCES, which defaults to AUTH_<PROVIDER>_CLIENT_ID (Entra ID: api://<client-id> and <client-id>) when it is not set. This is a fail-closed check and it is a breaking change for deployments that previously authenticated on an unverified token body: a mismatch returns 403 invalid audience.
Setting AUTH_<PROVIDER>_ALLOWED_AUDIENCES yourself is stronger than leaving it to the default. The derived value is waived for a provider that cannot verify JWTs, since a userinfo response declares no audience for it to match; a value you configure is never waived, even when it equals the derived one. Do not set it for a userinfo-only provider such as GitHub, or every bearer token from it is rejected with 403 invalid audience.
See UI & API Authentication for the full per-provider matrix.
Claim Mapping
AUTH_<PROVIDER>_USERNAME_CLAIM= # Claim for the session username. Default chain: upn, preferred_username
AUTH_<PROVIDER>_EMAIL_CLAIM= # Claim for the session email and ALLOWED_DOMAINS. Default chain: email, preferred_username, upn
AUTH_<PROVIDER>_GROUPS_CLAIM= # Claim for session groups and ALLOWED_GROUPS. Default chain: groups, roles, role, groupA configured claim is consulted first, with the default chain as fallback; the first claim present with a non-empty value wins.
AD FS example. Claim names depend on the relying party’s issuance rules; group membership commonly arrives under roles:
AUTH_OIDC_GROUPS_CLAIM=roles
AUTH_OIDC_USERNAME_CLAIM=upn
AUTH_OIDC_ALLOWED_GROUPS=kafka-adminsSee Claim Mapping for details.
MCP Server
PILOT_MCP_ENABLED=true # Enable Model Context Protocol serverSee Model Context Protocol for details.
License
LICENSE_STRING=eyJ... # JWT license token
LICENSE_FETCH_SUBSCRIPTION_ID=sub_123 # For automatic fetching
LICENSE_FETCH_TOKEN=tok_abc # Per-customer fetch token
LICENSE_FETCH_URL=https://license.calinora.io/api/license/fetch
LICENSE_FETCH_INTERVAL=1hSee Licensing for details.