Skip to Content
ReferenceChangelog

Changelog

All notable changes to Calinora Pilot are documented here.

0.27.0 (2026-10-07)

Balancing

  • Every broker gets its fair share: Pilot compares each broker with its fair share of every load and moves partitions when one stays more than 10% off, then evens that load until every broker is within 5%. The rule is the same on 3 brokers or 800, and a new, empty broker is filled in one plan. See Even Balance.
  • Smaller plans: Pilot leaves well-spread topics alone. A topic with too many leaders on one broker gets a leader-only repair, also in an otherwise balanced cluster. See Concentrated Topics.
  • Calmer rebalancing: short traffic spikes, idle loads and retention cuts do not move partitions. Pilot prefers leadership changes to data copies, does not move a partition again right after moving it, and applies a plan only once fresh measurements confirm it. See Recent Moves.
  • Compacted topics such as __consumer_offsets count at their steady size, so the log cleaner’s cuts do not move partitions, also after a Pilot restart. See Compacted Topics.
  • The plan you review is the plan that runs: Pilot keeps a plan while it holds, and the page tells you when a newer one is available. See Plan Stability.
  • When a broker cannot get closer, Pilot shows the load at its limit, with the cause and what would help. See At Its Limit.
  • New optional PILOT_BALANCE_FLOOR to ignore small differences, for example on quiet test clusters. See Minimum Difference.
  • New broker profiles (PILOT_BROKER_PROFILE or PILOT_BROKER_PROFILE_FILE) show each broker’s network and disk utilization; with network capacity set, Pilot also continues sooner after its own moves. See Broker Profiles.
  • Clearer balance summary and proposal review: per load, the broker farthest from its fair share, where the plan leaves it, and why a plan is or is not recommended. See Reading the Balance Summary.

Operations

  • The first plan after maintenance moves partitions back onto the returning broker. See Refill After Maintenance.
  • Bulk partition reassignment runs as a monitored operation that can be followed and cancelled. See Bulk Reassignment.
  • Applying proposals and reverts go through a review step that re-checks the cluster first; offset resets, quota, topic configuration and ACL changes get the same review step in the UI.
  • Reassignment progress shows the state of every partition, with a reason for each failure.
  • More reliable redistribution, rolling restarts and rate measurements.
  • Request quotas accept values above 100%, where 100% is one thread per broker.
  • The Pilot Agent forwards a log file from its end and does not replay existing lines on start. See Agents.
  • Built with Go 1.27.1.

Monitoring

  • Storage, lag and rack-safety views keep working while a broker is offline and show what could not be measured. See Measurement Coverage.
  • New metrics for balance state and rack-safety coverage, and two consumer-group gauges for when the last offset collection started and completed. See Metrics.
  • Pilot keeps each broker’s last reported rack through broker outages, failed metadata reads and its own restarts, so rack safety is judged on the racks it knows.

Audit and Security

  • Changes Pilot makes on its own, such as self-healing, now appear in the audit log. Audit events are retried when Kafka is briefly unavailable, and a new audit topic is replicated even when the broker default is one replica. See Audit Logging.
  • Security improvements for assistant approvals and access tokens.
  • Kafka connections can authenticate with a client certificate (mutual TLS): KAFKA_SSL_CERT_* and KAFKA_SSL_KEY_*, with KAFKA_SSL_KEY_PASSWORD for an encrypted key. See Mutual TLS.
  • Kafka connections support SASL SCRAM-SHA-256 and SCRAM-SHA-512 with KAFKA_SASL_USERNAME and KAFKA_SASL_PASSWORD. See SASL Authentication.

UI and Assistant

  • Redesigned interface, with navigation grouped into Cluster, Operations and Access.
  • New command palette on Cmd/Ctrl+K. The assistant moved to Cmd/Ctrl+Shift+K.
  • Ask Pilot opens beside the page, suggests questions about what you are looking at, and supports GPT-6. See AI Chat Assistant.
  • Views refresh automatically, and what-if covers replication-factor changes.
  • Message browser pages across partitions. Restart and agent upgrade controls are hidden for Strimzi-managed brokers, which restart through the broker image.
  • The assistant panel stays open across a browser refresh and restores the conversation, works with local and OpenAI-compatible models, and retries a provider request that fails on the way in. See AI Chat Assistant.

Upgrade notes

  • PILOT_BALANCE_THRESHOLD now applies to each broker: Pilot acts when a broker is more than twice the value from its fair share and evens to within the value. Review a pinned value: 10 now acts only beyond 20%. Pilot does not start with a value of 0 or below.
  • Clusters balanced under the previous rule get one catch-up plan after the upgrade. It may move partitions the previous rule moved shortly before.
  • Pilot spreads a topic’s leaders only when one broker leads too many of them, in a plan of its own if needed. The first plan after the upgrade may include a one-time leader-only repair.
  • While a broker is down and not in maintenance, balancing is paused and Pilot plans only repairs (status Balancing paused, API field balanceHold). Drain a broker that stays down with maintenance mode.
  • Let a running reassignment finish before restarting or upgrading Pilot: a start resets the replication throttles and does not resume tracking the moves it submitted.
  • Proposals explain balance per broker: new fairShare, topicSpread, benefitEvidence.tripped and benefitEvidence.topicBreaches, structuralLimits entries with brokerId, distance and cause, metrics.effective holding half the farthest broker’s distance, and the reason balance_held while balancing is paused. Status messages name the broker and the load. See API Fields.
  • Pilot needs DescribeConfigs on the cluster to read default topic configuration, instead of creating a temporary topic.
  • With OpenAI, GPT-5 and GPT-6 models use the Responses API. Set PILOT_CHAT_OPENAI_API=chat_completions for endpoints that only support Chat Completions. See OpenAI.
  • KAFKA_SSL_CERT_* and KAFKA_SSL_KEY_* now take effect under SSL and SASL_SSL: a client certificate or key that cannot be loaded stops Pilot at startup, so remove stale values before upgrading. See Mutual TLS.
  • Applying a balancing proposal through the API requires the current proposal, checked within the last two minutes (checkedAt), with decision.applicationAllowed; otherwise it returns 409. force=true overrides these checks, but not for a changed plan or a proposal that is no longer current. Activity self-healing uses the same check instead of requiring High confidence.
  • MCP: generate_proposal returns the current proposal and its decision instead of calculating a new one, and apply_proposal is blocked unless decision.applicationAllowed is true. See MCP.
  • POST /api/v1/topics/{topic}/partitions/bulk now executes live entries. Set preview: true on each entry for a preview. An empty partitions list returns 400.
  • POST /api/v1/topics/{topic}/preferred-leader-election only affects the topic in the path and returns 400 for a body naming another topic. Use POST /api/v1/preferred-leader-election for the whole cluster.
  • Monitoring, what-if and redistribution endpoints can return 503, or null values with a coverage field, while data is unavailable. /api/v1/ready returns 503 while Pilot’s view of the brokers is stale or unavailable, not only when no broker is reachable.
  • Reverting a reassignment from the audit log requires a review first (POST /api/v1/audit-log/revert/preview) and returns 202 once accepted. See Reassignment Review.
  • Audit events from MCP and the assistant use method MCP or CHAT (was chat-agent) and the user’s auth provider (was pilot). Pilot’s own changes use method INTERNAL and user pilot. Applies are recorded when submitted, and a later failure adds an event with status 500.
  • Alerting, see Metrics:
    • Lost audit events are pilot_audit_events_failure_total{error_type="dropped"}; the other error types now count attempts that are retried.
    • pilot_cluster_not_rack_aware_partitions is reported only for a complete rack assessment; check pilot_cluster_rack_assessment_status.
    • Broker counts keep the last successful discovery; alert on pilot_cluster_broker_observation_current == 0 for failed discovery.
    • The self-healing skip reason low_confidence is now data_not_ready or benefit_unconfirmed.
    • New self-healing skip reason partitions_in_cooldown: automatic balancing waits while a partition is still in its 30-minute cooldown.

0.26.0 (2026-09-04)

Behavior change for MCP clients: list-shaped tool results are now bounded summaries by default, some output shapes changed, and new drill-down parameters fetch detail. Per-tool limits and changed shapes: Output Limits and Drill-Down.

  • Self-healing no longer caps partitions moved per cycle by default: PILOT_HEAL_MAX_PARTITIONS_PER_RUN default changed from 500 to 0 (unlimited); set a value > 0 to restore a cap
  • Chat keeps its LLM context bounded: new PILOT_CHAT_MAX_CONTEXT_TOKENS and PILOT_CHAT_MAX_TOOL_RESULT_CHARS settings, JSON-aware truncation of oversized tool results, and recovery from context overflows
  • Fixed reassignments being silently skipped as “already at target” on stale Kafka metadata
  • Fixed the topic redistribution RACK_AWARE strategy (the UI default) assigning every partition of a topic to the same single broker per rack, which collapsed the topic onto one broker per rack and left the other rack members with no replicas. Replicas and leaders now spread across all brokers in each rack while keeping one replica per rack, and preview and apply output remains deterministic
  • Configurable OAuth/OIDC claim names: new AUTH_<PROVIDER>_USERNAME_CLAIM, AUTH_<PROVIDER>_EMAIL_CLAIM, and AUTH_<PROVIDER>_GROUPS_CLAIM choose which claims feed the session identity and the ALLOWED_DOMAINS / ALLOWED_GROUPS filters. Behavior change: the default groups lookup is now groups, roles, role, group (first non-empty wins), so an existing ALLOWED_GROUPS filter also matches tokens that carry membership under roles, role, or group. The login flow now also merges upn, unique_name, given_name, and the role-shaped claims from the id_token, so providers whose userinfo endpoint returns only sub (AD FS) resolve a full identity. See Claim Mapping
  • OAUTHBEARER token endpoint now trusts the CA from KAFKA_SSL_CA_CERT_FILE / KAFKA_SSL_CA_CERT_PEM in addition to the system roots, so a token endpoint signed by a private CA no longer fails with x509: certificate signed by unknown authority; KAFKA_SSL_INSECURE_SKIP_VERIFY applies to it too. KAFKA_SSL_CA_CERT_PEM is now honored for the broker connection as well (it was previously ignored). See Token Endpoint TLS
  • Kafka security configuration is validated at startup: a SASL protocol without credentials, an OAUTHBEARER setup without a token endpoint, or a reserved auth key in KAFKA_SASL_OAUTH_EXTENSIONS now fails fast at boot instead of failing on every broker handshake
  • Clearer proposal confidence statuses in the UI. “Settling…” is now “Reassignment in flight”, “Stabilizing” is now “Settling after moves”, and the catch-all “Warming up” has been replaced by cause-specific statuses: “Building history” (partial metric window), “Awaiting convergence” (recent generations still disagree), and “Building confidence” (bursty load or stale metadata). “Collecting data” now only appears while the metric window is genuinely empty. See Proposal Confidence
  • The rack-safety advisory now reports severity warning instead of high when the only finding is fewer racks than the replication factor (2 or more racks, complete rack metadata, no co-location detected), and that warning is gated on min.insync.replicas: it is suppressed entirely when every affected topic provably tolerates a single rack loss, states the concrete survivor-vs-min.insync.replicas shortfall (and that acks=all producers would stall) when a topic is exposed, and falls back to the generic rack-loss warning when min.insync.replicas cannot be determined. Unracked brokers, co-located partitions, or a single rack still report high. See Rack-Safety Advisory

0.25.0 (2026-08-03)

  • Calmer proposals: partitions without traffic are no longer shuffled between brokers when leaders are already well spread
  • Chat works with newer OpenAI models (the GPT-5 family) and Azure OpenAI deployments, with new docs for Azure OpenAI and corporate CA certificates
  • Fixed GitHub OAuth sign-in; GitHub users now have a stable identity in sessions and audit records
  • MCP and chat mutation tools are license-gated in parity with the REST API; read, proposal, and what-if tools stay free
  • Message browser can filter messages by key or value; refreshed logo, icon, and favicon with light/dark support
  • Pilot Agent Docker image base updated to Alpine 3.24; updated gRPC, MCP, and embedded UI libraries

0.24.0 (2026-06-17)

Proposal Engine

  • Better balance quality: topics that expand to more brokers no longer re-concentrate, per-topic replica and leader spread is preserved, and balance converges further after brokers rejoin the cluster.
  • Critical repairs (under-replicated partitions, rack violations, replication-factor shortfalls) are always kept in a proposal, and the proposal status shows when only critical fixes remain.
  • Clearer advisories that explain why a cluster cannot balance further.

Safety & Reliability

  • Every mutation runs a live cluster-health check first. Genuinely unsafe operations are blocked, softer warnings can be overridden, and overrides are audited. See Rebalancing Proposals.
  • Stronger pre-flight validation of replica sets, replication factor, and min.insync.replicas, with rack-safety and rack-collapse risks surfaced as warnings. See Maintenance Mode.
  • More reliable reassignments: moves are throttled, recover cleanly from deleted topics, unavailable brokers, and cancellation, and mutating operations no longer overlap.
  • Broader audit coverage, including license-gated audit revert and a record of who requested and who approved each action.

MCP

  • MCP mutations share the same execution path, health gating, and validation as the API, and return a clear reason when an action is blocked. See Execution Parity.
  • Approvals require a distinct approver: a client cannot approve its own action.

Chat

  • Chat-initiated mutations can be approved by the chat user.
  • Fixed a streaming crash with the OpenAI provider.

UI

  • Rack-safety and rack-collapse risks are now surfaced in the proposals and maintenance views.
  • More accurate data-movement estimates, and consistent message-rate and broker labels across the app.
  • Navigation, dark-mode, and labeling fixes.

Configuration

  • New PILOT_BROKER_EXPIRY setting to auto-expire brokers that have been permanently removed.
  • New PILOT_PROPOSAL_MAX_PARTITIONS setting to cap how many partitions a single proposal moves.
  • Stricter throttle-rate validation at startup. See Throttle configuration.

Removed

  • The quota analyzer (bottleneck analysis) has been removed across the API, UI, MCP tools, and Prometheus metrics. Quota management (view, create, update, and delete) is unchanged.

Dependencies & Runtime

  • Updated the Kafka client and Go toolchain.

0.23.0 (2026-05-28)

Security

  • Updated golang.org/x/net to v0.55.0 to address CVE-2026-39821.

0.22.0 (2026-05-26)

Proposal Engine

  • Faster proposal generation on large clusters. Candidate generation is now parallelized; the new PILOT_BALANCE_WORKERS env var caps the worker pool (default 0 auto-detects available CPUs and respects container CPU limits). Does not affect proposal output.
  • The rebalancing objective now uses seven metrics per broker: leader count, follower count, disk bytes, producer message rate, producer byte rate, consumer message rate, and consumer byte rate. The composite activity score remains as a dashboard “how busy is this broker?” indicator but no longer feeds the optimizer.
  • Smoother proposal confidence. The confidence label no longer flips back and forth between adjacent states, and “Up to date” now also requires recent generations to agree on the recommendation. A cluster whose recommendation keeps changing is held at “Warming up” until it converges.
  • Aligned “is this worth applying?” check. The “likely not beneficial” advisory shown in the UI and the self-healing activity loop now use the same threshold (>= 2 percentage points absolute improvement, or >= 10% relative improvement while still clearing 0.5pp). A structurally limited proposal will not auto-apply unless it clears one of these bars, and the advisory and the self-healing decision can no longer disagree.
  • More accurate rate calculations.

Pilot Agent

  • OAuth/OIDC token caching for SASL/OAUTHBEARER. Access tokens are reused across Kafka connections within the issuer’s refresh margin, with concurrent refreshes coalesced into a single fetch. Eliminates token-endpoint storms during broker reconnects.
  • Agent host metrics re-exported on Pilot’s /metrics endpoint. CPU, memory, file descriptors, disk, log dirs, network, and disk I/O are now exposed as pilot_agent_* gauges labelled by broker_id and node_id. Series are cleared when an agent disconnects so stale gauges do not trigger false alerts.
  • New lifecycle and inventory gauges on pilot_agent_*: broker_running (separates broker-down from agent-down), uptime_seconds, cert_expires_at_timestamp_seconds (for mTLS rotation alerting), upgrade_available, discovery_confidence, plus an info inventory gauge labelled with agent version, Kafka version, cluster ID, KRaft role, and distribution vendor. Disk and log-dir available_bytes, network link_speed_mbps, and disk I/O weighted_time_ms are also now exposed.

UI

  • Redesigned quota page. The single overview is split into focused panels for the quota table, quota editor, effective-quota lookup, active clients, and bottleneck analysis, with consistent units across all views.

0.21.0 (2026-04-30)

  • Pilot Agent: optional on-broker companion with mTLS gRPC, heartbeats, and system metrics
  • Pilot Agent: opt-in non-root mode via a dedicated pilot-agent system user plus a narrow sudoers allow-list (systemctl restart|stop|start|is-active|show) scoped to the detected Kafka unit
  • Pilot Agent: new pilot-agent detect and pilot-agent configure-sudoers subcommands for install-time and drift-recovery workflows
  • Pilot Agent: process-based Kafka unit detection (follows the kafka.Kafka JVM), so arbitrary unit names like confluent-kafka.service or vendor-custom names work without configuration
  • Pilot Agent: --kafka-unit and --kafka-group flags on install-agent.sh and matching fields under Advanced Settings in the UI deploy dialog
  • Pilot Agent: strict SSH host key verification for agent deploys (breaking change from implicit auto-accept). Deploys require hostKey, knownFingerprints, or an explicit allowInsecureHostKey opt-in. New PILOT_AGENT_SSH_KNOWN_FINGERPRINTS env var installs a fleet-wide fingerprint allowlist; the deploy dialog exposes per-hop fingerprints and an audited insecure-bypass checkbox
  • Rolling restart orchestrator with per-broker drain, stop, start, ISR catch-up, and leader restore phases, plus safety preflight and orphan detection
  • SSH-based agent deployment (single, bulk, and streaming with live progress) and connection test endpoint, gated behind license and HTTPS
  • In-app agent binary distribution with SHA256 integrity and remote self-upgrade via gRPC
  • Single-use enrollment tokens and auto-TLS certificate issuance for agents (CSR signing)
  • Live broker log streaming via SSE and buffered log search per broker
  • Personal access tokens (PATs) for MCP and CI clients, backed by a compacted state topic
  • What-if simulator and blast radius analysis for hypothetical cluster scenarios
  • Model Context Protocol (MCP) server with 67+ tools for AI agent integration
  • Message browser for inspecting Kafka messages from the UI and API
  • Proposal engine refactor with improved multi-phase pipeline
  • UI redesign with improved performance and navigation
  • Audit logger enhancements for better event tracking, including per-broker rolling-restart phases and agent deploy/upgrade actions

0.20.0 (2026-03-03)

  • UI redesign with new dashboard layout and improved navigation
  • Model Context Protocol (MCP) server implementation
  • Message browser for Kafka topic message inspection
  • Proposal engine refactor with multi-phase pipeline architecture
  • Audit logger improvements and bug fixes
  • UI performance optimizations

0.19.0 (2026-02-11)

  • Automatic license fetching with pay-as-you-go subscription support via Stripe
  • Simplified throttle configuration (removed dynamic throttle calculation)
  • Full build binary improvements (embedded UI, docs, and license notices)

0.18.0 (2026-01-23)

  • Live throttling configuration - update throttle rates and concurrency via API without restart
  • Quota API resilience fixes for broker outage scenarios
  • UI updates and bug fixes for proposal data display

0.17.0 (2026-01-19)

  • Client quota management - view, create, update, and delete Kafka client quotas
  • Log directory management - view and move log directories across brokers
  • Improved rack API with response caching
  • Self-healing and stale partition bug fixes
  • Switched to Debian 12 nonroot base image

0.16.0 (2026-01-12)

  • Structural traffic limit diagnostics in proposals - identify uncorrectable imbalances with remediation suggestions
  • Smarter rebalancing with improved cost estimation
  • Topic spread phase for even leader distribution per topic

0.15.0 (2026-01-11)

  • Streaming reassignment executor for efficient partition moves with managed throttling
  • Enhanced reassignment monitor with real-time progress tracking
  • UI improvements for proposal review and reassignment monitoring

0.14.0 (2026-01-08)

  • Enhanced proposal engine (v2) with two-level search optimization (focused + broad)
  • Improved scoring model with weighted squared excess and cost weighting

0.13.0 (2026-01-07)

  • Redesigned reassignment monitor with completion tracking and failure detection

0.12.0 (2026-01-07)

  • Reassignment cancellation support
  • Switched to Alpine-based Docker image

0.11.0 (2026-01-07)

  • Wave-based processing for partition reassignments (batched submission with per-broker limits)

0.10.0 (2026-01-06)

  • Instantaneous rate metrics for proposal generation
  • Improved stuck partition detection and handling

0.9.0 (2026-01-04)

  • Enhanced proposal engine with rack-awareness validation
  • Rack-aware partition writes (enforce rack constraints during reassignment)
  • KPI click-through in UI for quick navigation to relevant data

0.8.0 (2025-12-09)

  • Self-healing time windows - restrict automatic healing to specific hours
  • Rack awareness fixes for partition placement
  • Timeout handling improvements for large topic operations
  • Health monitor race condition fix

0.7.0 (2025-11-05)

  • Self-healing types - three independent loops (activity, critical fixes, RF increases)
  • Broker maintenance mode with partition evacuation
  • Replication factor increase feature (target RF configuration)
  • Stuck reassignment detection and resolution

0.6.0 (2025-10-13)

  • Wave-based batching for partition reassignments
  • Configurable balance threshold (PILOT_BALANCE_THRESHOLD)
  • Dynamic throttle rate management

0.5.0 (2025-10-10)

  • Topic redistribution via API
  • Dynamic throttling for reassignment operations

0.4.0 (2025-09-26)

  • OAuth2/OIDC authentication (Entra ID, Google, GitHub, Keycloak)
  • Kafka-backed audit logging (__pilot_audit_log topic)
  • Broker state persistence to Kafka (__pilot_broker_state topic)
  • URP pre-pass in proposal generation (fix under-replicated partitions first)
  • Swagger UI behind reverse proxy support
  • Renamed from KafkaPilot to Calinora Pilot

0.3.0 (2025-09-17)

  • Async proposal engine with background generation
  • Enhanced proposal scoring algorithm
  • UI data refresh improvements
  • Swagger UI served locally (no external CDN)
  • Default topic configuration fixes

0.2.0 (2025-09-04)

  • TLS/SSL support for Kafka connections
  • Docker Compose setup improvements

0.1.0 (2025-09-03)

  • Initial release
  • Kafka cluster monitoring with metadata-only sampling
  • Activity scoring based on watermark deltas and log directory sizes
  • Partition rebalancing proposal generation
  • React dashboard with broker, topic, and partition views
  • REST API for cluster operations
  • Docker image with embedded UI
  • Prometheus metrics endpoint
  • SASL/PLAIN authentication support
Last updated on