Skip to Content
ConfigurationAudit Logging

Audit Logging

Pilot records mutating operations to a Kafka topic for compliance and traceability. When authentication is enabled, audit events capture who performed what action and when.

How It Works

  • Mutating API requests (POST, PUT, DELETE), MCP and assistant tool calls, and changes Pilot makes on its own generate audit events
  • Events are produced to the __pilot_audit_log Kafka topic
  • Events include: action type, user identity, timestamp, request details, and result
  • Reassignments are recorded when they are submitted for execution, not when they finish. A failure after submission adds a separate event with status 500, error, query.moves and the monitored operation ID (query.bulkId, query.reassignmentId or query.executionId)
  • The audit topic is created automatically on startup with cleanup.policy=delete

Event Sources

The method field identifies where an event came from:

methodSourceuserId
POST, PUT, DELETEREST API requestAuthenticated user (AUDIT_USER_ID_CLAIM), empty without authentication
MCPMCP client tool callApproving user (unknown without authentication); requestedBy holds the requester when it differs
CHATAssistant tool callConversation owner who approved the action; requestedBy holds chat-llm:<owner>
INTERNALChange Pilot made on its ownpilot

provider is the auth provider of the user, pat for MCP sessions authenticated with a personal access token, and empty on INTERNAL events.

Large Change Lists

Events that list partition changes, such as reassignments.apply, selfheal.apply and broker.maintenance.set, carry one row per partition in after. ple.trigger from the REST API lists the partitions whose election ran or failed, with query.requested, query.elected, query.notNeeded and query.failed counts. From MCP and the assistant it records only the cluster or topic scope. Lists longer than 500 rows are written as several records that share a chunkGroup, with chunkIndex and chunkTotal set. The API and the Audit Log page return them merged into one event. If some chunks are not available, the merged event sets chunksMissing to their number and after holds only part of the rows. Consumers reading the topic directly must merge records by chunkGroup.

Changes Pilot Makes on Its Own

Pilot records changes it makes without a request under userId pilot and method INTERNAL, with an empty provider and no requestId. The Audit Log page shows them under Pilot.

ActionPathRecorded when
selfheal.apply/self-healing/<trigger>An automatic apply submits reassignments (status 202) or fails after its submission was recorded (status 500)
selfheal.cancel/self-healing/<type>/cancelThe execution window closes on a running self-heal and Pilot cancels its active reassignments
license.update/license/auto-fetchThe license auto-fetch changes the license state, ID, plan or expiry, or loads a token that is rejected
broker.expire/health-monitor/broker-expiryThe health monitor stops tracking a broker that stayed unavailable and hosted no replica for PILOT_BROKER_EXPIRY
agent.renew-cert/agents/renew-certPilot signs a renewed client certificate for a Pilot Agent
rolling-restart.*/rolling-restartEach step of a rolling restart; query.restartId identifies the restart
throttles.reconcile/throttles/reconcileAt startup with PILOT_THROTTLE_MANAGED=true, Pilot removes replication throttles left on a broker from an earlier run
reassignments.timeout-revert/reassignment-monitor/timeoutA tracked reassignment is still running 12 hours after it was last submitted to Kafka, and Pilot cancels it back to its old replicas

Automatic Applies

A selfheal.apply event is written when a self-healing loop or the rebalance after a maintenance drain submits its plan for execution, before any move is sent to Kafka. Follow reassignment progress for the outcome.

FieldContent
status202
proposalIdProposal the plan came from
brokerIdDrained broker (post-maintenance only)
topicEmpty, since a plan spans topics
query.triggeractivity, critical, rf or post-maintenance
query.healingTypeSelf-healing type whose window, cooldown and cap applied: the trigger, or activity for post-maintenance
query.movesNumber of partitions submitted
query.bulkId, query.drainBulkIdMonitored bulk operation, and the drain for post-maintenance
afterPer partition: topic, partitionId, oldReplicas and oldLeader read from live metadata at submission, newReplicas and newLeader as submitted

An apply that fails after its submission was recorded (status 202) writes one more event with status 500, error and query.moves set to the planned count. This holds even if no move reached Kafka, for example when the replication throttle cannot be applied and the run stops before its first move. An apply refused before submission, for example when every move is filtered out as leaderless during an outage, writes no event, because the audit log records changes to the cluster. The refusal is logged and, for self-healing loops, counted as pilot_selfhealing_runs_total{result="error"}.

A dry run (PILOT_HEAL_DRY_RUN=true) and a cycle skipped by a gate, such as the time window, cooldown or license, write nothing. Once Pilot begins shutting down it submits no further automatic applies, so none runs unrecorded. Use the self-healing metrics for skips.

selfheal.apply, throttles.reconcile and reassignments.timeout-revert events cannot be reverted from the Audit Log or with the MCP revert_audit_event tool.

Other Changes

  • selfheal.cancel is written once per cancelled self-heal and carries proposalId, query.healingType, query.cancelled (number of reassignments cancelled) and query.window (for example 22-5). A failed cancel has status 500 and error.
  • license.update carries state, licenseId, plan and expiresAt (RFC3339 UTC, or empty) in before and after. A rejected token has status 500 and error.
  • broker.expire carries the broker in brokerId and before, and query.ttl, query.unavailableSince and query.tombstoned (whether its record in __pilot_broker_state was removed). A failed removal has status 500 and error.
  • agent.renew-cert carries the agent in requestedBy, its broker in brokerId, query.commandId, and the old and new certificate serial and expiry in before and after. A signing failure has status 500 and error.
  • throttles.reconcile is written once per broker and carries the broker in brokerId and the removed values in before, keyed by config name (leader.replication.throttled.rate, follower.replication.throttled.rate and replica.alter.log.dirs.io.max.bytes.per.second). Pilot clears a broker whose own dynamic leader.replication.throttled.rate or follower.replication.throttled.rate is set; cluster-wide defaults and static settings stay in effect. A failed removal has status 500 and error, and the values may still apply.
  • reassignments.timeout-revert is written once per partition and carries topic, partition, query.reassignmentId, query.bulkId (for a move in a bulk operation), query.timeout and query.submittedAt (last submission to Kafka, RFC3339 UTC). before holds the target replicas the move was heading to and after the old replicas it was cancelled back to, each as {"replicas": [...]}. A failed cancel has status 500 and error, and the move may still be running.

Topic Configuration

The audit topic’s name, retention, partition count, and replication factor are all configurable via environment variables:

VariableDefaultDescription
AUDIT_ENABLEDtrueEnable audit event logging
AUDIT_TOPIC__pilot_audit_logKafka topic name
AUDIT_RETENTION_MS2592000000 (30 days)Topic retention in milliseconds
AUDIT_PARTITIONS1Number of partitions
AUDIT_REPLICATION_FACTOR0 (broker default)Replication factor for a new topic (0 uses the cluster default, see below)
AUDIT_USER_ID_CLAIMupnIdentity recorded as userId: upn (default), sub or email. Other values fall back to upn, then sub, then email
AUDIT_STORE_MAX_ITEMS10000Max audit events kept in memory for the UI

Pilot creates the topic on startup if it does not exist, and sets cleanup.policy=delete and retention.ms from AUDIT_RETENTION_MS on every startup, also on an existing topic.

With AUDIT_REPLICATION_FACTOR=0, a new audit topic uses the brokers’ default.replication.factor. If that is 1 on a cluster of two or more brokers, Pilot creates the topic with 3 replicas instead (2 on a two-broker cluster), so a single broker outage does not stop audit delivery. If creation fails, Pilot falls back to the broker default, then to 1. If creation with a set AUDIT_REPLICATION_FACTOR fails, Pilot retries with 1. An existing audit topic with one replica on a multi-broker cluster logs a warning at startup; Pilot does not change its replication factor.

Delivery

Pilot waits up to 5 seconds for the broker to confirm each event. Reassignments from self-healing, MCP and the assistant are recorded before they are sent to Kafka. Other changes are recorded right after Kafka accepts them. Rolling-restart steps are recorded in the background. A write that is not confirmed does not stop the change.

An unconfirmed event is kept in an in-memory retry queue (up to 1,024 events or 32 MiB) and retried with a delay that grows from 1 to 30 seconds. On shutdown, Pilot keeps delivering queued events for up to 10 seconds. An event is dropped and logged as an error when it is still undelivered after 10 minutes or at the end of shutdown, when the queue is full, or when the broker rejects it as too large or invalid. Each dropped event increments pilot_audit_events_failure_total{error_type="dropped"}.

Querying Audit Events

Via the UI

The Audit Log page in the Pilot dashboard provides a filterable view of all recorded events.

Via the API

GET /api/v1/audit-log

Supports query parameters for filtering:

ParameterDescription
actionFilter by action type
userIdFilter by user ID (case-insensitive substring)
afterInclusive start timestamp (RFC3339)
beforeInclusive end timestamp (RFC3339)
searchFree-text search
limitMaximum events to return (default: 500)
offsetPagination offset

Example:

curl "http://localhost:8080/api/v1/audit-log?action=reassignments.apply&after=2026-01-01T00:00:00Z"

Reverting Actions

Some audit events support reversal:

POST /api/v1/audit-log/revert

Reversal submits another mutation, requires a valid license, and creates its own audit record. It does not erase the original event or undo application side effects. The UI checks current resource values before presenting supported reversals.

Reassignment Review

Partition reassignment reversal has a separate preview at POST /api/v1/audit-log/revert/preview. Pilot resolves the stored audit event and compares its recorded original and target assignments with current Kafka placement. The review shows changes, partitions already at their original placement and blocked partitions.

Applying requires the stored event selector and reviewFingerprint from the preview. Client-supplied replica arrays do not define reversal targets. Pilot rebuilds the review before submission and rejects changed evidence. Reversal can be blocked by placement drift, unavailable evidence, replica or minimum-ISR constraints, maintenance state, or unfinished reassignment and restart operations.

A nonempty accepted reversal returns HTTP 202 with an executionId. Follow reassignment progress for outcomes rather than treating acceptance as completion. If every partition is already at its original assignment, the response is a no-op. See the OpenAPI specification  for request and response schemas.

Offset Reset Limits

Automatic consumer-group offset reversal is not available in the Audit Log UI. A successful reset event records the previous offsets, requested targets and Kafka request outcome. The offset-reset dialog separately rereads committed offsets, but those observations are not persisted as verified completion in the audit event. This limitation does not mean that the reset failed.

Kafka supports resetting offsets backwards. To restore a previous position, stop the group and review a new exact-offset reset against current offsets and retained log bounds. Records below the retained log start are gone; restoring an offset number cannot recover them or undo processing already performed.

The legacy direct reversal API still accepts previous offsets supplied by the caller. It does not provide the stored-event, drift and retention review described above for reassignments, and should not be treated as a verified undo operation.

Prometheus Metrics

MetricDescription
pilot_audit_events_totalTotal audit events attempted
pilot_audit_events_success_totalEvents successfully delivered
pilot_audit_events_failure_totalFailed deliveries (by error_type)
pilot_audit_delivery_duration_secondsDelivery latency

pilot_audit_events_failure_total counts first delivery attempts that failed (delivery_error, timeout, context_cancelled), which are retried, and events lost for good (dropped). Alert on lost events:

increase(pilot_audit_events_failure_total{error_type="dropped"}[15m]) > 0

Monitor the rate of first attempts that need a retry:

sum(rate(pilot_audit_events_failure_total{error_type!="dropped"}[5m])) / sum(rate(pilot_audit_events_total[5m]))

See Prometheus Metrics for the full metric reference.

Last updated on