---
title: "Audit Logging - Pilot Docs"
description: "Configure audit event logging to Kafka"
url: "https://docs.calinora.io/configuration/audit-logging/"
---

# Audit Logging

Pilot records mutating operations to a Kafka topic for compliance and traceability. When authentication is enabled, audit events capture who performed what action and when.

## How It Works

- Mutating API requests (POST, PUT, DELETE), MCP and assistant tool calls, and changes Pilot makes on its own generate audit events
- Events are produced to the `__pilot_audit_log` Kafka topic
- Events include: action type, user identity, timestamp, request details, and result
- Reassignments are recorded when they are submitted for execution, not when they finish. A failure after submission adds a separate event with status `500`, `error`, `query.moves` and the monitored operation ID (`query.bulkId`, `query.reassignmentId` or `query.executionId`)
- The audit topic is created automatically on startup with `cleanup.policy=delete`

### Event Sources

The `method` field identifies where an event came from:

| `method` | Source | `userId` |
| - | - | - |
| `POST`, `PUT`, `DELETE` | REST API request | Authenticated user (`AUDIT_USER_ID_CLAIM`), empty without authentication |
| `MCP` | MCP client tool call | Approving user (`unknown` without authentication); `requestedBy` holds the requester when it differs |
| `CHAT` | Assistant tool call | Conversation owner who approved the action; `requestedBy` holds `chat-llm:<owner>` |
| `INTERNAL` | [Change Pilot made on its own](https://docs.calinora.io/configuration/audit-logging/#changes-pilot-makes-on-its-own) | `pilot` |

`provider` is the auth provider of the user, `pat` for MCP sessions authenticated with a personal access token, and empty on `INTERNAL` events.

### Large Change Lists

Events that list partition changes, such as `reassignments.apply`, `selfheal.apply` and `broker.maintenance.set`, carry one row per partition in `after`. `ple.trigger` from the REST API lists the partitions whose election ran or failed, with `query.requested`, `query.elected`, `query.notNeeded` and `query.failed` counts. From MCP and the assistant it records only the cluster or topic scope. Lists longer than 500 rows are written as several records that share a `chunkGroup`, with `chunkIndex` and `chunkTotal` set. The API and the Audit Log page return them merged into one event. If some chunks are not available, the merged event sets `chunksMissing` to their number and `after` holds only part of the rows. Consumers reading the topic directly must merge records by `chunkGroup`.

## Changes Pilot Makes on Its Own

Pilot records changes it makes without a request under `userId` `pilot` and `method` `INTERNAL`, with an empty `provider` and no `requestId`. The Audit Log page shows them under **Pilot**.

| Action | Path | Recorded when |
| - | - | - |
| `selfheal.apply` | `/self-healing/<trigger>` | An automatic apply submits reassignments (status `202`) or fails after its submission was recorded (status `500`) |
| `selfheal.cancel` | `/self-healing/<type>/cancel` | The execution window closes on a running self-heal and Pilot cancels its active reassignments |
| `license.update` | `/license/auto-fetch` | The license auto-fetch changes the license state, ID, plan or expiry, or loads a token that is rejected |
| `broker.expire` | `/health-monitor/broker-expiry` | The health monitor stops tracking a broker that stayed unavailable and hosted no replica for `PILOT_BROKER_EXPIRY` |
| `agent.renew-cert` | `/agents/renew-cert` | Pilot signs a renewed client certificate for a Pilot Agent |
| `rolling-restart.*` | `/rolling-restart` | Each step of a [rolling restart](https://docs.calinora.io/features/agents/#broker-lifecycle-management); `query.restartId` identifies the restart |
| `throttles.reconcile` | `/throttles/reconcile` | At startup with `PILOT_THROTTLE_MANAGED=true`, Pilot removes replication throttles left on a broker from an earlier run |
| `reassignments.timeout-revert` | `/reassignment-monitor/timeout` | A tracked reassignment is still running 12 hours after it was last submitted to Kafka, and Pilot cancels it back to its old replicas |

### Automatic Applies

A `selfheal.apply` event is written when a [self-healing](https://docs.calinora.io/features/self-healing/) loop or the [rebalance after a maintenance drain](https://docs.calinora.io/features/maintenance-mode/#automatic-rebalance-after-a-drain) submits its plan for execution, before any move is sent to Kafka. Follow [reassignment progress](https://docs.calinora.io/features/reassignments/#progress-tracking) for the outcome.

| Field | Content |
| - | - |
| `status` | `202` |
| `proposalId` | Proposal the plan came from |
| `brokerId` | Drained broker (`post-maintenance` only) |
| `topic` | Empty, since a plan spans topics |
| `query.trigger` | `activity`, `critical`, `rf` or `post-maintenance` |
| `query.healingType` | Self-healing type whose window, cooldown and cap applied: the trigger, or `activity` for `post-maintenance` |
| `query.moves` | Number of partitions submitted |
| `query.bulkId`, `query.drainBulkId` | Monitored bulk operation, and the drain for `post-maintenance` |
| `after` | Per partition: `topic`, `partitionId`, `oldReplicas` and `oldLeader` read from live metadata at submission, `newReplicas` and `newLeader` as submitted |

An apply that fails after its submission was recorded (status `202`) writes one more event with status `500`, `error` and `query.moves` set to the planned count. This holds even if no move reached Kafka, for example when the replication throttle cannot be applied and the run stops before its first move. An apply refused before submission, for example when every move is filtered out as leaderless during an outage, writes no event, because the audit log records changes to the cluster. The refusal is logged and, for self-healing loops, counted as `pilot_selfhealing_runs_total{result="error"}`.

A dry run (`PILOT_HEAL_DRY_RUN=true`) and a cycle skipped by a gate, such as the time window, cooldown or license, write nothing. Once Pilot begins shutting down it submits no further automatic applies, so none runs unrecorded. Use the [self-healing metrics](https://docs.calinora.io/features/self-healing/#prometheus-metrics) for skips.

`selfheal.apply`, `throttles.reconcile` and `reassignments.timeout-revert` events cannot be reverted from the Audit Log or with the MCP `revert_audit_event` tool.

### Other Changes

- `selfheal.cancel` is written once per cancelled self-heal and carries `proposalId`, `query.healingType`, `query.cancelled` (number of reassignments cancelled) and `query.window` (for example `22-5`). A failed cancel has status `500` and `error`.
- `license.update` carries `state`, `licenseId`, `plan` and `expiresAt` (RFC3339 UTC, or empty) in `before` and `after`. A rejected token has status `500` and `error`.
- `broker.expire` carries the broker in `brokerId` and `before`, and `query.ttl`, `query.unavailableSince` and `query.tombstoned` (whether its record in `__pilot_broker_state` was removed). A failed removal has status `500` and `error`.
- `agent.renew-cert` carries the agent in `requestedBy`, its broker in `brokerId`, `query.commandId`, and the old and new certificate serial and expiry in `before` and `after`. A signing failure has status `500` and `error`.
- `throttles.reconcile` is written once per broker and carries the broker in `brokerId` and the removed values in `before`, keyed by config name (`leader.replication.throttled.rate`, `follower.replication.throttled.rate` and `replica.alter.log.dirs.io.max.bytes.per.second`). Pilot clears a broker whose own dynamic `leader.replication.throttled.rate` or `follower.replication.throttled.rate` is set; cluster-wide defaults and static settings stay in effect. A failed removal has status `500` and `error`, and the values may still apply.
- `reassignments.timeout-revert` is written once per partition and carries `topic`, `partition`, `query.reassignmentId`, `query.bulkId` (for a move in a bulk operation), `query.timeout` and `query.submittedAt` (last submission to Kafka, RFC3339 UTC). `before` holds the target replicas the move was heading to and `after` the old replicas it was cancelled back to, each as `{"replicas": [...]}`. A failed cancel has status `500` and `error`, and the move may still be running.

## Topic Configuration

The audit topic’s name, retention, partition count, and replication factor are all configurable via environment variables:

| Variable | Default | Description |
| - | - | - |
| `AUDIT_ENABLED` | `true` | Enable audit event logging |
| `AUDIT_TOPIC` | `__pilot_audit_log` | Kafka topic name |
| `AUDIT_RETENTION_MS` | `2592000000` (30 days) | Topic retention in milliseconds |
| `AUDIT_PARTITIONS` | `1` | Number of partitions |
| `AUDIT_REPLICATION_FACTOR` | `0` (broker default) | Replication factor for a new topic (0 uses the cluster default, see below) |
| `AUDIT_USER_ID_CLAIM` | `upn` | Identity recorded as `userId`: `upn` (default), `sub` or `email`. Other values fall back to upn, then sub, then email |
| `AUDIT_STORE_MAX_ITEMS` | `10000` | Max audit events kept in memory for the UI |

Pilot creates the topic on startup if it does not exist, and sets `cleanup.policy=delete` and `retention.ms` from `AUDIT_RETENTION_MS` on every startup, also on an existing topic.

With `AUDIT_REPLICATION_FACTOR=0`, a new audit topic uses the brokers’ `default.replication.factor`. If that is `1` on a cluster of two or more brokers, Pilot creates the topic with 3 replicas instead (2 on a two-broker cluster), so a single broker outage does not stop audit delivery. If creation fails, Pilot falls back to the broker default, then to 1. If creation with a set `AUDIT_REPLICATION_FACTOR` fails, Pilot retries with 1. An existing audit topic with one replica on a multi-broker cluster logs a warning at startup; Pilot does not change its replication factor.

## Delivery

Pilot waits up to 5 seconds for the broker to confirm each event. Reassignments from self-healing, MCP and the assistant are recorded before they are sent to Kafka. Other changes are recorded right after Kafka accepts them. Rolling-restart steps are recorded in the background. A write that is not confirmed does not stop the change.

An unconfirmed event is kept in an in-memory retry queue (up to 1,024 events or 32 MiB) and retried with a delay that grows from 1 to 30 seconds. On shutdown, Pilot keeps delivering queued events for up to 10 seconds. An event is dropped and logged as an error when it is still undelivered after 10 minutes or at the end of shutdown, when the queue is full, or when the broker rejects it as too large or invalid. Each dropped event increments `pilot_audit_events_failure_total{error_type="dropped"}`.

## Querying Audit Events

### Via the UI

The Audit Log page in the Pilot dashboard provides a filterable view of all recorded events.

### Via the API

```
GET /api/v1/audit-log
```

Supports query parameters for filtering:

| Parameter | Description |
| - | - |
| `action` | Filter by action type |
| `userId` | Filter by user ID (case-insensitive substring) |
| `after` | Inclusive start timestamp (RFC3339) |
| `before` | Inclusive end timestamp (RFC3339) |
| `search` | Free-text search |
| `limit` | Maximum events to return (default: 500) |
| `offset` | Pagination offset |

Example:

```bash
curl "http://localhost:8080/api/v1/audit-log?action=reassignments.apply&after=2026-01-01T00:00:00Z"
```

## Reverting Actions

Some audit events support reversal:

```
POST /api/v1/audit-log/revert
```

Reversal submits another mutation, requires a valid license, and creates its own audit record. It does not erase the original event or undo application side effects. The UI checks current resource values before presenting supported reversals.

### Reassignment Review

Partition reassignment reversal has a separate preview at `POST /api/v1/audit-log/revert/preview`. Pilot resolves the stored audit event and compares its recorded original and target assignments with current Kafka placement. The review shows changes, partitions already at their original placement and blocked partitions.

Applying requires the stored event selector and `reviewFingerprint` from the preview. Client-supplied replica arrays do not define reversal targets. Pilot rebuilds the review before submission and rejects changed evidence. Reversal can be blocked by placement drift, unavailable evidence, replica or minimum-ISR constraints, maintenance state, or unfinished reassignment and restart operations.

A nonempty accepted reversal returns HTTP 202 with an `executionId`. Follow [reassignment progress](https://docs.calinora.io/features/reassignments/#progress-tracking) for outcomes rather than treating acceptance as completion. If every partition is already at its original assignment, the response is a no-op. See the [OpenAPI specification](https://github.com/julianbergner/Pilot/blob/WIP/api-docs/openapi.yaml) for request and response schemas.

### Offset Reset Limits

**Automatic consumer-group offset reversal is not available in the Audit Log UI.** A successful reset event records the previous offsets, requested targets and Kafka request outcome. The offset-reset dialog separately rereads committed offsets, but those observations are not persisted as verified completion in the audit event. This limitation does not mean that the reset failed.

Kafka supports resetting offsets backwards. To restore a previous position, stop the group and review a new [exact-offset reset](https://docs.calinora.io/features/consumer-groups/#offset-reset) against current offsets and retained log bounds. Records below the retained log start are gone; restoring an offset number cannot recover them or undo processing already performed.

The legacy direct reversal API still accepts previous offsets supplied by the caller. It does not provide the stored-event, drift and retention review described above for reassignments, and should not be treated as a verified undo operation.

## Prometheus Metrics

| Metric | Description |
| - | - |
| `pilot_audit_events_total` | Total audit events attempted |
| `pilot_audit_events_success_total` | Events successfully delivered |
| `pilot_audit_events_failure_total` | Failed deliveries (by `error_type`) |
| `pilot_audit_delivery_duration_seconds` | Delivery latency |

`pilot_audit_events_failure_total` counts first delivery attempts that failed (`delivery_error`, `timeout`, `context_cancelled`), which are retried, and events lost for good (`dropped`). Alert on lost events:

```promql
increase(pilot_audit_events_failure_total{error_type="dropped"}[15m]) > 0
```

Monitor the rate of first attempts that need a retry:

```promql
sum(rate(pilot_audit_events_failure_total{error_type!="dropped"}[5m]))
/ sum(rate(pilot_audit_events_total[5m]))
```

See [Prometheus Metrics](https://docs.calinora.io/reference/metrics/#audit-logger) for the full metric reference.
