AI Chat Assistant
Pilot includes a built-in AI chat assistant that can help you monitor, diagnose, plan, and operate your Kafka cluster through natural language conversations.
Overview
The chat assistant uses Anthropic, OpenAI, or compatible local models combined with Pilot’s full tool catalog to answer questions about your cluster, run diagnostics, and execute operations - all with an approval gate for any mutations.
The assistant panel stays open after a browser refresh and restores the active conversation. Open or closed state, the history view and pinned page are remembered within the browser tab. Panel width is preserved, and an expanded assistant returns to the same page when collapsed. Conversation messages reload from Pilot; pinned proposal context refreshes from the current API decision.
Configuration
PILOT_CHAT_ENABLED=true
PILOT_CHAT_PROVIDER=anthropic # "anthropic" or "openai"
PILOT_CHAT_API_KEY=sk-ant-... # Required
PILOT_CHAT_MODEL=claude-haiku-4-5-20251001 # Model to useAnthropic models use the native Messages API.
Optional Settings
| Variable | Default | Description |
|---|---|---|
PILOT_CHAT_BASE_URL | "" | Override API endpoint (self-hosted models, Azure OpenAI) |
PILOT_CHAT_OPENAI_API | auto | OpenAI API selection: auto, responses, or chat_completions |
PILOT_CHAT_MAX_TOKENS | 4096 | Maximum response tokens |
PILOT_CHAT_MAX_CONTEXT_TOKENS | 30000 | Estimated context budget for history, tool catalog, and system prompt (min 4096) |
PILOT_CHAT_MAX_TOOL_RESULT_CHARS | 8000 | Maximum characters of a tool result kept in the conversation (min 1000) |
PILOT_CHAT_APPROVAL_TIMEOUT | 5m | Time window to approve/reject mutations |
PILOT_CHAT_RATE_LIMIT | 20 | Messages per minute per user |
PILOT_CHAT_MAX_CONVERSATIONS | 10 | Maximum conversations per user |
PILOT_CHAT_CONVERSATION_TTL | 24h | Conversation expiry |
Pilot does not expose a reasoning-effort setting. Responses and Anthropic requests use the selected model’s defaults. Forced Chat Completions may disable reasoning when the provider requires it for function tools.
OpenAI
Set PILOT_CHAT_PROVIDER=openai. The default PILOT_CHAT_OPENAI_API=auto uses Responses for GPT-5 and GPT-6 model names, including Astra, and Chat Completions for other models. Select responses or chat_completions explicitly when your endpoint requires it. This setting does not affect Anthropic.
PILOT_CHAT_ENABLED=true
PILOT_CHAT_PROVIDER=openai
PILOT_CHAT_API_KEY=<openai-api-key>
PILOT_CHAT_MODEL=gpt-6-astra
PILOT_CHAT_OPENAI_API=autoAzure OpenAI
Set PILOT_CHAT_PROVIDER=openai and point PILOT_CHAT_BASE_URL at your Azure resource with the trailing /openai path segment. Pilot appends /v1/responses or /v1/chat/completions according to the selected API:
PILOT_CHAT_PROVIDER=openai
PILOT_CHAT_BASE_URL=https://<resource>.openai.azure.com/openai
PILOT_CHAT_API_KEY=<azure-openai-api-key>
PILOT_CHAT_MODEL=<deployment-name> # Azure deployment name; must support tool calling
PILOT_CHAT_OPENAI_API=responses # For GPT-5/GPT-6 deployment aliasesThe deployment must support the selected API. Set responses explicitly for GPT-5/GPT-6 deployments with custom aliases to route directly to Responses.
See Environment Variables - Azure OpenAI for details, including corporate CA setup (SSL_CERT_FILE) for TLS-intercepting proxies.
Local and OpenAI-Compatible Models
Use the openai provider with a server and model that support streaming Chat Completions and function tools:
PILOT_CHAT_ENABLED=true
PILOT_CHAT_PROVIDER=openai
PILOT_CHAT_BASE_URL=http://localhost:11434
PILOT_CHAT_API_KEY=local # Replace if the server requires authentication
PILOT_CHAT_MODEL=<model-name>
PILOT_CHAT_OPENAI_API=chat_completionsUse an address reachable from Pilot and omit /v1 from the base URL; Pilot appends /v1/chat/completions. Pilot requires a non-empty API key, so use a placeholder such as local for servers without authentication. Compatibility depends on the server’s API and the model’s tool support.
Capabilities
Monitor & Diagnose
- “Show me the cluster health”
- “Which topics have under-replicated partitions?”
- “What’s the consumer lag for my-consumer-group?”
Plan & Simulate
- “Generate a rebalancing proposal”
- “What would happen if broker 3 goes down?”
- “Analyze the blast radius of losing rack-a”
Operate & Manage (Requires Approval)
- “Update the retention for topic X to 7 days”
- “Apply the latest rebalancing proposal”
- “Put broker 3 into maintenance mode”
Audit & Control
- “Show me the last 10 audit events”
- “What changes were made today?”
Mutation Approval
Any operation that modifies cluster state requires explicit user approval:
- The assistant proposes an action
- The user sees the exact operation and can approve or reject
- If approved within the timeout window, the operation executes
- The result is reported back in the conversation
This prevents accidental mutations from AI hallucinations or misunderstandings.
Only the user who owns the conversation can approve or reject its pending actions. The audit log records that user as the approver, with method CHAT.
Approved operations run independently of the chat connection, with a two-hour execution limit. Closing chat or changing pages does not cancel them; use the Execution view to cancel a running reassignment. After reconnecting, you can continue the same conversation and ask about the operation’s status.
While processing a message, Pilot shows a progress indicator. Temporary provider failures can retry up to twice before a response starts. Pending approval notifications repeat automatically without creating another action or extending the approval deadline.
If the stream disconnects or stops responding, Pilot shows an error. Chat requests and cluster operations are not automatically resubmitted. Check the conversation and operation status before repeating a mutation.
API Endpoints
| Method | Path | Description |
|---|---|---|
GET | /api/v1/chat/status | Check if chat is enabled |
POST | /api/v1/chat/conversations | Create a conversation |
GET | /api/v1/chat/conversations | List conversations |
GET | /api/v1/chat/conversations/{id} | Get conversation with messages |
DELETE | /api/v1/chat/conversations/{id} | Delete a conversation |
POST | /api/v1/chat/conversations/{id}/messages | Send message (SSE streaming) |
POST | /api/v1/chat/conversations/{id}/approve/{toolCallId} | Approve mutation |
POST | /api/v1/chat/conversations/{id}/reject/{toolCallId} | Reject mutation |
Responses are streamed via Server-Sent Events (SSE) for real-time output.