Skip to Content
FeaturesAI Chat Assistant

AI Chat Assistant

Pilot includes a built-in AI chat assistant that can help you monitor, diagnose, plan, and operate your Kafka cluster through natural language conversations.

Overview

The chat assistant uses Anthropic, OpenAI, or compatible local models combined with Pilot’s full tool catalog to answer questions about your cluster, run diagnostics, and execute operations - all with an approval gate for any mutations.

The assistant panel stays open after a browser refresh and restores the active conversation. Open or closed state, the history view and pinned page are remembered within the browser tab. Panel width is preserved, and an expanded assistant returns to the same page when collapsed. Conversation messages reload from Pilot; pinned proposal context refreshes from the current API decision.

Configuration

PILOT_CHAT_ENABLED=true PILOT_CHAT_PROVIDER=anthropic # "anthropic" or "openai" PILOT_CHAT_API_KEY=sk-ant-... # Required PILOT_CHAT_MODEL=claude-haiku-4-5-20251001 # Model to use

Anthropic models use the native Messages API.

Optional Settings

VariableDefaultDescription
PILOT_CHAT_BASE_URL""Override API endpoint (self-hosted models, Azure OpenAI)
PILOT_CHAT_OPENAI_APIautoOpenAI API selection: auto, responses, or chat_completions
PILOT_CHAT_MAX_TOKENS4096Maximum response tokens
PILOT_CHAT_MAX_CONTEXT_TOKENS30000Estimated context budget for history, tool catalog, and system prompt (min 4096)
PILOT_CHAT_MAX_TOOL_RESULT_CHARS8000Maximum characters of a tool result kept in the conversation (min 1000)
PILOT_CHAT_APPROVAL_TIMEOUT5mTime window to approve/reject mutations
PILOT_CHAT_RATE_LIMIT20Messages per minute per user
PILOT_CHAT_MAX_CONVERSATIONS10Maximum conversations per user
PILOT_CHAT_CONVERSATION_TTL24hConversation expiry

Pilot does not expose a reasoning-effort setting. Responses and Anthropic requests use the selected model’s defaults. Forced Chat Completions may disable reasoning when the provider requires it for function tools.

OpenAI

Set PILOT_CHAT_PROVIDER=openai. The default PILOT_CHAT_OPENAI_API=auto uses Responses for GPT-5 and GPT-6 model names, including Astra, and Chat Completions for other models. Select responses or chat_completions explicitly when your endpoint requires it. This setting does not affect Anthropic.

PILOT_CHAT_ENABLED=true PILOT_CHAT_PROVIDER=openai PILOT_CHAT_API_KEY=<openai-api-key> PILOT_CHAT_MODEL=gpt-6-astra PILOT_CHAT_OPENAI_API=auto

Azure OpenAI

Set PILOT_CHAT_PROVIDER=openai and point PILOT_CHAT_BASE_URL at your Azure resource with the trailing /openai path segment. Pilot appends /v1/responses or /v1/chat/completions according to the selected API:

PILOT_CHAT_PROVIDER=openai PILOT_CHAT_BASE_URL=https://<resource>.openai.azure.com/openai PILOT_CHAT_API_KEY=<azure-openai-api-key> PILOT_CHAT_MODEL=<deployment-name> # Azure deployment name; must support tool calling PILOT_CHAT_OPENAI_API=responses # For GPT-5/GPT-6 deployment aliases

The deployment must support the selected API. Set responses explicitly for GPT-5/GPT-6 deployments with custom aliases to route directly to Responses.

See Environment Variables - Azure OpenAI for details, including corporate CA setup (SSL_CERT_FILE) for TLS-intercepting proxies.

Local and OpenAI-Compatible Models

Use the openai provider with a server and model that support streaming Chat Completions and function tools:

PILOT_CHAT_ENABLED=true PILOT_CHAT_PROVIDER=openai PILOT_CHAT_BASE_URL=http://localhost:11434 PILOT_CHAT_API_KEY=local # Replace if the server requires authentication PILOT_CHAT_MODEL=<model-name> PILOT_CHAT_OPENAI_API=chat_completions

Use an address reachable from Pilot and omit /v1 from the base URL; Pilot appends /v1/chat/completions. Pilot requires a non-empty API key, so use a placeholder such as local for servers without authentication. Compatibility depends on the server’s API and the model’s tool support.

Capabilities

Monitor & Diagnose

  • “Show me the cluster health”
  • “Which topics have under-replicated partitions?”
  • “What’s the consumer lag for my-consumer-group?”

Plan & Simulate

  • “Generate a rebalancing proposal”
  • “What would happen if broker 3 goes down?”
  • “Analyze the blast radius of losing rack-a”

Operate & Manage (Requires Approval)

  • “Update the retention for topic X to 7 days”
  • “Apply the latest rebalancing proposal”
  • “Put broker 3 into maintenance mode”

Audit & Control

  • “Show me the last 10 audit events”
  • “What changes were made today?”

Mutation Approval

Any operation that modifies cluster state requires explicit user approval:

  1. The assistant proposes an action
  2. The user sees the exact operation and can approve or reject
  3. If approved within the timeout window, the operation executes
  4. The result is reported back in the conversation

This prevents accidental mutations from AI hallucinations or misunderstandings.

Only the user who owns the conversation can approve or reject its pending actions. The audit log records that user as the approver, with method CHAT.

Approved operations run independently of the chat connection, with a two-hour execution limit. Closing chat or changing pages does not cancel them; use the Execution view to cancel a running reassignment. After reconnecting, you can continue the same conversation and ask about the operation’s status.

While processing a message, Pilot shows a progress indicator. Temporary provider failures can retry up to twice before a response starts. Pending approval notifications repeat automatically without creating another action or extending the approval deadline.

If the stream disconnects or stops responding, Pilot shows an error. Chat requests and cluster operations are not automatically resubmitted. Check the conversation and operation status before repeating a mutation.

API Endpoints

MethodPathDescription
GET/api/v1/chat/statusCheck if chat is enabled
POST/api/v1/chat/conversationsCreate a conversation
GET/api/v1/chat/conversationsList conversations
GET/api/v1/chat/conversations/{id}Get conversation with messages
DELETE/api/v1/chat/conversations/{id}Delete a conversation
POST/api/v1/chat/conversations/{id}/messagesSend message (SSE streaming)
POST/api/v1/chat/conversations/{id}/approve/{toolCallId}Approve mutation
POST/api/v1/chat/conversations/{id}/reject/{toolCallId}Reject mutation

Responses are streamed via Server-Sent Events (SSE) for real-time output.

Last updated on