Skip to Content
FeaturesWhat-If Simulator

What-If Simulator

The what-if simulator lets you test hypothetical cluster scenarios without making any changes. Understand the impact of broker failures, rack outages, traffic changes, and configuration adjustments before they happen.

What-If Simulation

Simulate scenarios and see how the cluster would respond:

  • Broker failures - what happens if one or more brokers go offline?
  • Rack failures - what is the impact of losing an entire rack?
  • Adding brokers - how would new brokers improve balance?
  • RF changes - what happens if you change the replication factor for specific topics?
  • Traffic multipliers - how would traffic increases affect cluster balance?

API

POST /api/v1/what-if/simulate
{ "brokerFailures": [3], "rackFailures": [], "addBrokers": [{"id": 10, "rack": "rack-d"}], "topicRFChanges": {"my-topic": 3}, "trafficMultipliers": {"high-throughput-topic": 2.0} }

The response includes the simulated cluster state, affected partitions, and balance metrics. Balance metrics are relative variances. The modeled rebalance follows proposal rules: metrics with no traffic are not balanced, and with PILOT_BALANCE_FLOOR set, neither are differences below the minimum difference.

For topics whose cleanup.policy contains compact, the simulation weighs each partition at its typical size over its log-cleaner cycle, the same size proposals balance. diskVariance, diskBytesBefore and diskBytesAfter can therefore differ from the Brokers view on brokers holding compacted replicas. See Compacted Topics.

Minimum-ISR Availability

Producer-write impact requires each affected topic’s effective min.insync.replicas. If the required configuration is unavailable, the simulation returns 503 Service Unavailable instead of assuming a value of 1. Restore configuration access and retry before relying on a producer-availability result.

The rack health assessment measures whether assigned replicas can tolerate losing one rack. It does not establish that those replicas are currently in sync. What-if producer-write impact also considers the ISR remaining after the simulated failure; a placement that tolerates a rack loss can still have a current ISR problem.

Blast Radius Analysis

Analyze the blast radius of losing a specific entity:

POST /api/v1/blast-radius/analyze
{ "entityType": "broker", "entityId": "3" }

The response shows which partitions, topics, and consumer groups would be affected. Disk totals (diskBytes, totalDiskBytes) are measured sizes of the data on the affected replicas, compacted topics included. The failureSimulation block uses the same cleaner-cycle size as Simulate, so its diskBytesBefore, diskBytesAfter and diskVariance can differ from these totals on brokers holding compacted replicas.

Entity Types

  • Broker - impact of a broker going offline
  • Rack - impact of an entire rack failure
  • Topic - dependencies and consumers affected by a topic issue

Use Cases

  • Capacity planning - test adding brokers before provisioning hardware
  • Disaster recovery - understand rack failure impact and plan accordingly
  • Change validation - verify RF changes or topic configurations before applying
  • On-call preparation - pre-analyze failure scenarios for runbooks
Last updated on