Is your Cassandra cluster lagging? Before you smash the “add more nodes” button and balloon your cloud bill, stop.

Diagnosing performance drops in a masterless, distributed database like Apache Cassandra is notoriously difficult. Too often, engineering teams throw expensive CPU and RAM at latency spikes, only to find the root cause was a misconfigured compaction strategy, JVM garbage collection pauses, or a silent buildup of query-killing tombstones.

In this session, we’re sharing the data-driven insights and diagnostic frameworks we use to keep thousands of production nodes running flawlessly around the clock.

What you’ll learn:

  • How to decode read/write metrics to isolate the exact bottleneck (before spending a dime on hardware).
  • Why simply increasing your JVM heap size often backfires and degrades cluster performance.
  • How to hunt down and safely clear tombstone-heavy tables that are dragging your queries to a crawl.
  • How rigorous maintenance strategies, early-warning indicators, and active drift-detection can shift your focus from reactive firefighting to calm, predictable operations.

Maximize your Cassandra clusters’ reliability and give your engineering team the freedom to innovate. Join us to discover how to run a highly optimized, lightning-fast, and effortlessly predictable Cassandra fleet.

Speaker: Steve Richardson, Cloud Solution Architect, NetApp Instaclustr