Cassandra consulting · 24×7 support · Remote DBA · Apache Cassandra 4.x and 5.0, K8ssandra, Astra DB, Amazon Keyspaces

Cassandra Consulting Measured in Partition Sizes, Tombstones per Read and Repair Completion, Not Opinions

MinervaDB Cassandra consulting is delivered by senior engineers who have modelled partitions that stayed bounded for years, chosen compaction strategies per table rather than per cluster, and run repair inside gc_grace_seconds on multi-datacenter rings. The practice covers Apache Cassandra 4.0, 4.1 and 5.0 on VMs, bare metal and Kubernetes, and the managed forms, DataStax Astra DB and Amazon Keyspaces, with their divergence written down. Every recommendation names the nodetool output, system table or log line behind it; every change ships with a rollback; and the same team carries the 24×7 watch afterwards. Vendor-neutral by principle: we resell no managed service, so the answer can be self-managed, Kubernetes, Astra, Keyspaces, or a different engine altogether.

4.0 · 4.1 · 5.0Cassandra lines under active MinervaDB Cassandra consulting, 5.0.9 current
900+enterprises supported across every major engine
46cities with on-site delivery presence
15 minS1 acknowledgement, 24×7×365
200+years of combined leadership experience

01 · Why MinervaDB for Cassandra consulting

Cassandra scales linearly, and so do the consequences of a bad partition key

Cassandra delivers what it promises when the data model matches the queries, compaction matches the write pattern and repair completes on time. Most emergencies we are called into are one of those three, discovered at scale.

Vendor-neutral

No Astra or Keyspaces resale, no distribution partnership, no licence commission. Cassandra consulting recommendations come from nodetool tablehistograms, tpstats, the GC log and your invoice, including when the answer is fewer nodes, a managed service, or that this workload never belonged on a wide-column store.

Senior engineers only

Every engagement is led by principal-level Cassandra engineers who have rewritten a wide-partition table under load, recovered resurrected data after a missed repair, and tuned a JVM out of stop-the-world pauses. No junior bench learning on your time.

Real 24×7 operations

Follow-the-sun coverage across APAC, EMEA and the Americas with a senior engineer on watch, S1 acknowledged within 15 minutes, and every action timestamped in a shared incident system. NoSQL consulting and support →

Measurement first

Nothing is claimed until the number has moved: p99 client latency from proxyhistograms, SSTables per read, tombstones scanned, pending compactions, repair completion inside gc_grace, GC pause time, cost per node. Estimates are labelled as estimates.

02 · Cassandra consulting services

Six disciplines around the Cassandra estate, one accountable team

Most Cassandra consulting engagements start with the health check and grow into the discipline the evidence points at. All six are delivered by the same engineers, self-managed or managed.

Cassandra consulting service map: data modelling, performance, compaction and repair, multi-datacenter topology, upgrades, security and cloud around the production Apache Cassandra estate with health check and 24x7 remote DBA

Figure 1. The MinervaDB Cassandra consulting service map: data modelling, performance, compaction and repair, multi-datacenter topology, upgrades, security and cloud, entered through the health check and sustained by 24×7 remote DBA operations.

Data modelling

Cassandra consulting starts at the partition: query-first table design with partition keys chosen from access patterns and cardinality, clustering order for the reads that matter, bounded partitions (p99 under 100 MB), controlled denormalisation, and a review of secondary indexes, materialized views and Storage-Attached Indexes on 5.0 against the workload that will use them.

Performance engineering

Read and write path tuning from nodetool tablestats, tablehistograms, tpstats and request tracing; key and chunk cache sizing; JVM heap, off-heap and garbage-collector selection from the GC log; driver policies (token-aware, speculative retry, paging) and consistency-level trade-offs stated with the latency they buy.

Compaction, repair and tombstones

Compaction strategy per table (STCS, LCS, TWCS, UCS on 5.0) from the write pattern, gc_grace_seconds and TTL design so deletes purge and never resurrect, Reaper-scheduled incremental repair that completes inside the grace window, and tombstone thresholds that surface bad access patterns before they page anyone.

Multi-datacenter topology and DR

Cassandra consulting for availability: NetworkTopologyStrategy with replication factor per datacenter, rack-aware placement, snitch configuration, LOCAL_QUORUM as the default consistency, cross-DC replication verified, and drills for node, rack and datacenter loss that produce measured RPO and RTO alongside timed snapshot restores with Medusa or sstableloader.

Upgrades and lifecycle

Exits from end-of-life 3.x, rolling upgrades from 4.0 and 4.1 to 5.0 one node at a time with SSTable format upgrades and driver and JDK alignment, and a patch calendar so the estate stays inside the project’s three-supported-lines window. Data modernization →

Security, cloud and Kubernetes

TLS on client and internode traffic, role-based access, audit logging, encryption at rest, and the evidence pack for GDPR, HIPAA, PCI DSS and SOC 2; K8ssandra and cass-operator on Kubernetes; the divergence map for DataStax Astra DB and Amazon Keyspaces with the exit path written before migration.

03 · Cassandra consulting method

The read and write path, and the metric that exposes each stage

Every Cassandra latency problem lives at a specific point in the write or read path. Cassandra consulting at MinervaDB names that point from the engine’s own counters before changing anything.

Cassandra consulting read and write path: coordinator, commit log, memtable, SSTables, hints, bloom filter, caches, read repair, with the nodetool metric that exposes each stage

Figure 2. The Cassandra write and read path on a replica node, from commit log and memtable through SSTables, bloom filters, caches and read repair, with the nodetool metric that exposes each stage.

A Cassandra consulting baseline is captured over a window that includes peak traffic, the repair schedule and a compaction cycle: client request latency from nodetool proxyhistograms, per-table read and write latency, SSTables per read and partition-size percentiles from tablehistograms, dropped mutations and pending tasks from tpstats, pending compactions, bloom-filter false-positive ratios, hints in flight, GC pause distribution and disk I/O wait per node. Attribution names the stage: too many SSTables per read, a wide partition, tombstones scanned, a flush storm, a coordinator waiting on a slow replica across the WAN, or a garbage collector sized for the wrong heap.

Changes are made one at a time and are reversible: a compaction strategy change on one table with a documented reversal, a gc_grace adjustment paired with a repair, a JVM flag with its restart scope stated, a driver policy behind a configuration switch. Blast radius and rollback are written before execution.

# Read amplification and partition size, per table (each node)
nodetool tablehistograms my_keyspace.events
# Percentile  Read Lat  Write Lat  SSTables  Partition Size
#             (micros)   (micros)               (bytes)
# 50%           182.79      29.52        2           1916
# 95%          1358.10      88.15        6         105778
# 99%          5839.59     219.34       12        1358102  <- wide
# Max         25109.16    1955.67       20       52066354

# Thread-pool pressure and dropped messages
nodetool tpstats | egrep \
  'Pool Name|Compaction|MemtableFlush|Native-Transport|ReadStage|MutationStage|Dropped|MUTATION|READ'

# Compaction backlog and incremental repair sessions (4.0+)
nodetool compactionstats
nodetool repair_admin list

# Tombstones and partition health per table
nodetool tablestats my_keyspace.events | egrep \
  'tombstones per slice|partition maximum bytes|Bloom filter false ratio|SSTable count'

04 · Cassandra consulting for compaction, repair and tombstones

The three settings that decide whether Cassandra is healthy at 10 TB per node

Compaction strategy, gc_grace_seconds and the repair schedule are coupled. Cassandra consulting at MinervaDB sets them together, per table, from the write pattern and the deletion pattern.

Cassandra consulting compaction, repair and tombstone lifecycle with strategy selection per table (STCS, LCS, TWCS, UCS) and what MinervaDB sets and watches

Figure 3. Cassandra consulting view of the compaction, repair and tombstone lifecycle: from write or delete through flush, compaction, the grace window and repair to purge, with compaction strategy selection per table and what MinervaDB sets and watches.

Strategy per table

Cassandra consulting chooses size-tiered for write-heavy tables without a TTL pattern, levelled for read-heavy and update-heavy tables where bounded SSTables per read matter more than compaction I/O, time-window for time-series with a uniform TTL, and on 5.0 the unified strategy where one tunable replaces the choice. The default is never left in place on a production table without a reason written down.

Repair inside the grace window

Cassandra consulting schedules Reaper-driven incremental repair per keyspace that completes inside gc_grace_seconds, parallelism sized to compaction headroom, a full repair after topology changes, and an alert when a keyspace’s last successful repair approaches the grace boundary. A missed repair is how deleted data comes back.

Tombstones as a design signal

tombstone_warn_threshold and tombstone_failure_threshold tuned so that collection deletes, range deletes and overwrite-heavy access patterns surface in the logs before they surface as timeouts, and the offending table redesigned rather than the threshold raised.

05 · Cassandra consulting for multi-datacenter topology and DR

Availability is a topology and a consistency level, proven by killing nodes on purpose

Cassandra’s availability story holds only when replication factor, rack placement and consistency level agree with each other. Cassandra consulting verifies that they do, then breaks the cluster on a schedule.

Cassandra consulting multi-datacenter topology: two datacenters with rack-aware placement, NetworkTopologyStrategy RF 3, LOCAL_QUORUM, cross-DC replication and the node, rack, datacenter and restore drills

Figure 4. Cassandra consulting multi-datacenter topology: two datacenters with three racks each, NetworkTopologyStrategy with RF 3 per datacenter, LOCAL_QUORUM reads and writes, asynchronous cross-datacenter replication, and the node, rack, datacenter and restore drills that prove it.

Drill What is injected What is measured Evidence artefact
Node loss One node per rack stopped or decommissioned p99 latency and unavailable errors during and after; hinted handoff and repair catch-up time proxyhistograms before, during, after; nodetool status and hint metrics
Rack loss All nodes in one rack stopped LOCAL_QUORUM still satisfied with RF 3 across three racks; snitch and rack assignment behave as designed nodetool describecluster, getendpoints per sample key
Datacenter loss Application traffic pointed at the second datacenter Staleness at switch (RPO) from cross-DC replication lag; time to serve (RTO) including driver failover Driver metrics, per-DC latency, replication lag samples
Restore Keyspace restored to a scratch cluster from snapshot plus incrementals with Medusa or sstableloader Elapsed restore time on production-size data; row-count and checksum parity Restore report with the observed RTO written into the runbook

06 · Cassandra consulting for versions and upgrades

Cassandra versions we plan against, and how a rolling upgrade is run

Verified against the Apache Cassandra project’s release and support policy on 28 September 2026. The project supports the latest three release lines; 3.0 and 3.11 have been end of life since 5 September 2024.

Cassandra consulting version lifecycle table for Apache Cassandra 3.x, 4.0, 4.1, 5.0 and 6.0 with the rolling upgrade method

Figure 5. Apache Cassandra version lifecycle as verified on 28 September 2026 and the rolling upgrade method: pre-flight, one node at a time, SSTable upgrade, verification against the baseline.

Release line Latest patch Status What Cassandra consulting does about it
Cassandra 3.0 and 3.11 3.0.32, 3.11.19 (February 2025) End of life since 5 September 2024 Upgrade 3.11 to 4.0 or 4.1 first, then to 5.0; SSTable format, driver and JDK migration rehearsed on a copy
Cassandra 4.0 4.0.21 (August 2026) Supported, oldest supported line Plan the move to 4.1 or 5.0; it leaves support when the next major ships
Cassandra 4.1 4.1.12 (August 2026) Supported, stable production target Stay current on patches; evaluate 5.0 for Storage-Attached Indexes, unified compaction, trie memtables and vector search
Cassandra 5.0 5.0.9 (August 2026) Current major, supported Target for new builds and upgrades from 4.x on JDK 11 or 17
Cassandra 6.0 Pre-release Accord transactions in development Tracked; no customer estate moves until GA and a patch series exist

Sources: the Apache Cassandra 5.0 announcement and the project’s download and support pages. Confirm the exact patch release at engagement start; the project ships patches across all three supported lines together.

07 · Cassandra consulting across platforms

Self-managed, Kubernetes, Astra DB or Keyspaces: decided per workload with the exit path written first

Each platform removes a different share of the operational work and takes a different share of control. Cassandra consulting at MinervaDB maps the divergence before recommending one.

Cassandra consulting platform decision map: self-managed, K8ssandra on Kubernetes, DataStax Astra DB and Amazon Keyspaces with divergence and exit paths

Figure 6. Cassandra deployment options and the managed-service divergence map: self-managed, K8ssandra on Kubernetes, DataStax Astra DB and Amazon Keyspaces, with what each removes and constrains, and how MinervaDB decides.

Self-managed

Cassandra consulting for self-managed estates: Apache Cassandra 4.1 or 5.0 on VMs or bare metal with the JDK aligned, full control of compaction, repair, snitch and JVM, and full ownership of the pager, which is where MinervaDB remote DBA usually comes in.

Kubernetes

K8ssandra or cass-operator managing pods and persistent volumes with Reaper and Medusa built in; storage class, anti-affinity and disruption budgets decide availability, and a restore drill proves them. Databases on Kubernetes →

DataStax Astra DB

Serverless Cassandra with vector search, no nodetool, no JVM and no compaction tuning, priced per request and storage. The exit runs through DSBulk or CDC, and the data model must stay Cassandra-shaped to keep that door open.

Amazon Keyspaces

CQL-compatible managed service that is not Cassandra internals: lightweight-transaction semantics, type support and capacity modes differ, and the feature gaps are mapped before any migration. Exit through DSBulk with the compatibility list in hand.

08 · Cassandra health check and performance audit

The fixed-scope entry point to Cassandra consulting at MinervaDB

Read-only, evidence-based and delivered as findings your own engineers can verify. No change is made to production during the audit.

What is reviewed

Cassandra consulting audits cover version and patch currency; topology, snitch and rack placement against replication factor and consistency levels; data model and partition-size percentiles; compaction strategy per table and pending compactions; repair schedule against gc_grace; tombstone ratios; JVM and GC behaviour; cache hit rates; hints and dropped messages; backup and restore evidence; security posture; and, on managed services, tier fit and cost per unit of work.

What you receive

Findings ranked P0 to P2, each with the observation, the nodetool, system-table or log evidence, the recommended change with its rollback and the metric it is expected to move; a prioritised remediation plan; and a versioned report your team keeps whether or not MinervaDB does the remediation. Findings typically arrive within days of read-only access.

Standing caveat: every recommendation on this page is tested on a non-production ring against production-representative data before it is applied to production, with a verified snapshot taken first and a disaster-recovery posture exercised by a timed restore.

09 · Cassandra consulting rates

Transparent Cassandra consulting rates

MinervaDB operates as a virtual corporation, so you pay for senior engineering hours rather than office overhead. Emergency support is included in every retainer.

Remote consulting · US $300 per hour

Hourly Cassandra consulting delivered remotely worldwide: data modelling, compaction and repair engineering, multi-datacenter design, JVM and read-path tuning, upgrade and migration planning, technical advisory, available on short notice.

Remote DBA retainer · from US $4,500 per quarter

Ongoing Cassandra DBA operations with a four-hour monthly minimum: 24×7 monitoring and alerting, incident response with root-cause analysis, repair and compaction calendar, backup verification and restore drills, patching and version upgrades, node and rack health, and a monthly performance and SLO report. Emergency support always included.

On-site consulting · US $500 per hour

Architecture and data-modelling workshops, multi-datacenter design sessions, executive and engineering briefings, implementation and cutover presence, team training and on-site incident response, available in 46 cities worldwide. Travel applies.

10 · FAQ

Cassandra consulting questions we are asked most

Short answers to what engineering leaders ask before the first call.

What does MinervaDB Cassandra consulting cover?

Data modelling and partition design, read and write path performance engineering, compaction strategy and repair scheduling, tombstone control, multi-datacenter topology and consistency design, upgrades from 3.x and 4.x to 5.0, security and compliance, and platform decisions across self-managed, K8ssandra on Kubernetes, DataStax Astra DB and Amazon Keyspaces. Every finding cites the nodetool output, system table or log line behind it, and every change ships with a rollback.

Do you provide 24×7 Cassandra support?

Yes. A senior Cassandra engineer is on watch across APAC, EMEA and the Americas with S1 (production outage, data-integrity event or security incident) acknowledged within 15 minutes, S2 within 12 hours, S3 within 24 hours and S4 within 48 hours, with a repair and compaction calendar, patching, restore drills and a monthly SLO report included in remote DBA retainers.

How do you ensure Cassandra high availability and disaster recovery?

By making replication factor, rack placement and consistency level agree, NetworkTopologyStrategy with RF 3 across three racks per datacenter and LOCAL_QUORUM as the default, then proving it: node, rack and datacenter loss drills that record p99 latency and unavailable errors, cross-datacenter replication lag as the RPO, driver failover time as the RTO, and a quarterly timed keyspace restore with Medusa or sstableloader.

Why does deleted data come back in Cassandra, and how do you prevent it?

A delete writes a tombstone that must reach every replica before it is purged after gc_grace_seconds. If repair does not complete inside that window on every node, a replica that missed the delete resurrects the row. Cassandra consulting sets gc_grace and TTLs per table, schedules incremental repair with Reaper to finish inside the window, and alerts when a keyspace's last successful repair approaches the boundary.

Which compaction strategy should we use?

It depends on the table, not the cluster: size-tiered for write-heavy tables without TTL, levelled for read-heavy and update-heavy tables, time-window for time-series with a uniform TTL, and on Cassandra 5.0 the unified strategy where one parameter tunes between tiered and levelled. The choice is made from the write pattern, SSTables per read and compaction headroom measured on your nodes.

Which Cassandra versions and platforms do you support?

Apache Cassandra 4.0, 4.1 and 5.0, the three lines the project supports, on VMs, bare metal and Kubernetes with K8ssandra or cass-operator, plus DataStax Astra DB and Amazon Keyspaces with their divergence from open-source Cassandra mapped. Estates still on 3.0 or 3.11, end of life since September 2024, are upgraded through 4.x to 5.0 with SSTable, driver and JDK migration rehearsed first.

Should we move to Astra DB or Amazon Keyspaces?

Sometimes. Astra removes nodetool, the JVM and compaction tuning and prices per request; Keyspaces is CQL-compatible but not Cassandra internals, with lightweight-transaction and type differences. MinervaDB resells neither, so the assessment compares cost per unit of work at your measured p99 over twelve months, maps the feature gaps, and writes the exit path before any migration.

How much does Cassandra consulting cost?

Remote Cassandra consulting is US $300 per hour and on-site consulting is US $500 per hour with travel applying. Remote DBA retainers start at US $4,500 per quarter with a four-hour monthly minimum and emergency support always included. There are no advance payments and no long-term lock-in.

Talk to a senior Cassandra consulting engineer

Bring nodetool tablehistograms and tpstats from a peak hour, the last repair run’s Reaper report, a GC log, and the current infrastructure or managed-service invoice to the first call. We will tell you which stage of the path is the constraint, what moving it is worth, and what we would change first.