Cassandra consulting · 24×7 support · Remote DBA · Apache Cassandra 4.x and 5.0, K8ssandra, Astra DB, Amazon Keyspaces
Cassandra Consulting Measured in Partition Sizes, Tombstones per Read and Repair Completion, Not Opinions
MinervaDB Cassandra consulting is delivered by senior engineers who have modelled partitions that stayed bounded for years, chosen compaction strategies per table rather than per cluster, and run repair inside gc_grace_seconds on multi-datacenter rings. The practice covers Apache Cassandra 4.0, 4.1 and 5.0 on VMs, bare metal and Kubernetes, and the managed forms, DataStax Astra DB and Amazon Keyspaces, with their divergence written down. Every recommendation names the nodetool output, system table or log line behind it; every change ships with a rollback; and the same team carries the 24×7 watch afterwards. Vendor-neutral by principle: we resell no managed service, so the answer can be self-managed, Kubernetes, Astra, Keyspaces, or a different engine altogether.
01 · Why MinervaDB for Cassandra consulting
Cassandra scales linearly, and so do the consequences of a bad partition key
Cassandra delivers what it promises when the data model matches the queries, compaction matches the write pattern and repair completes on time. Most emergencies we are called into are one of those three, discovered at scale.
Vendor-neutral
No Astra or Keyspaces resale, no distribution partnership, no licence commission. Cassandra consulting recommendations come from nodetool tablehistograms, tpstats, the GC log and your invoice, including when the answer is fewer nodes, a managed service, or that this workload never belonged on a wide-column store.
Senior engineers only
Every engagement is led by principal-level Cassandra engineers who have rewritten a wide-partition table under load, recovered resurrected data after a missed repair, and tuned a JVM out of stop-the-world pauses. No junior bench learning on your time.
Real 24×7 operations
Follow-the-sun coverage across APAC, EMEA and the Americas with a senior engineer on watch, S1 acknowledged within 15 minutes, and every action timestamped in a shared incident system. NoSQL consulting and support →
Measurement first
Nothing is claimed until the number has moved: p99 client latency from proxyhistograms, SSTables per read, tombstones scanned, pending compactions, repair completion inside gc_grace, GC pause time, cost per node. Estimates are labelled as estimates.
02 · Cassandra consulting services
Six disciplines around the Cassandra estate, one accountable team
Most Cassandra consulting engagements start with the health check and grow into the discipline the evidence points at. All six are delivered by the same engineers, self-managed or managed.
Figure 1. The MinervaDB Cassandra consulting service map: data modelling, performance, compaction and repair, multi-datacenter topology, upgrades, security and cloud, entered through the health check and sustained by 24×7 remote DBA operations.
Data modelling
Cassandra consulting starts at the partition: query-first table design with partition keys chosen from access patterns and cardinality, clustering order for the reads that matter, bounded partitions (p99 under 100 MB), controlled denormalisation, and a review of secondary indexes, materialized views and Storage-Attached Indexes on 5.0 against the workload that will use them.
Performance engineering
Read and write path tuning from nodetool tablestats, tablehistograms, tpstats and request tracing; key and chunk cache sizing; JVM heap, off-heap and garbage-collector selection from the GC log; driver policies (token-aware, speculative retry, paging) and consistency-level trade-offs stated with the latency they buy.
Compaction, repair and tombstones
Compaction strategy per table (STCS, LCS, TWCS, UCS on 5.0) from the write pattern, gc_grace_seconds and TTL design so deletes purge and never resurrect, Reaper-scheduled incremental repair that completes inside the grace window, and tombstone thresholds that surface bad access patterns before they page anyone.
Multi-datacenter topology and DR
Cassandra consulting for availability: NetworkTopologyStrategy with replication factor per datacenter, rack-aware placement, snitch configuration, LOCAL_QUORUM as the default consistency, cross-DC replication verified, and drills for node, rack and datacenter loss that produce measured RPO and RTO alongside timed snapshot restores with Medusa or sstableloader.
Upgrades and lifecycle
Exits from end-of-life 3.x, rolling upgrades from 4.0 and 4.1 to 5.0 one node at a time with SSTable format upgrades and driver and JDK alignment, and a patch calendar so the estate stays inside the project’s three-supported-lines window. Data modernization →
Security, cloud and Kubernetes
TLS on client and internode traffic, role-based access, audit logging, encryption at rest, and the evidence pack for GDPR, HIPAA, PCI DSS and SOC 2; K8ssandra and cass-operator on Kubernetes; the divergence map for DataStax Astra DB and Amazon Keyspaces with the exit path written before migration.
03 · Cassandra consulting method
The read and write path, and the metric that exposes each stage
Every Cassandra latency problem lives at a specific point in the write or read path. Cassandra consulting at MinervaDB names that point from the engine’s own counters before changing anything.
Figure 2. The Cassandra write and read path on a replica node, from commit log and memtable through SSTables, bloom filters, caches and read repair, with the nodetool metric that exposes each stage.
A Cassandra consulting baseline is captured over a window that includes peak traffic, the repair schedule and a compaction cycle: client request latency from nodetool proxyhistograms, per-table read and write latency, SSTables per read and partition-size percentiles from tablehistograms, dropped mutations and pending tasks from tpstats, pending compactions, bloom-filter false-positive ratios, hints in flight, GC pause distribution and disk I/O wait per node. Attribution names the stage: too many SSTables per read, a wide partition, tombstones scanned, a flush storm, a coordinator waiting on a slow replica across the WAN, or a garbage collector sized for the wrong heap.
Changes are made one at a time and are reversible: a compaction strategy change on one table with a documented reversal, a gc_grace adjustment paired with a repair, a JVM flag with its restart scope stated, a driver policy behind a configuration switch. Blast radius and rollback are written before execution.
# Read amplification and partition size, per table (each node) nodetool tablehistograms my_keyspace.events # Percentile Read Lat Write Lat SSTables Partition Size # (micros) (micros) (bytes) # 50% 182.79 29.52 2 1916 # 95% 1358.10 88.15 6 105778 # 99% 5839.59 219.34 12 1358102 <- wide # Max 25109.16 1955.67 20 52066354 # Thread-pool pressure and dropped messages nodetool tpstats | egrep \ 'Pool Name|Compaction|MemtableFlush|Native-Transport|ReadStage|MutationStage|Dropped|MUTATION|READ' # Compaction backlog and incremental repair sessions (4.0+) nodetool compactionstats nodetool repair_admin list # Tombstones and partition health per table nodetool tablestats my_keyspace.events | egrep \ 'tombstones per slice|partition maximum bytes|Bloom filter false ratio|SSTable count'
04 · Cassandra consulting for compaction, repair and tombstones
The three settings that decide whether Cassandra is healthy at 10 TB per node
Compaction strategy, gc_grace_seconds and the repair schedule are coupled. Cassandra consulting at MinervaDB sets them together, per table, from the write pattern and the deletion pattern.
Figure 3. Cassandra consulting view of the compaction, repair and tombstone lifecycle: from write or delete through flush, compaction, the grace window and repair to purge, with compaction strategy selection per table and what MinervaDB sets and watches.
Strategy per table
Cassandra consulting chooses size-tiered for write-heavy tables without a TTL pattern, levelled for read-heavy and update-heavy tables where bounded SSTables per read matter more than compaction I/O, time-window for time-series with a uniform TTL, and on 5.0 the unified strategy where one tunable replaces the choice. The default is never left in place on a production table without a reason written down.
Repair inside the grace window
Cassandra consulting schedules Reaper-driven incremental repair per keyspace that completes inside gc_grace_seconds, parallelism sized to compaction headroom, a full repair after topology changes, and an alert when a keyspace’s last successful repair approaches the grace boundary. A missed repair is how deleted data comes back.
Tombstones as a design signal
tombstone_warn_threshold and tombstone_failure_threshold tuned so that collection deletes, range deletes and overwrite-heavy access patterns surface in the logs before they surface as timeouts, and the offending table redesigned rather than the threshold raised.
05 · Cassandra consulting for multi-datacenter topology and DR
Availability is a topology and a consistency level, proven by killing nodes on purpose
Cassandra’s availability story holds only when replication factor, rack placement and consistency level agree with each other. Cassandra consulting verifies that they do, then breaks the cluster on a schedule.
Figure 4. Cassandra consulting multi-datacenter topology: two datacenters with three racks each, NetworkTopologyStrategy with RF 3 per datacenter, LOCAL_QUORUM reads and writes, asynchronous cross-datacenter replication, and the node, rack, datacenter and restore drills that prove it.
| Drill | What is injected | What is measured | Evidence artefact |
|---|---|---|---|
| Node loss | One node per rack stopped or decommissioned | p99 latency and unavailable errors during and after; hinted handoff and repair catch-up time | proxyhistograms before, during, after; nodetool status and hint metrics |
| Rack loss | All nodes in one rack stopped | LOCAL_QUORUM still satisfied with RF 3 across three racks; snitch and rack assignment behave as designed | nodetool describecluster, getendpoints per sample key |
| Datacenter loss | Application traffic pointed at the second datacenter | Staleness at switch (RPO) from cross-DC replication lag; time to serve (RTO) including driver failover | Driver metrics, per-DC latency, replication lag samples |
| Restore | Keyspace restored to a scratch cluster from snapshot plus incrementals with Medusa or sstableloader |
Elapsed restore time on production-size data; row-count and checksum parity | Restore report with the observed RTO written into the runbook |
06 · Cassandra consulting for versions and upgrades
Cassandra versions we plan against, and how a rolling upgrade is run
Verified against the Apache Cassandra project’s release and support policy on 28 September 2026. The project supports the latest three release lines; 3.0 and 3.11 have been end of life since 5 September 2024.
Figure 5. Apache Cassandra version lifecycle as verified on 28 September 2026 and the rolling upgrade method: pre-flight, one node at a time, SSTable upgrade, verification against the baseline.
| Release line | Latest patch | Status | What Cassandra consulting does about it |
|---|---|---|---|
| Cassandra 3.0 and 3.11 | 3.0.32, 3.11.19 (February 2025) | End of life since 5 September 2024 | Upgrade 3.11 to 4.0 or 4.1 first, then to 5.0; SSTable format, driver and JDK migration rehearsed on a copy |
| Cassandra 4.0 | 4.0.21 (August 2026) | Supported, oldest supported line | Plan the move to 4.1 or 5.0; it leaves support when the next major ships |
| Cassandra 4.1 | 4.1.12 (August 2026) | Supported, stable production target | Stay current on patches; evaluate 5.0 for Storage-Attached Indexes, unified compaction, trie memtables and vector search |
| Cassandra 5.0 | 5.0.9 (August 2026) | Current major, supported | Target for new builds and upgrades from 4.x on JDK 11 or 17 |
| Cassandra 6.0 | Pre-release | Accord transactions in development | Tracked; no customer estate moves until GA and a patch series exist |
Sources: the Apache Cassandra 5.0 announcement and the project’s download and support pages. Confirm the exact patch release at engagement start; the project ships patches across all three supported lines together.
07 · Cassandra consulting across platforms
Self-managed, Kubernetes, Astra DB or Keyspaces: decided per workload with the exit path written first
Each platform removes a different share of the operational work and takes a different share of control. Cassandra consulting at MinervaDB maps the divergence before recommending one.
Figure 6. Cassandra deployment options and the managed-service divergence map: self-managed, K8ssandra on Kubernetes, DataStax Astra DB and Amazon Keyspaces, with what each removes and constrains, and how MinervaDB decides.
Self-managed
Cassandra consulting for self-managed estates: Apache Cassandra 4.1 or 5.0 on VMs or bare metal with the JDK aligned, full control of compaction, repair, snitch and JVM, and full ownership of the pager, which is where MinervaDB remote DBA usually comes in.
Kubernetes
K8ssandra or cass-operator managing pods and persistent volumes with Reaper and Medusa built in; storage class, anti-affinity and disruption budgets decide availability, and a restore drill proves them. Databases on Kubernetes →
DataStax Astra DB
Serverless Cassandra with vector search, no nodetool, no JVM and no compaction tuning, priced per request and storage. The exit runs through DSBulk or CDC, and the data model must stay Cassandra-shaped to keep that door open.
Amazon Keyspaces
CQL-compatible managed service that is not Cassandra internals: lightweight-transaction semantics, type support and capacity modes differ, and the feature gaps are mapped before any migration. Exit through DSBulk with the compatibility list in hand.
08 · Cassandra health check and performance audit
The fixed-scope entry point to Cassandra consulting at MinervaDB
Read-only, evidence-based and delivered as findings your own engineers can verify. No change is made to production during the audit.
What is reviewed
Cassandra consulting audits cover version and patch currency; topology, snitch and rack placement against replication factor and consistency levels; data model and partition-size percentiles; compaction strategy per table and pending compactions; repair schedule against gc_grace; tombstone ratios; JVM and GC behaviour; cache hit rates; hints and dropped messages; backup and restore evidence; security posture; and, on managed services, tier fit and cost per unit of work.
What you receive
Findings ranked P0 to P2, each with the observation, the nodetool, system-table or log evidence, the recommended change with its rollback and the metric it is expected to move; a prioritised remediation plan; and a versioned report your team keeps whether or not MinervaDB does the remediation. Findings typically arrive within days of read-only access.
Standing caveat: every recommendation on this page is tested on a non-production ring against production-representative data before it is applied to production, with a verified snapshot taken first and a disaster-recovery posture exercised by a timed restore.
09 · Cassandra consulting rates
Transparent Cassandra consulting rates
MinervaDB operates as a virtual corporation, so you pay for senior engineering hours rather than office overhead. Emergency support is included in every retainer.
Remote consulting · US $300 per hour
Hourly Cassandra consulting delivered remotely worldwide: data modelling, compaction and repair engineering, multi-datacenter design, JVM and read-path tuning, upgrade and migration planning, technical advisory, available on short notice.
Remote DBA retainer · from US $4,500 per quarter
Ongoing Cassandra DBA operations with a four-hour monthly minimum: 24×7 monitoring and alerting, incident response with root-cause analysis, repair and compaction calendar, backup verification and restore drills, patching and version upgrades, node and rack health, and a monthly performance and SLO report. Emergency support always included.
On-site consulting · US $500 per hour
Architecture and data-modelling workshops, multi-datacenter design sessions, executive and engineering briefings, implementation and cutover presence, team training and on-site incident response, available in 46 cities worldwide. Travel applies.
10 · FAQ
Cassandra consulting questions we are asked most
Short answers to what engineering leaders ask before the first call.
What does MinervaDB Cassandra consulting cover?
Data modelling and partition design, read and write path performance engineering, compaction strategy and repair scheduling, tombstone control, multi-datacenter topology and consistency design, upgrades from 3.x and 4.x to 5.0, security and compliance, and platform decisions across self-managed, K8ssandra on Kubernetes, DataStax Astra DB and Amazon Keyspaces. Every finding cites the nodetool output, system table or log line behind it, and every change ships with a rollback.
Do you provide 24×7 Cassandra support?
Yes. A senior Cassandra engineer is on watch across APAC, EMEA and the Americas with S1 (production outage, data-integrity event or security incident) acknowledged within 15 minutes, S2 within 12 hours, S3 within 24 hours and S4 within 48 hours, with a repair and compaction calendar, patching, restore drills and a monthly SLO report included in remote DBA retainers.
How do you ensure Cassandra high availability and disaster recovery?
By making replication factor, rack placement and consistency level agree, NetworkTopologyStrategy with RF 3 across three racks per datacenter and LOCAL_QUORUM as the default, then proving it: node, rack and datacenter loss drills that record p99 latency and unavailable errors, cross-datacenter replication lag as the RPO, driver failover time as the RTO, and a quarterly timed keyspace restore with Medusa or sstableloader.
Why does deleted data come back in Cassandra, and how do you prevent it?
A delete writes a tombstone that must reach every replica before it is purged after gc_grace_seconds. If repair does not complete inside that window on every node, a replica that missed the delete resurrects the row. Cassandra consulting sets gc_grace and TTLs per table, schedules incremental repair with Reaper to finish inside the window, and alerts when a keyspace's last successful repair approaches the boundary.
Which compaction strategy should we use?
It depends on the table, not the cluster: size-tiered for write-heavy tables without TTL, levelled for read-heavy and update-heavy tables, time-window for time-series with a uniform TTL, and on Cassandra 5.0 the unified strategy where one parameter tunes between tiered and levelled. The choice is made from the write pattern, SSTables per read and compaction headroom measured on your nodes.
Which Cassandra versions and platforms do you support?
Apache Cassandra 4.0, 4.1 and 5.0, the three lines the project supports, on VMs, bare metal and Kubernetes with K8ssandra or cass-operator, plus DataStax Astra DB and Amazon Keyspaces with their divergence from open-source Cassandra mapped. Estates still on 3.0 or 3.11, end of life since September 2024, are upgraded through 4.x to 5.0 with SSTable, driver and JDK migration rehearsed first.
Should we move to Astra DB or Amazon Keyspaces?
Sometimes. Astra removes nodetool, the JVM and compaction tuning and prices per request; Keyspaces is CQL-compatible but not Cassandra internals, with lightweight-transaction and type differences. MinervaDB resells neither, so the assessment compares cost per unit of work at your measured p99 over twelve months, maps the feature gaps, and writes the exit path before any migration.
How much does Cassandra consulting cost?
Remote Cassandra consulting is US $300 per hour and on-site consulting is US $500 per hour with travel applying. Remote DBA retainers start at US $4,500 per quarter with a four-hour monthly minimum and emergency support always included. There are no advance payments and no long-term lock-in.
Talk to a senior Cassandra consulting engineer
Bring nodetool tablehistograms and tpstats from a peak hour, the last repair run’s Reaper report, a GC log, and the current infrastructure or managed-service invoice to the first call. We will tell you which stage of the path is the constraint, what moving it is worth, and what we would change first.