MinervaDB · MongoDB Support · 24×7 consultative support, remote DBA and managed operations for replica sets, sharded clusters, Percona Server for MongoDB and Atlas
MongoDB Support Engineered From the Storage Engine Up: 24×7 Consultative Cover for Mission-Critical Estates
MinervaDB MongoDB Support is a standing engineering relationship, not a ticket queue. Principal-level engineers who work daily inside WiredTiger, the replication protocol and the query planner hold your pager 24×7×365, and they spend the quiet hours removing the reasons it rings: cache pressure, unbounded arrays, wrongly ordered indexes, hot shards and oplog windows that quietly shrank below your detection time. Every recommendation arrives with the explain output, serverStatus counters or oplog evidence that justifies it.
Why MinervaDB
Why enterprises choose MinervaDB for MongoDB Support
MongoDB carries a large share of the operational weight in a modern estate: customer profiles, catalogues, carts, sessions, telemetry, feature stores and the event collections behind analytics pipelines. It keeps that position only while the document model, the index set and the cluster topology still match the workload that actually runs. When they drift apart, nothing announces itself as a MongoDB problem; it surfaces as checkout timeouts, abandoned sessions and dashboards that stopped refreshing.
Our MongoDB Support practice exists for that gap, and MongoDB Support from MinervaDB is built to close it permanently rather than incident by incident. We work at the level of the storage engine, the replication protocol and the query planner, and we translate what we find there into changes your teams can ship the same week. The people on the MongoDB Support rota are the people who diagnose the problem; there is no coordinator relaying messages between you and an engineer you never meet.
Vendor neutrality is a principle rather than a slogan. We do not resell MongoDB licences, cloud credits or hardware, so a recommendation to add a shard, move to Atlas, stay self-managed or move a workload off MongoDB entirely is judged on one criterion only: whether it makes your data layer faster, safer or cheaper to run. Our MongoDB consulting and MongoDB optimization practices sit behind the support contract for the project-shaped work.
Evidence before change
Nothing is changed in production on a hunch under MongoDB Support. Every change ticket carries the explain plan, the serverStatus delta, the profiler sample or the oplog timestamp that justifies it, plus the verification query that proves it worked.
Principal engineers on the rota
The duty engineer works daily on replica set behaviour, sharding topology, WiredTiger internals and driver-level failure handling, and a principal joins any Severity-1 that touches replication, sharding or the storage engine.
Whole-estate scope
MongoDB Support covers Community and Enterprise Advanced on bare metal, virtual machines and Kubernetes operators, Percona Server for MongoDB, Atlas, and the hybrid estates where one application spans several of them.
Reversible, staged change
Blast radius and rollback path are written down before any change to a running cluster. Binary upgrades and featureCompatibilityVersion are always separate tickets, because FCV is the only real rollback a major upgrade has.
Reference architecture
The MongoDB architecture our MongoDB Support team covers end to end
Supporting MongoDB properly means understanding every layer between the driver call and the block device, because latency and data loss can originate in any of them. A connection pool that is too small produces the same user-visible symptom as an under-provisioned cache. A read preference that quietly points at a lagging secondary looks exactly like an application bug. The diagram below is the mental model our engineers apply on every engagement.
Figure 1. The sharded cluster reference architecture MinervaDB MongoDB Support instruments layer by layer: drivers, mongos, the config server replica set, per-shard replica sets and WiredTiger, with the observability plane that generates the alerts.
Two properties of this picture matter more to MongoDB Support than any single component. The config server replica set is the authority for the chunk map, so if it is unhealthy, routing degrades even when every shard is perfectly fine. And every shard is itself a replica set, which means replication health has to be assessed per shard rather than as one cluster-wide number: a majority commit point that has stalled on shard B is invisible in an average across shards A, B and C.
The defaults MongoDB ships with are reasonable for a laptop and rarely correct for a production estate. The parameters we validate, document and put under change control on every MongoDB Support engagement are the ones below; each has a current value, a proposed value, a unit and a reload-or-restart requirement in the change record.
# mongod.conf — the settings MongoDB Support puts under change control (illustrative values; restart required)
storage:
wiredTiger:
engineConfig:
cacheSizeGB: 24 # sized to the measured working set; MUST be explicit inside a cgroup-limited container
collectionConfig:
blockCompressor: zstd # snappy where latency dominates, zstd where storage and I/O bandwidth dominate
journal:
commitIntervalMs: 100 # tighten only where j:true acknowledgement latency requires it
replication:
replSetName: rs_orders
oplogSizeMB: 65536 # write-rate × (detection + decision time); never the default 5% of disk
net:
tls:
mode: requireTLS
certificateKeyFile: /etc/mongo/tls/server.pem
security:
authorization: enabled
clusterAuthMode: x509
setParameter:
transactionLifetimeLimitSeconds: 30
cursorTimeoutMillis: 600000
Layer 4
Replica sets, elections and the durability contract
A three-member replica set is the smallest unit of durability in MongoDB, and almost every availability incident our MongoDB Support team is called into is a story about how that unit behaved under stress. The protocol is sound. What varies between estates is whether the settings around it were ever tuned to a written outage budget, and whether anyone has measured what actually happens when a primary disappears.
Figure 2. The failover timeline at default settings, the rollback exposure of w:1, and the write and read concern matrix we agree per collection under MongoDB Support.
With default settings a lost primary is detected through missed heartbeats, an election begins once electionTimeoutMillis (10 seconds) expires, the candidate with the highest optime wins a majority vote, and the new primary opens for writes after catch-up. That sequence typically costs ten to twelve seconds of write availability, and MongoDB Support starts by measuring it on your topology. Whether that is acceptable is a business decision, not a database default, and we tune the timeouts against the outage the business can actually absorb, then prove the drivers retry cleanly through it.
The most expensive misunderstanding in MongoDB is the belief that a successful write is a durable write. Acknowledgement means exactly what the write concern says it means and nothing more: a write acknowledged at w:1 that has not reached a majority can be discarded when the old primary rejoins and rolls back, after the application has already told the customer it succeeded. MongoDB Support audits every collection for that exposure and closes it where correctness is not negotiable.
// Engineered replica set: deterministic priority ladder, hidden backup source,
// delayed member for logical corruption, majority floor as the default write concern
cfg = rs.conf()
cfg.members[0].priority = 3 // preferred primary, largest failure domain
cfg.members[1].priority = 2
cfg.members[2].priority = 1
cfg.members[3].priority = 0 // hidden backup source: never elected, never read
cfg.members[3].hidden = true
cfg.members[4].priority = 0 // delayed member: one hour of live rollback capacity
cfg.members[4].hidden = true
cfg.members[4].secondaryDelaySecs = 3600
cfg.settings.electionTimeoutMillis = 5000 // tuned to the written outage budget, not left at 10000
rs.reconfig(cfg)
db.adminCommand({ setDefaultRWConcern: 1,
defaultWriteConcern: { w: "majority", j: true, wtimeout: 5000 },
defaultReadConcern: { level: "majority" } })
// Verify: no member shares a priority, and lag is within the agreed bound
rs.status().members.map(m => ({ name: m.name, state: m.stateStr,
lagSec: (rs.status().members.find(p => p.stateStr === "PRIMARY").optimeDate - m.optimeDate) / 1000 }))
Layer 5
WiredTiger cache, eviction and checkpoints: where MongoDB Support finds unexplained latency
When a cluster is described as mysteriously slow, the answer is usually inside WiredTiger. Cache occupancy, dirty content ratio, eviction pressure and checkpoint duration together explain the large majority of latency complaints we investigate, and none of them are visible from an application trace.
Figure 3. The WiredTiger write path, the eviction and dirty thresholds at their defaults, and the serverStatus signals MinervaDB MongoDB Support alerts on by trend rather than by breach.
The cache defaults to half of RAM minus one gigabyte, which is frequently wrong inside a container where the process reads host memory rather than the cgroup limit. Once occupancy passes the eviction target (80 percent), background workers start reclaiming pages. Once it passes the eviction trigger (95 percent), application threads are conscripted into eviction and p99 latency changes shape within seconds. Dirty content behaves the same way on a tighter scale: above 5 percent reconciliation turns aggressive, above 20 percent writers are throttled until the ratio recovers.
A write-heavy batch job that was harmless last quarter can push a cluster across that line as data volume grows, with no code change and no configuration change to blame. That is why MongoDB Support alerts are on the trend towards the trigger, and why the cache is sized against the measured working set (indexes, hot documents and the history store) rather than against a rule of thumb.
// The WiredTiger triage query MongoDB Support runs first on any latency incident
const wt = db.serverStatus().wiredTiger, c = wt.cache;
({
cacheUsedPct: (c["bytes currently in the cache"] / c["maximum bytes configured"] * 100).toFixed(1),
dirtyPct: (c["tracked dirty bytes in the cache"] / c["maximum bytes configured"] * 100).toFixed(1),
appThreadEvictPg: c["pages evicted by application threads"],
bytesReadIntoCache: c["bytes read into cache"],
historyStoreBytes: c["history store table size"] ?? c["bytes belonging to the history store table in the cache"],
lastCheckpointMs: wt.transaction["transaction checkpoint most recent time (msecs)"],
readTicketsAvail: db.serverStatus().wiredTiger.concurrentTransactions?.read?.available,
writeTicketsAvail: db.serverStatus().wiredTiger.concurrentTransactions?.write?.available
})
// Healthy: cacheUsedPct < 80, dirtyPct < 5, appThreadEvictPg ≈ 0 between samples, checkpoint in seconds.
// appThreadEvictPg climbing between two samples a minute apart is the signal that user queries are paying for eviction.
Layer 2
Index strategy, the query planner and aggregation pipelines
Almost every performance problem we are engaged on under MongoDB Support reduces to one of three things: a query with no usable index, a compound index whose field order does not match the query, or an index that exists but is never chosen because the plan cache learned a bad lesson under different data conditions. The planner normalises each query into a shape, looks for a cached plan and, where none exists, races candidate plans over a works budget before caching the winner. That plan persists until a replan is triggered, which is exactly why a plan chosen against last quarter’s data distribution can outlive its usefulness.
Figure 4. The planner pipeline, the E-S-R compound index rule, the explain signals that drive each optimisation lever, and the shard-versus-merger split of an aggregation pipeline.
Field order in compound indexes follows the equality, sort, range guideline. Equality predicates lead, fields that satisfy the sort come next, and range predicates come last. Get the order wrong and MongoDB will still use the index, but it will add a blocking in-memory sort that spills once it exceeds 100 MB, turning a fast query into an unpredictable one. The levers MongoDB Support applies, in order of impact, are: add the missing compound index, reorder an existing one, make the query covered, retire indexes with near-zero usage in $indexStats over a full business cycle, replace leading-wildcard regular expressions with collation-aware or search indexes, and replace large skip values with range-based pagination.
Aggregation is where MongoDB does its most valuable work, where MongoDB Support spends much of its performance time, and where it is easiest to write something that is fast on a test dataset and ruinous in production. A $match that runs first can be served by an index and streams; a $sort without a supporting index and a $group both buffer, capped at 100 MB unless allowDiskUse converts the hard failure into a slow disk-bound success. On a sharded cluster the pipeline is split, and pushing $match and $project before the split is the cheapest optimisation available because it reduces both per-shard work and the bytes that cross the network to the merger.
// 1. Find the query shapes that hurt: profiler at 100 ms, then rank by keysExamined per document returned
db.setProfilingLevel(1, { slowms: 100, sampleRate: 1.0 })
db.system.profile.aggregate([
{ $match: { ns: "shop.orders", op: { $in: ["query", "command"] } } },
{ $group: { _id: "$queryHash", n: { $sum: 1 }, avgMs: { $avg: "$millis" },
ratio: { $avg: { $divide: ["$keysExamined", { $max: ["$nreturned", 1] }] } } } },
{ $sort: { ratio: -1 } }, { $limit: 10 } ])
// 2. Confirm the plan before and after: no SORT stage, no FETCH on a covered query
db.orders.find({ tenant_id: 42, status: "open", created_at: { $gte: ISODate("2026-09-01") } })
.sort({ created_at: -1 }).explain("executionStats").executionStats
// 3. Build to E-S-R, rolling, in the agreed window
db.orders.createIndex({ tenant_id: 1, status: 1, created_at: -1 },
{ name: "ix_orders_tenant_status_created" })
// 4. Retire what nothing uses, after a full business cycle of evidence
db.orders.aggregate([{ $indexStats: {} }, { $match: { "accesses.ops": { $lt: 10 } } },
{ $project: { name: 1, "accesses.ops": 1, "accesses.since": 1 } }])
Layer 7
Backup, point-in-time recovery and disaster readiness under MongoDB Support
A backup you have never restored is a hypothesis. We treat recovery as a measured capability: a written recovery point objective, a written recovery time objective, and rehearsals that produce real numbers against both. Point-in-time recovery depends on two things being true at once: a snapshot at or before the recovery target, and an oplog that still contains every operation from that snapshot forward. The second condition is the one that quietly fails as write volume grows.
Figure 6. Snapshots plus continuous oplog capture, the oplog window measured against detection time, the rehearsed five-step restore, and a three-domain disaster-recovery topology.
That is why the oplog window is a first-class recovery metric in our MongoDB Support scorecard rather than a replication detail. If corruption is detected six hours after it happened and the oplog holds four hours, point-in-time recovery is impossible at any price. We size local.oplog.rs against realistic detection and decision time and alert on the trend rather than the breach. Snapshots are taken from a hidden member so production nodes are untouched, sharded clusters are snapshotted consistently across shards rather than one shard at a time, and every restore is timed on production-scale data with the wall-clock result recorded.
The MongoDB Support tooling is chosen per estate: Percona Backup for MongoDB for self-managed replica sets and sharded clusters, Ops Manager or Cloud Manager where Enterprise Advanced is licensed, Atlas continuous backup for managed deployments, or volume snapshots with oplog capture where the platform provides them. The rehearsal is not optional under any of them. The number we insist on knowing is how long a full restore of your largest shard actually takes, on the hardware you have today.
# Oplog window: the metric MongoDB Support trends against your detection time
mongosh --quiet --eval '
const o = db.getSiblingDB("local").oplog.rs;
const first = o.find().sort({ $natural: 1 }).limit(1).next().ts.getTime();
const last = o.find().sort({ $natural: -1 }).limit(1).next().ts.getTime();
print("oplog window hours:", ((last - first) / 3600).toFixed(1))'
# Percona Backup for MongoDB: consistent snapshot + continuous oplog, then a timed point-in-time rehearsal
pbm config --set pitr.enabled=true --set pitr.oplogSpanMin=10
pbm backup --type physical --wait
pbm status # snapshot list and PITR ranges: prove the target ts is covered
time pbm restore --time "2026-09-29T20:09:59" --wait # into the isolated rehearsal replica set
# Verify before any cutover decision: counts, checksums, application smoke tests, then record the wall-clock RTO
mongosh "mongodb://${MDB_USER}:${MDB_PASSWORD}@rehearsal-0,rehearsal-1,rehearsal-2/?replicaSet=rs_rehearsal" \
--eval 'db.getSiblingDB("shop").orders.countDocuments({ created_at: { $lt: ISODate("2026-09-29T20:10:00Z") } })'
Incident response
Severity matrix, escalation and the first hour of a MongoDB Support incident
When a production cluster is in trouble, the value of a support contract is decided in the first sixty minutes: who answers, what they do before they touch anything, how quickly a principal engineer is involved, and whether evidence is captured before a restart destroys it. Our discipline is containment before diagnosis and diagnosis before change. A restart may clear a symptom, and it will also destroy the state that explains it, which guarantees a second outage later.
Every MongoDB Support Severity-1 closes with a written root cause analysis and tracked preventive actions, and every action lands in the next monthly service review. The evidence capture below runs before anything is restarted; it takes under a minute and it is the difference between a root cause and a guess. For estates that need the pager held rather than shared, MinervaDB 24×7 emergency DBA coverage and enterprise database management extend the same severity matrix to fully managed operations.
// First-hour evidence capture: run BEFORE any restart, stepdown or index change
const t = new Date().toISOString().replace(/[:.]/g, "-");
const snap = {
currentOp: db.adminCommand({ currentOp: true, "secs_running": { $gte: 5 } }),
serverStatus: db.serverStatus(),
rsStatus: rs.status(),
hostInfo: db.hostInfo(),
slowest: db.getSiblingDB("shop").system.profile.find().sort({ millis: -1 }).limit(50).toArray()
};
fs.writeFileSync(`/var/tmp/mdb-evidence-${t}.json`, JSON.stringify(snap));
// then: contain (kill the runaway op, shed the batch job, add read capacity), diagnose, change, write up
Coverage
Platforms, versions and deployment models covered by MongoDB Support
MongoDB Support from MinervaDB covers MongoDB wherever it runs, including the awkward hybrid estates where a single application spans a self-managed cluster and a managed service. Support is never conditional on moving to a platform we prefer, and the version coverage runs from end-of-life 4.x and 5.x estates that still need a safe, hop-by-hop upgrade path through the current 7.0 and 8.0 series and the rapid releases beyond them.
| Dimension | Covered | What MongoDB Support adds |
|---|---|---|
| Distributions | MongoDB Community, MongoDB Enterprise Advanced, Percona Server for MongoDB, MongoDB Atlas, Amazon DocumentDB (migration off, or compatibility assessment) | Feature-parity review per distribution; licensing-neutral recommendations; Atlas cost and configuration review against the same golden signals |
| Versions | Upgrades from 4.x and 5.x; production operation on 7.0 and 8.0 long-term series; rapid releases through 8.3 assessed per feature before adoption, per the MongoDB versioning policy | Hop-by-hop upgrade plans with a soak period between binary upgrade and featureCompatibilityVersion, and driver compatibility matrices proven in staging |
| Deployment | Bare metal and colocation, VMware and private cloud, AWS, Microsoft Azure, Google Cloud, Kubernetes operators including Percona Operator for MongoDB and OpenShift | cgroup-aware cache sizing, storage-class latency validation for checkpoints, pod anti-affinity aligned to failure domains |
| Workload features | Time series collections, change streams into Kafka, multi-document transactions, Atlas Search and vector search, Queryable Encryption, Ops Manager and Cloud Manager | Bucketing and TTL design for telemetry, resume-token handling for consumers, transaction lifetime and retry verification |
| Adjacent estates | Polyglot environments alongside our PostgreSQL support, MySQL support, MariaDB support and wider NoSQL practice | One severity matrix and one service review across every engine in the estate |
Engagement models
How MongoDB Support engagements are structured
Different problems need different commercial shapes of MongoDB Support. A regulated platform with a permanent on-call requirement needs something very different from a team that needs three weeks of index and schema work before a launch. All of the models below combine inside one agreement, and pricing is structured per cluster rather than per node, so scaling horizontally does not multiply your support cost and development environments do not attract production pricing.
24×7 consultative MongoDB Support retainer
Named engineers, the severity matrix above, unlimited incident volume, golden-signal monitoring wired into your stack and a monthly service review. Your team retains operational control; we provide the expertise, the pager cover and the remediation plans.
Fully managed MongoDB operations
MinervaDB holds the pager and owns uptime, performance and cost outcomes directly while you keep full administrative access. Many customers start on consultative support and move specific clusters into managed operations later.
Fixed-scope MongoDB Support health check
A structured review of schema, indexes, configuration, topology, backup posture and security in two to three weeks, delivered as a prioritised remediation plan with the evidence behind each finding and an executive summary.
Architecture and migration engagement
Shard key design, replica set topology for your failure domains, resharding plans, relational-to-document modelling, and migrations onto or off MongoDB with dual reads and feature-flag cutover.
Support credits
A pre-purchased block of principal-level engineering hours with no expiry, drawn for incidents or projects as needed. Suited to variable workloads that need expert capacity available but not always active.
Training and enablement
Workshops on MongoDB for relational teams, schema design at scale, WiredTiger and replication internals for performance engineers, and incident drills run against your own topology.
Every MongoDB Support engagement begins the same way: a short discovery call, a rapid assessment against the seven layers above, and an executive summary with a risk matrix, a cost impact model and a prioritised roadmap. You see the findings before you commit to the work. As with every MinervaDB recommendation, test before applying to production and maintain a robust disaster-recovery posture throughout.
FAQ
Frequently asked questions about MinervaDB MongoDB Support
The questions we are asked most often before a support agreement starts.
What does MinervaDB MongoDB Support include?
Round-the-clock incident response under the severity matrix (Severity-1 acknowledged in 15 minutes), golden-signal monitoring across driver, query, schema, replication, WiredTiger, sharding and recovery layers, schema and index reviews with explain evidence, shard key and replica set design, backup and point-in-time recovery assurance with timed rehearsals, security hardening with audit evidence, and a written root cause analysis after every critical incident.
How quickly do you respond to a production MongoDB outage?
Severity-1 incidents are acknowledged within 15 minutes, 24 hours a day, every day of the year. A duty engineer contains first, captures evidence before any restart, and a principal engineer joins within the hour for anything touching replication, sharding or the storage engine. Severity 2, 3 and 4 carry 12, 24 and 48 hour targets respectively.
Does MongoDB Support cover MongoDB Atlas as well as self-managed clusters?
Yes. MongoDB Support covers MongoDB Community and Enterprise Advanced on bare metal, virtual machines and Kubernetes, Percona Server for MongoDB, and Atlas deployments, including hybrid estates where one application spans a self-managed cluster and a managed service. Atlas is reviewed against the same golden signals and cost model as a self-managed cluster.
How do you diagnose unexplained MongoDB latency?
With evidence rather than theories: WiredTiger cache occupancy and dirty ratio, pages evicted by application threads, checkpoint duration, read and write ticket queueing, replication lag, slow query and profiler samples, plan cache and replan events, and connection pool behaviour. In most engagements the cause is cache pressure, a missing or wrongly ordered index, or a blocking aggregation stage that has started spilling to disk.
Can you help us choose or change a shard key?
Yes. We analyse cardinality, frequency and monotonicity over a real data sample, replay actual query shapes against each candidate key, and quantify write distribution and read amplification before anything is sharded. Where an existing key already causes a hot shard, we model the cost and duration of resharding against the ongoing penalty of leaving it, and since MongoDB 8.0 we can also unshard or move collections that should never have been sharded.
Does MongoDB Support provide point-in-time recovery?
Yes. MongoDB Support implements consistent snapshots plus continuous oplog capture, typically with Percona Backup for MongoDB on self-managed estates, so recovery is to a nominated second rather than to the last backup. The oplog window is treated as a first-class recovery metric and alerted on its trend, because point-in-time recovery becomes impossible once the window closes. Restores are rehearsed on a schedule and the wall-clock results are recorded.
Which MongoDB versions does MongoDB Support cover?
Production operation on the 7.0 and 8.0 long-term series, rapid releases assessed per feature, and upgrade paths from end-of-life 4.x and 5.x estates. Every major upgrade is sequenced hop by hop with the binary upgrade and the featureCompatibilityVersion change kept as separate tickets, because FCV is the only real rollback a major upgrade has.
Will you work alongside our existing DBAs and platform team?
Yes, and that is the normal shape of an engagement. Your team keeps operational control and administrative access; we add principal-level depth, pager cover, reviews and remediation plans, and we document every change so that knowledge transfers rather than accumulates on our side. Where a team has no MongoDB specialist at all, the fully managed model takes the pager entirely.
Put principal-level MongoDB Support behind your production clusters
Speak with a MinervaDB engineer who has recovered real clusters under real pressure. The first conversation is a technical one about your topology, your golden signals and your recovery numbers, never a sales pitch.