MongoDB observability on the current self-managed release, MongoDB 8.3 (8.3.11 shipped on 2026-09-11), no longer suffers from a shortage of signals. The server now reports how its Cost-Based Ranker chose a plan, how much memory each query tracked at peak, how long operations waited for admission, and how far a secondary's oplog fetcher is behind, in seconds, without any arithmetic on optimes. The hard part of MongoDB observability has moved. It is no longer collection; it is deciding which few of several thousand series are allowed to wake a person at three in the morning.
This post is the reference design we use for a MongoDB observability platform on 8.3: four native sources, a collection topology that survives the loss of any single pipeline, a slow-operation dataset that ranks query shapes, and an alerting model that pages only on what users feel. Every configuration and query below is complete and runnable. Where we could not verify a detail against MongoDB's documentation, we say so instead of guessing.
Why MongoDB observability is a platform decision, not a dashboard
Most MongoDB observability we inherit is a Grafana dashboard with forty panels and no owner. It shows cache usage, connections and opcounters, and it is consulted after an incident rather than before one. That is not MongoDB observability; it is a screensaver. A MongoDB observability platform answers three questions on its own schedule: is the service meeting its latency objective right now, which cause is most likely to break that objective next, and which query shape changed when it did.
Answering those questions needs three kinds of MongoDB observability data, each with different economics. Counters and gauges from serverStatus are cheap and continuous but describe the whole server, never a single query. Slow-operation log entries describe individual queries but only the tail. FTDC captures everything at one-second resolution but lives on the node and is not meant for querying. The platform's job is to put each where it is cheapest to use and to join them by cluster, member and query shape.
The native signal sources in MongoDB 8.3
There are six MongoDB observability sources worth wiring, and each answers a different question. We keep this table on the first page of every observability design document because it prevents the most common mistake: trying to answer a per-query question with a server-wide counter.
| Source | Resolution | Answers | What changed recently |
|---|---|---|---|
serverStatus | Whatever you scrape at, typically 15 s | Throughput, cache pressure, queueing, connections, replication fetch lag | 8.3 adds metrics.query.cbr.*, queues.execution.* delinquency, metrics.ttl.* and selective output with none: 1 |
FTDC (diagnostic.data) | 1 s | What the server looked like in the seconds before an event, after the fact | 8.3 raises the directory default to 500 MiB and captures connPoolStats for mongod |
| Structured JSON log | Per operation above slowms | Which query shape was slow, with which plan, examining how much | 8.3 adds slow in-progress entries and originalQueryShapeHash for operations routed through mongos |
Profiler and currentOp | Per operation, on demand | What is running now, and how much memory it holds | 8.3 adds inUseTrackedMemBytes and peakTrackedMemBytes |
$queryStats | Per query shape, cumulative | Aggregate latency, CPU and examined counts for every shape, not just slow ones | Atlas M10+ only on 8.x; the 9.0 notes describe it on by default for reads and writes |
| Host and cgroup | 15 s | CPU steal, memory limits, device latency, filesystem headroom | Unchanged, and still the half of MongoDB observability that storage incidents live in |
FTDC deserves a specific note because it is widely misunderstood. According to the FTDC documentation, the files never contain query samples, predicates, results or user data; they hold serverStatus, replSetGetStatus, the oplog's collStats, connPoolStats and host metrics once per second. That makes FTDC safe to archive off-host as part of MongoDB observability, and the archive is the only one-second history of a node that died. The MongoDB 8.3 release notes and the serverStatus reference list the full set of fields; the sections below cover the ones we collect.
Collection topology for MongoDB observability: one exporter per member
We run one exporter per mongod, on the same host, connected to 127.0.0.1 with a direct connection. Scraping through mongos or through a replica-set connection string looks tidier and is wrong for MongoDB observability: the driver picks a member, and the metrics you receive describe whichever node it picked. Replication fetch lag, cache pressure and admission queueing are per-member properties. A primary under eviction pressure and a healthy secondary average out to a healthy-looking line.
The exporter we deploy is Percona's mongodb_exporter, which derives metric names mechanically from serverStatus paths, so new 8.3 fields appear as new series without an exporter release. It needs only clusterMonitor on admin and read on local. Give the MongoDB observability identity exactly that and nothing more; a monitoring identity with write privileges is a finding in any audit we run.
// Least-privilege identity for the exporter: MongoDB observability needs to read, never to write
use admin
db.createUser({
user: "svc_mongodb_exporter",
pwd: passwordPrompt(), // supplied from the secret store, never inline
roles: [
{ role: "clusterMonitor", db: "admin" }, // serverStatus, replSetGetStatus, top, getDiagnosticData
{ role: "read", db: "local" } // oplog window: first and last entry of local.oplog.rs
],
mechanisms: ["SCRAM-SHA-256"]
})
// Verify: the role set is exactly what we granted, nothing inherited
db.getUser("svc_mongodb_exporter", { showPrivileges: false }).rolesThe unit below pins the collectors we use. Two of them deserve restraint. collstats and indexstats emit a series per collection and per index, and an unbounded list on a multi-tenant cluster with thousands of collections will do more damage to your Prometheus than to MongoDB. List the collections your MongoDB observability actually alerts or plans capacity on, and let everything else be found through the slow log.
# /etc/systemd/system/mongodb_exporter.service — MongoDB observability collector
# One exporter per mongod, on the same host, talking to 127.0.0.1 with directConnection.
[Unit]
Description=Percona mongodb_exporter (MongoDB observability, local member only)
After=network-online.target mongod.service
[Service]
User=mongodb_exporter
# /etc/mongodb_exporter/env holds a single line, rendered by the secret store at deploy time:
# MONGODB_URI=mongodb://svc_mongodb_exporter:${MONGODB_EXPORTER_PASSWORD}@127.0.0.1:27017/admin?tls=true
EnvironmentFile=/etc/mongodb_exporter/env
ExecStart=/usr/local/bin/mongodb_exporter \
--mongodb.direct-connect \
--mongodb.global-conn-pool \
--collector.diagnosticdata \
--collector.replicasetstatus \
--collector.topmetrics \
--collector.collstats \
--collector.indexstats \
--mongodb.collstats-colls=shop.orders,shop.events \
--mongodb.indexstats-colls=shop.orders,shop.events \
--web.listen-address=:9216
Restart=on-failure
RestartSec=5
[Install]
WantedBy=multi-user.targetThe cluster label is attached once, in the scrape configuration, so every recording rule can aggregate by it. The second job scrapes the slow-operation counter that the log pipeline exposes; it becomes the numerator of the latency objective later in this post.
# prometheus.yml (fragment): the cluster label is attached here, once, so every rule can group by it
scrape_configs:
- job_name: mongodb
scrape_interval: 15s
scrape_timeout: 10s
static_configs:
- targets: ["rs-orders-0:9216", "rs-orders-1:9216", "rs-orders-2:9216"]
labels: { cluster: "orders-prod", role: "replset" }
- job_name: mongodb-slow-ops # Vector's log_to_metric counter, see the log pipeline below
static_configs:
- targets: ["rs-orders-0:9598", "rs-orders-1:9598", "rs-orders-2:9598"]
labels: { cluster: "orders-prod" }The MongoDB 8.3 fields worth collecting and deriving
A field earns a place in the MongoDB observability platform when it maps to one of the four golden signals and has a known direction of failure. Figure 2 is the map we use. It is deliberately small: anything not on it is dashboard material and never an alert.
Four of the 8.3 additions change how we build the platform. metrics.query.cbr.* counts Cost-Based Ranker invocations, the time spent ranking and the number of candidate plans, which matters because 8.3 makes multi-planning with a CBR backup the default plan selection mechanism for eligible queries. queues.execution.* exposes admission, deprioritisation and delinquency counters, so queueing is measured rather than inferred from latency. metrics.repl.network.oplogFetcherLagSeconds reports fetch lag directly. And inUseTrackedMemBytes with peakTrackedMemBytes appear in currentOp, the profiler, explain and the slow query log.
Before any of this reaches the MongoDB observability stack, we read it by hand. The script below takes two samples ten seconds apart using 8.3's none: 1 option, which suppresses every optional serverStatus section and lets you opt back in. On a busy primary that difference in payload size is noticeable, and it is the right habit for any ad hoc poller.
// mongosh: a ten-second, two-sample read of the 8.3 MongoDB observability fields that matter.
// none:1 (new in 8.3) drops every optional serverStatus section; we then opt back in to the ones we read.
// Confirm the section list on your build with Object.keys(sample()) before depending on it.
const sample = () => db.adminCommand({
serverStatus: 1, none: 1,
opLatencies: 1, wiredTiger: 1, metrics: 1, queues: 1, connections: 1
});
const n = (v) => Number(v ?? 0); // NumberLong -> Number for arithmetic
const a = sample(); sleep(10000); const b = sample();
const d = (f) => n(f(b)) - n(f(a));
const wt = b.wiredTiger.cache;
printjson({
readMeanMicros: d(s => s.opLatencies.reads.latency) / Math.max(1, d(s => s.opLatencies.reads.ops)),
writeMeanMicros: d(s => s.opLatencies.writes.latency) / Math.max(1, d(s => s.opLatencies.writes.ops)),
cacheFillPct: 100 * n(wt["bytes currently in the cache"]) / n(wt["maximum bytes configured"]),
cacheDirtyPct: 100 * n(wt["tracked dirty bytes in the cache"]) / n(wt["maximum bytes configured"]),
appThreadEvictionsPerSec: d(s => s.wiredTiger.cache["pages evicted by application threads"]) / 10,
cbrRankingsPerSec: d(s => s.metrics.query?.cbr?.count) / 10, // 8.3+
cbrChoseWinningPlanDelta: d(s => s.metrics.query?.cbr?.choseWinningPlan), // 8.3+
oplogFetcherLagSeconds: b.metrics.repl?.network?.oplogFetcherLagSeconds, // 8.3+, secondaries
connections: b.connections.current + " / " + (b.connections.current + b.connections.available)
});Two MongoDB observability cautions from experience. Means from opLatencies are means: they hide the tail completely, which is why the latency objective below is built from the slow log rather than from these counters. And the CBR counters mean little in isolation. Trend them per hour and look for step changes after a deploy or an upgrade; that is the signal, not the absolute value.
Slow-operation logs as a MongoDB observability dataset
Every slow operation MongoDB logs is a structured JSON object with component COMMAND and message id 51803. It already carries the namespace, the plan summary, the examined counts, queryShapeHash (since 8.0), planCacheShapeHash, and on 8.3 the operation's peak tracked memory. Metrics cannot tell you which query regressed. The slow-operation dataset can, provided it is kept, parsed and made aggregatable instead of rotated away.
mongod.log to a ranked list of query shapes. The record shows field shape only; values are elided.The MongoDB observability log pipeline depends on a logging contract in mongod.conf. The important line is slowOpThresholdMs: we set it equal to the latency threshold in the service objective, so that every slow-log entry is, by definition, one bad event. The profiler stays off; the slow log is written at profiling level 0.
# mongod.conf (fragment) — the logging contract the MongoDB observability pipeline depends on operationProfiling: mode: off # profiler collection stays off; the slow log still records slow ops slowOpThresholdMs: 100 # = the latency SLO threshold, so every slow entry is one bad event slowOpSampleRate: 1.0 # sampling below 1.0 makes the SLO numerator lie systemLog: destination: file path: /var/log/mongodb/mongod.log logAppend: true logRotate: reopen # logrotate owns rotation; Vector follows the reopened file setParameter: diagnosticDataCollectionDirectorySizeMB: 500 # 8.3 default; state it so drift is visible # Slow in-progress entries (8.3) are enabled with the mongod start-up option: # --defaultSlowInProgMS <ms> set above the longest legitimate operation, e.g. a nightly batch
| Parameter | Default on 8.3 | Our value | Unit | How it applies |
|---|---|---|---|---|
operationProfiling.slowOpThresholdMs | 100 | The SLO threshold (100 in the examples) | ms | Runtime with db.setProfilingLevel(0, { slowms: … }); persist in the config file |
operationProfiling.slowOpSampleRate | 1.0 | 1.0 | fraction | Runtime with setProfilingLevel; persist in the config file |
diagnosticDataCollectionDirectorySizeMB | 500 | 500, stated explicitly | MiB | setParameter; persist in the config file |
--defaultSlowInProgMS | Check db.getProfilingStatus() on your build | Above the longest legitimate operation | ms | Start-up option; restart required |
We ship the log with Vector. The transform keeps only slow-query entries, projects fourteen fields and discards filter literals by construction, because it never copies attr.command. On Enterprise builds that handle regulated data, enable redactClientLogData on the server as well; defence in depth costs nothing here. The same events feed a counter on port 9598, which Prometheus scrapes as the SLO numerator.
# /etc/vector/mongodb.toml — tail, parse, shape, then fan out to ClickHouse and to a Prometheus counter
[sources.mongod_log]
type = "file"
include = ["/var/log/mongodb/mongod.log"]
read_from = "end"
[transforms.slow_ops]
type = "remap"
inputs = ["mongod_log"]
drop_on_abort = true
source = '''
e, err = parse_json(.message)
if err != null { abort }
e = object!(e)
if e.c != "COMMAND" || e.id != 51803 { abort } # 51803 = "Slow query"
a = object(e.attr) ?? {}
. = {
"ts": parse_timestamp(e.t."$date", "%+") ?? now(),
"cluster": "${MONGO_CLUSTER_NAME}",
"member": get_hostname() ?? "unknown",
"ns": to_string(a.ns) ?? "",
"app_name": to_string(a.appName) ?? "",
"query_shape_hash": to_string(a.queryShapeHash) ?? "",
"plan_cache_shape_hash": to_string(a.planCacheShapeHash) ?? "",
"plan_summary": to_string(a.planSummary) ?? "",
"duration_ms": to_int(a.durationMillis) ?? 0,
"planning_time_us": to_int(a.planningTimeMicros) ?? 0,
"keys_examined": to_int(a.keysExamined) ?? 0,
"docs_examined": to_int(a.docsExamined) ?? 0,
"n_returned": to_int(a.nreturned) ?? 0,
"peak_tracked_mem_bytes": to_int(a.peakTrackedMemBytes) ?? null
}
'''
[sinks.clickhouse]
type = "clickhouse"
inputs = ["slow_ops"]
endpoint = "https://${CH_HOST}:8443"
database = "observability"
table = "mongodb_slow_ops"
skip_unknown_fields = true
auth.strategy = "basic"
auth.user = "${CH_USER}"
auth.password = "${CH_PASSWORD}"
[transforms.slow_ops_counter]
type = "log_to_metric"
inputs = ["slow_ops"]
[[transforms.slow_ops_counter.metrics]]
type = "counter"
field = "ns"
name = "mongodb_slow_ops_total"
tags.cluster = "{{ cluster }}"
tags.member = "{{ member }}"
[sinks.prom]
type = "prometheus_exporter"
inputs = ["slow_ops_counter"]
address = "0.0.0.0:9598"The table is ordered by cluster, shape and time, so a per-shape question reads a narrow range of a single day's partition. ttl_only_drop_parts makes retention a partition drop rather than a row-level mutation.
-- ClickHouse: MongoDB observability store, one row per slow operation, ordered for per-shape scans
CREATE TABLE observability.mongodb_slow_ops
(
ts DateTime64(3, 'UTC'),
cluster LowCardinality(String),
member LowCardinality(String),
ns LowCardinality(String),
app_name LowCardinality(String),
query_shape_hash String,
plan_cache_shape_hash String,
plan_summary String,
duration_ms UInt32,
planning_time_us UInt64,
keys_examined UInt64,
docs_examined UInt64,
n_returned UInt64,
peak_tracked_mem_bytes Nullable(UInt64)
)
ENGINE = MergeTree
PARTITION BY toYYYYMMDD(ts)
ORDER BY (cluster, query_shape_hash, ts)
TTL toDateTime(ts) + INTERVAL 30 DAY DELETE
SETTINGS index_granularity = 8192,
ttl_only_drop_parts = 1;The daily ranking is the most useful single query in our MongoDB observability platform. keys_per_doc_returned is the index-health signal we trend per shape; distinct_plans above one means the shape has run under more than one plan in the window, which on 8.3 is where we look first for Cost-Based Ranker effects.
-- Top 20 query shapes by total slow time in the last day, with the index-health and plan-flip signals
SELECT
query_shape_hash,
any(ns) AS ns,
count() AS slow_execs,
sum(duration_ms) AS total_ms,
quantileTDigest(0.95)(duration_ms) AS p95_ms,
round(sum(keys_examined) / greatest(sum(n_returned), 1), 1) AS keys_per_doc_returned,
max(peak_tracked_mem_bytes) AS peak_tracked_mem_bytes,
uniqExact(plan_summary) AS distinct_plans
FROM observability.mongodb_slow_ops
WHERE cluster = {cluster:String}
AND ts >= now64(3) - INTERVAL 1 DAY
GROUP BY query_shape_hash
ORDER BY total_ms DESC
LIMIT 20;Remember what this table is: the tail. It sees only operations above the threshold, so a shape that averages 3 ms and never crosses 100 ms is invisible here, which is exactly why $queryStats on Atlas, and the 9.0 change to collect it by default, matter. For upgrades we keep a second query that isolates every shape whose plan set changed across a timestamp. It is how we compare 8.0 and 8.3 behaviour on the same workload without guessing; our MongoDB performance regression testing write-up covers the capture side.
-- Plan-shape diff across an upgrade (for example 8.0 -> 8.3 and the Cost-Based Ranker):
-- every shape whose plan summary set changed, with its tail latency before and after
SELECT
query_shape_hash,
any(ns) AS ns,
groupUniqArrayIf(plan_summary, ts < {upgrade_ts:DateTime64(3)}) AS plans_before,
groupUniqArrayIf(plan_summary, ts >= {upgrade_ts:DateTime64(3)}) AS plans_after,
quantileTDigestIf(0.95)(duration_ms, ts < {upgrade_ts:DateTime64(3)}) AS p95_before_ms,
quantileTDigestIf(0.95)(duration_ms, ts >= {upgrade_ts:DateTime64(3)}) AS p95_after_ms
FROM observability.mongodb_slow_ops
WHERE cluster = {cluster:String}
AND ts >= {upgrade_ts:DateTime64(3)} - INTERVAL 7 DAY
GROUP BY query_shape_hash
HAVING arraySort(plans_before) != arraySort(plans_after)
AND length(plans_before) > 0
ORDER BY p95_after_ms - p95_before_ms DESC
LIMIT 50;Turning the slow log into a MongoDB observability SLO
The opLatencies counters only yield means, and a mean-latency alert fires either constantly or never. A MongoDB observability objective needs good and bad events. The design above supplies both without histograms: bad events are slow-log entries at a threshold equal to the objective, counted by Vector; total events are opcounters for the operation types the objective covers. The ratio is the error rate, and the objective turns it into a budget.
With a 99.9% objective over 30 days, we page using the multi-window burn-rate pattern from the Google SRE workbook: 14.4 times the budget rate over one hour confirmed over five minutes, and six times over six hours confirmed over thirty minutes. Those multipliers are the workbook's, chosen so that a fast burn pages when two per cent of a thirty-day budget has gone in an hour. We did not invent them and we do not tune them per cluster.
# mongodb-observability.rules.yml — recording rules first, alerts only on recorded series.
# Metric names follow mongodb_exporter's default naming, derived from serverStatus paths.
# Confirm every name against curl -s localhost:9216/metrics on your exporter build before loading.
groups:
- name: mongodb-observability-recording
interval: 30s
rules:
- record: mongodb:read_latency_mean_micros:rate5m
expr: |
rate(mongodb_ss_opLatencies_latency{op_type="reads"}[5m])
/ clamp_min(rate(mongodb_ss_opLatencies_ops{op_type="reads"}[5m]), 1)
- record: mongodb:wt_cache_fill:ratio
expr: mongodb_ss_wt_cache_bytes_currently_in_the_cache / mongodb_ss_wt_cache_maximum_bytes_configured
- record: mongodb:wt_cache_dirty:ratio
expr: mongodb_ss_wt_cache_tracked_dirty_bytes_in_the_cache / mongodb_ss_wt_cache_maximum_bytes_configured
# SLO: 99.9% of operations complete under slowOpThresholdMs (100 ms). Bad events = slow-log entries.
- record: mongodb:slo_bad_ratio:rate5m
expr: |
sum by (cluster) (rate(mongodb_slow_ops_total[5m]))
/ clamp_min(sum by (cluster) (rate(mongodb_ss_opcounters{legacy_op_type=~"query|insert|update|delete|getmore"}[5m])), 1)
- record: mongodb:slo_bad_ratio:rate1h
expr: |
sum by (cluster) (rate(mongodb_slow_ops_total[1h]))
/ clamp_min(sum by (cluster) (rate(mongodb_ss_opcounters{legacy_op_type=~"query|insert|update|delete|getmore"}[1h])), 1)
- record: mongodb:slo_bad_ratio:rate30m
expr: |
sum by (cluster) (rate(mongodb_slow_ops_total[30m]))
/ clamp_min(sum by (cluster) (rate(mongodb_ss_opcounters{legacy_op_type=~"query|insert|update|delete|getmore"}[30m])), 1)
- record: mongodb:slo_bad_ratio:rate6h
expr: |
sum by (cluster) (rate(mongodb_slow_ops_total[6h]))
/ clamp_min(sum by (cluster) (rate(mongodb_ss_opcounters{legacy_op_type=~"query|insert|update|delete|getmore"}[6h])), 1)
- name: mongodb-observability-alerts
rules:
- alert: MongoDBLatencySLOFastBurn
expr: mongodb:slo_bad_ratio:rate1h > (14.4 * 0.001) and mongodb:slo_bad_ratio:rate5m > (14.4 * 0.001)
labels: { severity: page }
annotations:
summary: "{{ $labels.cluster }}: 2% of the 30-day latency budget spent in the last hour"
runbook: "Step 1: rank shapes for the last hour in ClickHouse; step 2: serverStatus two-sample read on the primary"
- alert: MongoDBLatencySLOSlowBurn
expr: mongodb:slo_bad_ratio:rate6h > (6 * 0.001) and mongodb:slo_bad_ratio:rate30m > (6 * 0.001)
labels: { severity: page }
- alert: MongoDBWiredTigerCacheFillTrend
expr: predict_linear(mongodb:wt_cache_fill:ratio[30m], 3600) > 0.95 and mongodb:wt_cache_fill:ratio > 0.80
for: 15m
labels: { severity: ticket }
- alert: MongoDBAppThreadEviction
expr: rate(mongodb_ss_wt_cache_pages_evicted_by_application_threads[10m]) > 0
for: 30m
labels: { severity: ticket }
- alert: MongoDBOplogFetcherLag # 8.3+ field; threshold = your agreed staleness bound
expr: mongodb_ss_metrics_repl_network_oplogFetcherLagSeconds > 30
for: 10m
labels: { severity: ticket }A word on metric names, which trips up more MongoDB observability rollouts than any design flaw. The exporter derives metric names from serverStatus paths, and naming conventions have shifted between exporter releases and its compatibility mode. The names above match its default mode; check them against your own /metrics output before loading the file, and treat a rule that evaluates to no data as a failed deploy, not a quiet cluster.
Alert tiers: what pages, what opens a ticket, what stays on a graph
A MongoDB observability platform lives or dies on alert discipline. We allow exactly three tiers, and the tier is decided by whether the user can feel the condition now, not by how alarming the metric looks.
In MongoDB observability, pages are reserved for objective burn, a missing primary beyond the election budget, and an oplog window shorter than the time it takes you to detect and decide on a logical corruption, because that last one silently removes point-in-time recovery. Everything that predicts a burn but has not caused one yet is a ticket: cache fill trending toward the eviction trigger, application threads doing eviction work, rising admission delinquency on 8.3, connection counts near the configured ceiling, and a new query shape entering the top twenty. The 8.3 CBR, tracked-memory and TTL counters live on graphs until you have a quarter of baseline data.
One rule governs every alert we write: if the responder cannot name the first command to run from the alert text alone, it is not ready to page anyone. That is why the fast-burn rule above carries its runbook step in the annotation. When an alert does fire, our MongoDB performance troubleshooting guide for 8.3 covers the diagnostic sequence that follows.
What MongoDB observability costs the server
MongoDB observability has a price, and we would rather state it than pretend otherwise. None of the figures below are benchmarks; they are the mechanisms by which each source consumes resources, so that you can measure them on your own workload.
| Source | Where the cost lands | How we bound it |
|---|---|---|
| Exporter scrape | One serverStatus plus optional collectors per scrape interval | 15 s interval; bounded collstats and indexstats lists |
| Slow log | Log I/O proportional to the number of operations above slowms | Threshold equals the objective; a burst of slow logging is itself the incident signal |
| Profiler level 1 or 2 | Writes to system.profile on the hot path | Off by default; enabled per incident with a sample rate. 8.3 adds profiler.totalAbandonedWrites and throttling parameters |
| FTDC | Up to the directory limit on the data volume | 500 MiB default on 8.3; archived off-host and never deleted by our tooling |
$queryStats (Atlas) | An in-memory store capped at 1% of system memory | Read with transformIdentifiers when results leave the cluster |
The FTDC archive is small enough to be a shell script, and it is the piece of MongoDB observability most teams skip. When a node is lost, its diagnostic.data directory is lost with it, together with the only one-second record of what happened. The script copies closed files only, never touches the source directory, and verifies its own freshness.
#!/usr/bin/env bash
# /usr/local/sbin/ftdc-archive.sh — MongoDB observability evidence: copy closed FTDC files off-host every 15 minutes (cron).
# Read-only on the source: nothing in diagnostic.data is modified or deleted by this script.
set -euo pipefail
SRC="/var/lib/mongodb/diagnostic.data"
DST="s3://${FTDC_BUCKET}/$(hostname -s)/"
find "$SRC" -maxdepth 1 -type f -name 'metrics.*' ! -name 'metrics.interim' -mmin +5 -print0 |
xargs -0 -r -I{} aws s3 cp --only-show-errors --no-progress {} "$DST"
# Verification: the newest archived object should be less than 30 minutes old
aws s3 ls "$DST" | sort | tail -n 1MongoDB 9.0: what changes for the observability platform
At the time of writing, the MongoDB 9.0 release notes describe the release as production ready and rolling out incrementally to Atlas clusters on the latest-version track, with Enterprise Advanced and Community availability "coming soon". We therefore build on 8.3 field names today and plan the 9.0 changes as an exporter and dashboard update, not a redesign.
Three 9.0 changes matter to MongoDB observability. The notes state that $queryStats collects statistics by default for reads and writes, sampled at 1% through internalQueryStatsSampleRate and internalQueryStatsWriteCmdSampleRate. On 8.x the $queryStats stage requires Atlas M10 or above; a default-on version closes the blind spot of a slow-log-only dataset for shapes that never cross the threshold. New metrics.changeStreams.* counters make change stream consumers observable from the server side. And a per-operation memory limit arrives with metrics.query.operationsFailedDueToMemoryLimit, which belongs on the ticket tier from day one. Whether the $queryStats default applies to self-managed builds is something we will confirm when those builds ship, not before.
FAQ: MongoDB observability on 8.3
Is the Atlas monitoring UI enough for MongoDB observability?
For Atlas-only estates it covers collection and dashboards well, and $queryStats is available there on M10 and above. It does not give you a service level objective tied to your own latency threshold, a query-shape history you control across upgrades, or a single view across Atlas and self-managed clusters. Most estates we see need all three, which is why we treat Atlas as one MongoDB observability source rather than the platform.
Why one exporter per member instead of one per cluster?
Because replication lag, cache pressure and admission queueing are properties of a member. An exporter connected through mongos or a replica-set URI reports whichever member the driver selected, and averages hide the one member that is in trouble. Per-member collection is the first rule of MongoDB observability.
Can the profiler replace the slow log for this pipeline?
Not safely. The profiler writes to system.profile on the hot path and is a capped collection that rotates under load. The slow log is written at profiling level 0, costs log I/O only, and is already structured JSON. We enable the profiler per incident, with a sample rate, and switch it off afterwards; it is a diagnostic tool, not a MongoDB observability feed.
How do we monitor the Cost-Based Ranker after upgrading to 8.3?
Collect metrics.query.cbr.* from serverStatus and trend it hourly, and run the plan-shape diff query across the upgrade timestamp. A shape whose plan summary changed and whose p95 rose is the evidence to act on; the counters alone only tell you the ranker is busy. MongoDB observability for the optimiser always needs both.
What retention should the MongoDB observability stores have?
For MongoDB observability we keep metrics at least as long as the longest objective window plus one comparison period, slow-operation events long enough to span an upgrade cycle, and FTDC archives for as long as you would want to answer a vendor escalation. The DDL above uses thirty days for slow events; that number is a starting point, not a recommendation for regulated estates.
Where to go from here
Build MongoDB observability in the order the incidents arrive: the exporter and host metrics first, the slow-operation dataset second, the objective and its burn-rate pages third, the FTDC archive the same afternoon. Test every rule and every configuration change on a staging cluster before applying it to production, and keep your backup and disaster recovery posture current while you do; observability tells you something is wrong, it does not undo it.
If you would rather have this MongoDB observability platform built and operated for you, our MongoDB Support team runs it as part of 24×7 cover, and you can book a technical conversation with an engineer who has built it before.
Running this in production?
MinervaDB provides MongoDB Consulting, MongoDB Support, NoSQL Database Support and NoSQL Consulting with 24x7 coverage and a 15-minute S1 response. Talk to an engineer.