Every storage engine pays for its writes three times, and RocksDB is the engine that lets the operator choose how. A key written once is written again when its memtable flushes, and again each time the SST file holding it is compacted into a lower level; that is write amplification.
The same key may be looked for in the memtable, in every level’s files whose range covers it, and in the block cache before it is found; that is read amplification. And until compaction removes the older versions, the key occupies space in several files at once; that is space amplification. The three cannot be minimised together, and RocksDB’s several hundred options are, almost all of them, a way of deciding which two to favour.
This page organises RocksDB around that three-way trade. Each of the three amplifications gets a section that states what produces it, the statistic that measures it, the option that moves it and what that option costs in the other two. A fourth section covers the read path that the three shape (Bloom filters, the block cache, and the levels), and a fifth covers the engines that embed RocksDB (MyRocks, MongoRocks, the vector and key-value stores) and inherit the trade whether or not their operators know it.
The archive it introduces covers the LSM-tree architecture, write throughput on multicore servers, Bloom filters for ad-tech latency, RocksDB against InnoDB for write-heavy workloads, MongoRocks thread handling, and the platform-level posts on where an LSM store belongs in an estate.
Write amplification in RocksDB: how many times a byte is written
Write amplification is total bytes written to storage divided by bytes written by the application, and on a levelled compaction with the default fanout of ten it is commonly ten to thirty. The mechanism is the LSM tree itself: the write goes to the write-ahead log and the memtable, the memtable flushes to a level-0 file, and each compaction from level N to level N+1 rewrites the overlapping files in N+1, which are up to ten times larger.
the deep dive into RocksDB’s LSM-tree architecture is the archive’s post on the structure, and how the LSM tree influences high-throughput writes is the companion on why a structure that rewrites data so often is still faster for writes than a B-tree: every write it makes is sequential.
The statistic is rocksdb.compact.write.bytes plus rocksdb.flush.write.bytes against rocksdb.bytes.written, all in the engine’s statistics, and the per-level breakdown is in the compaction stats block of rocksdb.stats (the W-Amp column). The options that move it are the compaction style (kCompactionStyleUniversal trades write amplification for space and read amplification), max_bytes_for_level_multiplier (a smaller fanout means more levels and less rewriting per level, at a read cost), and level0_file_num_compaction_trigger together with the memtable size, which decide how much data accumulates before the first compaction runs.
why RocksDB suits high write throughput better than InnoDB is the archive’s comparison, and its honest part is the amplification accounting: InnoDB writes each page in place (plus the doublewrite buffer and the redo log), RocksDB writes each key many times but never in place, and which is cheaper depends on the storage: on flash, where random writes are expensive and sequential writes are cheap, RocksDB wins on throughput and often on endurance despite the higher byte count.
Read amplification in RocksDB: how many places a key is looked for
Read amplification is the number of I/O operations (or, more usefully, the number of files and blocks consulted) per point lookup or range scan. A point lookup checks the memtables, then level 0 (where files overlap, so every file may hold the key), then one file per level below; a range scan merges iterators across all of them. The levelled layout keeps read amplification bounded (one file per level from L1 down) at the write cost above; universal compaction lets files overlap more and pays here instead.
The statistics are rocksdb.number.db.seek and rocksdb.number.db.next against the block-cache and filter counters (rocksdb.block.cache.hit, rocksdb.block.cache.miss, rocksdb.bloom.filter.useful), and the per-level file count and size in rocksdb.stats. Level 0 is the lever: level0_file_num_compaction_trigger caps how many overlapping files a read has to check, and level0_slowdown_writes_trigger and level0_stop_writes_trigger are the back-pressure that keeps the cap honest by throttling writes when compaction falls behind. A server whose reads got slow after a write burst has usually crossed the slowdown trigger; the fix is compaction throughput (threads, I/O), not a read setting.
// the three amplifications, read from a running RocksDB (C++ API; the same properties are exposed by
// MyRocks as SHOW ENGINE ROCKSDB STATUS and by most embedders as a stats endpoint)
std::string stats;
db->GetProperty("rocksdb.stats", &stats); // per-level table: Files, Size, Rd(GB), Wr(GB), W-Amp, Rn, Rnp1
db->GetProperty("rocksdb.levelstats", &stats); // file count and bytes per level: the space amplification shape
db->GetProperty("rocksdb.estimate-live-data-size", &stats); // live bytes; total SST bytes / this = space amp
db->GetProperty("rocksdb.num-running-compactions", &stats);
db->GetProperty("rocksdb.compaction-pending", &stats); // 1 = compaction debt is accumulating
db->GetProperty("rocksdb.actual-delayed-write-rate", &stats); // non-zero = the slowdown trigger is active
// options that trade one amplification for another (illustrative values; set from measurement, not from here)
rocksdb::Options o;
o.compaction_style = rocksdb::kCompactionStyleLevel; // universal: less write amp, more space/read amp
o.max_bytes_for_level_multiplier = 10; // smaller: less write amp per level, more levels to read
o.level0_file_num_compaction_trigger = 4; // higher: fewer compactions, more L0 files per read
o.write_buffer_size = 64 * 1024 * 1024; // larger memtable: fewer flushes, more memory, longer recovery
o.max_background_jobs = 8; // compaction throughput: the fix for slowdown-trigger stalls
rocksdb::BlockBasedTableOptions t;
t.filter_policy.reset(rocksdb::NewBloomFilterPolicy(10, false)); // ~1% false positives at 10 bits/key
t.block_cache = rocksdb::NewLRUCache(8ULL * 1024 * 1024 * 1024); // sized to the hot key range, not to RAM
o.table_factory.reset(rocksdb::NewBlockBasedTableFactory(t));
The comment on every option is the same: set it from the statistic above it, on the workload in question. An operator who copies the block wholesale has chosen levelled compaction with a fanout of ten, which is the default, and has learned nothing about the workload.
Space amplification in RocksDB: how many copies of a key exist
Space amplification is total SST bytes divided by live data bytes, and it is the amplification most often ignored until the disk fills. Every overwrite and delete leaves the old version in place until compaction reaches it, tombstones occupy space until they reach the bottom level, and universal compaction can hold several full copies of the data during a large compaction. Levelled compaction keeps space amplification near 1.1 in steady state; universal compaction can reach 2 or more, which is the price it pays for its lower write amplification.
The statistic is rocksdb.estimate-live-data-size against the sum of SST file sizes, and the tombstone pressure is visible in rocksdb.estimate-num-keys against the count the application expects. The options are the compaction style, ttl and periodic_compaction_seconds (which force old files through compaction so that deletes are reclaimed on a schedule rather than when the level happens to fill), and the deletion strategy on the application side: a workload that deletes in ranges should use DeleteRange so that one tombstone covers thousands of keys.
configuring RocksDB for write throughput on multicore servers covers the compaction thread pool that determines how quickly reclaimed space is actually returned, and its rule is that compaction threads are sized to the write rate, not to the core count.
The RocksDB read path: Bloom filters, block cache and what they cost
Read amplification is bounded by the layout but paid by the read path, and two structures decide how much of it reaches the disk. The Bloom filter on each SST file answers “is this key possibly here” with a configurable false-positive rate (ten bits per key gives about one percent), so that a point lookup skips most files without reading them; rocksdb.bloom.filter.useful is the count of reads it saved. The block cache holds decompressed data and index blocks, and its hit ratio is the read path’s equivalent of a buffer pool’s.
RocksDB Bloom filters for ultra-low-latency mobile ad tech is the archive’s worked example, where the workload is point lookups on keys that mostly do not exist (a device id that has not been seen) and the Bloom filter is the whole design: a lookup that touches no SST file at all is the only one fast enough.
The post’s measurement discipline is the reusable part: bits per key were chosen by plotting bloom.filter.useful against filter memory, and the block cache was sized to the hot key range rather than to the host’s RAM, because a cache larger than the working set is memory taken from the page cache that serves the compaction reads.
The trade both structures make is memory: filters and index blocks for a large database can run to gigabytes, and pinning them in the cache (cache_index_and_filter_blocks with pin_l0_filter_and_index_blocks_in_cache) is the setting that decides whether a cold read pays one I/O or three. It is a memory-for-read-amplification trade, which is the same shape as everything else on this page.
Engines that embed RocksDB, and the trade they inherit
RocksDB is rarely operated directly. It is the storage engine under MyRocks (MySQL and MariaDB), MongoRocks, several distributed SQL and key-value databases, Kafka Streams’ state stores and a growing number of vector and search engines, and each of them exposes some of the options above under its own names and hides the rest. The operator of the embedding engine inherits the three amplifications regardless.
optimising MongoRocks for dynamic thread handling is the archive’s post on one embedding, where the compaction and flush thread pools compete with the database’s own worker threads for cores and the slowdown trigger shows up as MongoDB write latency with no MongoDB cause.
understanding vector databases and five NoSQL data platforms that changed data processing are the survey posts, and the reason they sit in this archive is that several of the platforms they describe are LSM stores underneath, so that their write throughput, their read latency after a burst and their disk usage over time all follow the trade above.
from chaos to clarity: a case study of a failed CDP and why internet companies trust MinervaDB for open-source database operations are the practice-level framing: a platform built on an LSM engine that nobody measured for amplification is a platform whose disk fills and whose reads slow on a schedule its operators cannot predict, and the measurement is the remedy.
The three RocksDB amplifications, side by side
| Amplification | Measured by | Options that reduce it | What the reduction costs |
|---|---|---|---|
| Write (bytes written to storage per byte written by the app) | compact.write.bytes + flush.write.bytes vs bytes.written; W-Amp per level in rocksdb.stats |
Universal compaction; smaller level multiplier; larger memtable; higher L0 trigger | Space amplification (universal); more levels to read; memory and recovery time; more L0 files per read |
| Read (files and blocks consulted per lookup or scan) | number.db.seek, number.db.next; block.cache.hit/miss; bloom.filter.useful; L0 file count |
Levelled compaction; low L0 trigger; more bits per key in the filter; pinned index and filter blocks; larger block cache | Write amplification (levelled); more compaction I/O; filter and cache memory |
| Space (SST bytes per live byte) | Sum of SST sizes vs estimate-live-data-size; estimate-num-keys vs expected |
Levelled compaction; periodic_compaction_seconds and ttl; DeleteRange; compaction throughput |
Write amplification (levelled); compaction CPU and I/O on a schedule |
Any row’s “costs” column is another row’s amplification; that is the whole trade, and a configuration review is a statement of which row the workload can afford to lose on.

Tuning RocksDB from the statistics, in order
A RocksDB review starts by reading the three amplifications from the statistics and writing them down with the workload’s mix (read/write ratio, point versus range, key size, value size, delete rate) beside them.
The second step is to decide which amplification the workload can afford: a write-heavy ingest store with reads only by the compaction itself can accept a high read amplification and should be tuned for write; a lookup service with a fixed dataset can accept write amplification during its nightly load and should be tuned for read; a store on expensive flash whose dataset grows without bound has to be tuned for space, whatever the cost in the other two.
The third step is one option at a time, with the statistic that justified it read again after the change on the same workload. The blast radius of a RocksDB option depends on where it lives: memtable and cache sizes take effect on reopen and affect memory immediately; compaction style cannot be changed in place without a rewrite; the L0 triggers and thread counts are dynamic on most versions and are the safest first move.
The rollback is the previous option value, and for the compaction style it is a restore, which is the reason that decision is made on a copy first.
Version notes: the option names and their defaults have moved across RocksDB releases (the block-based table format, the default filter implementation, max_background_jobs replacing the separate compaction and flush thread settings, and the periodic compaction default), and every embedding engine pins its own RocksDB version and exposes its own subset; confirm the embedded version before applying a setting from an archive post.
Test every change on a copy under a replay of the production write mix, keep the previous options file with the ticket, and keep the DR posture in view: an LSM store’s backup is a checkpoint plus the WAL, and a checkpoint taken while compaction debt is high restores to a store that stalls on its first write burst.
What a RocksDB stall looks like from outside
The three amplifications are internal accounting; what the application sees is a stall. The signature is a write latency that is flat for minutes and then steps up by an order of magnitude for seconds, repeating on a period that tracks the memtable flush rate, with rocksdb.actual-delayed-write-rate non-zero during each step.
That is the slowdown trigger doing its job: level 0 has accumulated more files than the trigger allows and writes are being throttled until compaction catches up. The remedy is never to raise the trigger, which only moves the stall and worsens read amplification meanwhile; it is to give compaction the threads and the I/O bandwidth to keep level 0 clear, and to check that the disk is not already saturated by the compaction it is doing.
The second signature is a read latency that degrades slowly over days with no change in the workload, which is space amplification arriving as read amplification: tombstones and stale versions accumulating in the lower levels, so that every scan merges more data to return the same rows. periodic_compaction_seconds is the setting that bounds how long a file can sit un-compacted, and it is the one most embeddings leave at its default.
Where this RocksDB archive sits
This archive is the LSM-engine layer of minervadb.com. The InnoDB archive is the B-tree engine it is most often compared with, the Linux archive covers the host-level evidence (I/O wait, page cache, THP) that compaction stalls show up in, the database performance tuning archive places the three amplifications inside the cross-engine five-bottleneck method, and the data strategy archive is where the choice between an LSM and a B-tree engine for a workload is recorded. The authoritative reference for every option named here is the RocksDB tuning guide on the project wiki.
For a review of a RocksDB-backed store that reads the three amplifications before touching an option, an engine selection between LSM and B-tree for a write-heavy workload, or 24×7 support for MyRocks, MongoRocks and the other embeddings, the MinervaDB database consulting practice starts from this triangle, and states for every recommendation which amplification it reduces and which one it pays with.