Kafka is sold as a set of guarantees and operated as a set of trade-offs, and the distance between the two is where most production incidents live. Durability, ordering, exactly-once delivery, retention, cross-region survival and throughput are all things Kafka can promise, but none of them is on by default in the strong form, and every one of them is paid for in a specific configuration parameter with a specific cost in latency, disk, network or operational complexity. A cluster that was built for throughput and is later asked for durability does not get it by wishing.
This page takes six Kafka guarantees one at a time and, for each, states what the guarantee means precisely, the producer, broker and consumer settings that make it real, what those settings cost, and the archive post that shows the guarantee being engineered on a real cluster. The order runs from the guarantee most clusters get wrong first (durability) to the one most clusters over-buy (throughput), and the table at the end puts the settings side by side so that a configuration review can be read against it.
The archive it introduces covers Kafka internals, cluster setup with multiple brokers, producer configuration and performance tuning, exactly-once semantics, tiered storage and multi-region replication, Kafka as a replication bridge between databases, and the DBA’s view of the log as a database with different physics.
Kafka guarantee 1: durability, or which acknowledgement means written
The first guarantee is that an acknowledged write survives a broker failure. It has three settings and all three have to agree: the producer’s acks (1 means the leader has it in memory, all means every in-sync replica has it), the topic’s min.insync.replicas (how many replicas count as “all”), and the broker’s replication.factor for the topic. acks=all with min.insync.replicas=1 is a durability guarantee against nothing, because a single replica satisfies it, and it is one of the two most common misconfigurations a review finds.
The cost is latency: every acknowledged write waits for the slowest in-sync replica, and the producer’s p99 tracks the replication network rather than the leader’s disk. The second cost is availability under failure: with replication.factor=3 and min.insync.replicas=2, losing two brokers makes the partition unwritable, which is the correct behaviour and surprises teams who expected the cluster to degrade rather than refuse.
the Kafka education series on internals is the archive’s post on the replication protocol behind these settings, the high-water mark, and why a follower that falls behind is removed from the in-sync set. setting up a multi-broker Kafka cluster is where the replication factor is chosen, and the post’s rule is three brokers minimum for any topic that carries data anyone would miss.
Kafka guarantee 2: ordering, within a partition and nowhere else
The second guarantee is that events arrive in the order they were sent, and the precise version is: within one partition, from one producer, with idempotence enabled. Across partitions there is no ordering, which means the partitioning key decides what is ordered: all events for one customer on one partition are ordered, and events for one customer spread across partitions by round-robin are not. Without enable.idempotence=true, a retry after a transient failure can reorder two batches from the same producer, and max.in.flight.requests.per.connection above five with idempotence off is the second common misconfiguration.
The cost is partition skew: a key that carries most of the traffic (one large customer, one busy device) puts most of the traffic on one partition, and that partition’s leader broker becomes the cluster’s bottleneck. Kafka performance tuning: producer configuration is the archive’s post on the producer settings that carry this guarantee (idempotence, in-flight requests, the partitioner, linger.ms and batch size), and on measuring skew before choosing a key: the per-partition byte rate from the broker’s BytesInPerSec metrics, plotted per partition, tells whether a key is safe before the topic is created.
Kafka guarantee 3: exactly-once, from producer to consumer and no further
The third guarantee is that an event is processed once, and the precise version is narrower than the phrase: an idempotent producer writes each batch once even under retry, a transactional producer writes to several partitions atomically, and a consumer with isolation.level=read_committed reads only committed transactions. Together these give exactly-once within Kafka, and for a consume-transform-produce loop (Kafka Streams, or a hand-written consumer that produces to another topic inside the same transaction) the guarantee holds end to end.
The boundary is the sink. A consumer that writes to PostgreSQL, ClickHouse, Elasticsearch or an object store is outside the transaction, and exactly-once there requires the sink to be idempotent (an upsert on a business key, or the consumer offset stored in the same database transaction as the row).
Kafka exactly-once semantics: the ultimate guide is the archive’s post on all three mechanisms and on the boundary, and its cost accounting is the important part: transactions add a coordinator round-trip per commit, read_committed consumers see events later (they wait for the transaction to commit), and the transaction timeout has to be longer than the slowest processing step or the transaction is aborted mid-batch.
# the six guarantees, read back from a running cluster (Kafka 3.x CLI; adapt paths and bootstrap)
BOOTSTRAP="${KAFKA_BOOTSTRAP}"
TOPIC="${KAFKA_TOPIC}"
# 1. durability: replication factor, min.insync.replicas, and whether every partition is fully in sync
kafka-topics.sh --bootstrap-server "$BOOTSTRAP" --describe --topic "$TOPIC"
kafka-configs.sh --bootstrap-server "$BOOTSTRAP" --describe --entity-type topics --entity-name "$TOPIC"
kafka-topics.sh --bootstrap-server "$BOOTSTRAP" --describe --under-min-isr-partitions
kafka-topics.sh --bootstrap-server "$BOOTSTRAP" --describe --under-replicated-partitions
# 2. ordering and 3. exactly-once: the producer side lives in the client config; verify it there
grep -E 'enable.idempotence|max.in.flight|acks|transactional.id' "${PRODUCER_CONFIG}"
# 4. retention and tiered storage: per-topic retention and whether remote storage is enabled
kafka-configs.sh --bootstrap-server "$BOOTSTRAP" --describe --entity-type topics --entity-name "$TOPIC" \
| grep -E 'retention.ms|retention.bytes|remote.storage.enable|local.retention.ms'
# 5. cross-region: MirrorMaker 2 lag per replicated partition (from the MM2 connector's metrics endpoint)
curl -s "${MM2_METRICS_URL}" | grep -E 'replication-latency-ms|record-count' | head
# 6. throughput vs consumer lag: the number that says whether the guarantee is being kept for readers
kafka-consumer-groups.sh --bootstrap-server "$BOOTSTRAP" --describe --group "${KAFKA_GROUP}"
Nothing above changes state; every command is a read, and the outputs are the evidence a configuration review is written from. The two describe flags in block 1 (--under-min-isr-partitions and --under-replicated-partitions) should both return nothing on a healthy cluster, and a non-empty result is the first alert a Kafka monitoring stack needs.
Kafka guarantee 4: retention, and the log that outlives its consumers
The fourth guarantee is that an event can be read again after it was first consumed: for a new consumer bootstrapping from the beginning, for a replay after a bug, for an audit. The setting is per-topic retention (retention.ms and retention.bytes, whichever is reached first), and the cost is disk on every replica: a topic at 1 GB/hour with seven-day retention and a replication factor of three is half a terabyte of broker disk for that one topic, and the disk is the fast, expensive kind because the same volumes serve the live writes.
Tiered storage changes the cost. Kafka tiered storage and multi-region replication is the archive’s post on KIP-405 (production-ready in Kafka 3.9, with 4.x hardening it): closed log segments are copied to object storage, local.retention.ms keeps only the recent segments on broker disk, and a consumer reading old offsets is served from the remote tier at object-storage latency.
The trade is exactly that latency: a replay from a month ago is now slower, and a consumer that lags by days reads from S3 rather than from an SSD. The post’s rule is to size local retention to the slowest healthy consumer’s lag plus a margin, and to alert when any consumer’s lag crosses into the remote tier.
Kafka guarantee 5: survival of a region, and what replication does not promise
The fifth guarantee is that the log survives the loss of a data centre or a cloud region, and it is the one where the vocabulary matters most. Intra-cluster replication (the replication factor) protects against broker loss within one cluster; it does not protect against the loss of the region the cluster is in.
Cross-region survival needs either a stretched cluster (brokers in several regions, one cluster, with replica.selector.class and rack awareness so that replicas land in different regions, paid for in write latency on every acknowledged write) or two clusters with asynchronous replication between them (MirrorMaker 2 or a vendor equivalent, paid for in a non-zero RPO and in consumer offset translation at failover).
The tiered storage post above covers both topologies and the decision between them, and its position is the one MinervaDB takes on an engagement: a stretched cluster only where the inter-region latency is low enough that acks=all stays inside the producer’s budget, and two clusters with MM2 everywhere else, with the failover rehearsed quarterly and the offset translation tested as part of the rehearsal.
using Apache Kafka to replicate data between databases is the related use of the same machinery, where the log is the bridge in a migration or a cross-region database replication and its own durability and ordering guarantees become the database’s RPO.
Kafka guarantee 6: throughput, and the guarantee most clusters over-buy
The sixth guarantee is that the cluster keeps up, and it is the one most often provisioned for and least often measured. Kafka’s throughput is bounded by the slowest of the broker’s disk sequential write rate, the replication network, the number of partitions (for parallelism) and the consumer’s processing rate; a cluster that is “slow” is almost always one of the last two, not the brokers.
The evidence is consumer lag per group, plotted over time, against the produce rate per topic; lag that grows during the day and recovers at night is a consumer with too little parallelism, and lag that never recovers is a consumer that cannot keep up at all.
improving Apache Kafka performance and scalability is the archive’s post on the broker-side levers (partition count, num.io.threads and num.network.threads, log segment size, compression) and on the trade each one makes against the guarantees above: more partitions means more parallelism and also more open files, more replication traffic and a longer leader election after a broker failure.
mastering Apache Kafka for DBAs is the archive’s translation of all six guarantees into database terms: acks=all is synchronous_commit=on, the partition is the shard, retention is the WAL archive, and consumer lag is replication lag. why CTOs choose MinervaDB for scalability is the customer-side statement of the standard the review is held to.
The six Kafka guarantees, side by side
| Guarantee | Settings that make it real | What it costs | What it does not cover |
|---|---|---|---|
| 1. Durability | acks=all, min.insync.replicas=2, replication.factor=3, all agreeing |
Write latency tracks the slowest in-sync replica; partition refuses writes below min ISR | Loss of the region; a single in-sync replica |
| 2. Ordering | enable.idempotence=true, in-flight requests at most 5, a partition key that is the ordering unit |
Partition skew on a hot key; one leader broker as the bottleneck | Anything across partitions |
| 3. Exactly-once | Idempotent + transactional producer, read_committed consumer, offsets in the transaction |
Coordinator round-trip per commit; consumers see events later; transaction timeout tuning | Any sink outside Kafka unless it is idempotent |
| 4. Retention | retention.ms / retention.bytes; with tiering, remote.storage.enable and local.retention.ms |
Disk on every replica; with tiering, object-storage latency on old offsets | Compacted topics keep only the last value per key |
| 5. Region survival | Stretched cluster with rack-aware replicas, or two clusters with MirrorMaker 2 | Inter-region latency on every write, or a non-zero RPO and offset translation at failover | The replication factor alone covers none of this |
| 6. Throughput | Partition count, broker thread pools, segment size, compression, consumer parallelism | More partitions: more files, more replication traffic, longer elections | A consumer that cannot keep up, which is the usual cause |
The figures implied here (three replicas, in-sync minimum of two, five in-flight requests) are Kafka’s own documented recommendations rather than benchmark results; the right values for a cluster come from the producer’s latency budget and the consumers’ lag tolerance, measured on that cluster.

Reading a Kafka configuration against the guarantees
A configuration review at MinervaDB is a walk down the table: for each guarantee, what the workload needs (stated by its owner, in the same measurable form as any other contract), what the configuration provides, and the gap. The two most common findings have already been named: durability settings that do not agree with each other, and ordering assumed across partitions. The next two are retention set by default (seven days, whether or not any consumer could ever need a replay) and a cross-region story that consists of the replication factor.
Every change that closes a gap is staged the way a database change is. A change to min.insync.replicas or the replication factor is applied per topic, on a non-critical topic first, with the under-min-ISR and under-replicated checks run before and after; a producer change (acks, idempotence) is deployed to one producer instance and its p99 compared with the others.
A retention change is applied in the direction of keeping more before keeping less, since shortening retention deletes data and cannot be undone. The rollback for every change is the previous value, written into the ticket before the change is made, and the blast radius of a broker-level setting (thread pools, segment sizes) is the whole cluster, so those are changed one broker at a time with a rolling restart.
Version notes: the KRaft controller replaced ZooKeeper as the default in Kafka 3.3 and ZooKeeper mode was removed in 4.0, so a cluster-setup post that describes a ZooKeeper ensemble applies to 3.x clusters only; tiered storage reached production readiness in 3.9; the exactly-once protocol was strengthened in 3.x (KIP-890) and the defaults for idempotence changed in 3.0 (on by default for the Java producer).
Confirm the broker and client versions before applying any setting from an archive post, test every change on a staging cluster carrying a replay of production traffic, and keep the DR posture in view: a Kafka cluster’s backup is its replication, and replication that has never been failed over is a backup that has never been restored.
Kafka as the bridge out of an incumbent platform
The archive’s two posts that are not about Kafka’s own configuration are about what the guarantees are for. Oracle Exadata cost optimisation: migrating to an open-source stack uses the log as the replication bridge in a migration off a proprietary platform, and the migration’s RPO is guarantee one, its ordering of changes per row is guarantee two, and its ability to pause and resume the cutover is guarantee four.
the fractional chief data officer and real-time analytics is the ownership question: a Kafka cluster that carries the enterprise’s events needs an owner for the six guarantees in the same way a database needs an owner for its backups, and in an organisation without a data leader that ownership defaults to whoever last restarted a broker.
Both posts make the same point from opposite ends: the guarantees are only worth engineering when a consumer depends on them, and a review that starts from the consumers’ needs rather than from the cluster’s defaults produces a shorter, cheaper list of changes.
Where this Kafka archive sits
This archive is the log layer of minervadb.com. The data engineering archive covers the contracts that ride on these guarantees, the data strategy archive is where the event backbone is chosen, the monitoring archive is where consumer lag and under-replicated partitions become alerts, and the ClickHouse ingestion side of the same log is covered on ChistaDATA. The authoritative reference for every parameter named here is the configuration section of the Apache Kafka documentation.
For a Kafka configuration review read against these six guarantees, a cross-region or tiered-storage design with its failover rehearsed, or 24×7 support for the log alongside the databases on either side of it, the MinervaDB database consulting practice starts from this table, and states for every recommendation which guarantee it strengthens and what it costs.