Apache Kafka · Kafka consulting, performance engineering, KRaft migration and 24×7 support
Kafka Consulting, Performance Engineering and 24×7 Support
Kafka consulting from MinervaDB is distributed-log engineering measured from broker, controller and client metrics, not a slide deck. We design and review Kafka platforms, tune the producer, replication and consumer paths until throughput and tail latency are numbers rather than opinions, run ZooKeeper-to-KRaft and 3.x-to-4.x migrations with rollback checkpoints, and then operate the cluster around the clock with named engineers. This page is deliberately technical: it is the checklist of what a competent Kafka partner should already know before touching your cluster.
01 · The platform
Kafka in 2026: what changed, and why it reshapes Kafka consulting
The 4.x line is a different platform to operate from the 2.x and 3.x clusters most estates were built on. ZooKeeper is gone, the consumer protocol moved into the broker, queues arrived as a first-class consumption model, and the release cadence is three versions a year. Kafka consulting that still assumes the 2021 platform costs its clients money.
Kafka 4.0 (March 2025) was the first release to run entirely without ZooKeeper; 3.9 is the last line that supports it and the mandatory bridge for any migration. 4.0 also made the new consumer rebalance protocol (KIP-848) the server-side default, introduced Eligible Leader Replicas (KIP-966) in preview, and moved brokers, Connect and the tools to Java 17 with clients on Java 11.
Kafka 4.2 (February 2026) took share groups, the queue semantics of KIP-932, to general availability alongside the Kafka Streams server-side rebalance protocol (KIP-1071), dead-letter queues in Streams exception handlers and a corrected metric naming scheme. Kafka 4.3 (May 2026, 4.3.1 in June) added broker and log-directory cordoning (KIP-1066), follower bootstrap from the tiered offset (KIP-1023), a minimum ISR for the remote-log metadata topic, and phase one of the classic rebalance protocol deprecation (KIP-1274).
Deprecations now have a date. The classic rebalance protocol, the streams-scala module and the old MirrorMaker metric names are all marked for removal in Kafka 5.0. An estate that has not moved its consumers to group.protocol=consumer has an upgrade cliff ahead of it, and that is a Kafka consulting finding we now write into every review.
The managed services track the same line. Confluent Platform 8.3 ships Kafka 4.3 with community patches to June 2027; 8.0 on Kafka 4.0 left community support in June 2026. Amazon MSK Express brokers took Kafka 4.2 in July 2026, and MSK Replicator can now replicate from external Kafka clusters into either Express or Standard brokers, which changes the shape of a migration into AWS. Confluent Cloud, Redpanda, Aiven and WarpStream-style diskless designs each carry a divergence map we keep current, because a Kafka consulting recommendation that ignores what the managed offering cannot do is not a recommendation.
Version and date claims above follow the Apache Kafka release page, the 4.2.0 and 4.3.0 announcements and the Confluent Platform interoperability table, all re-verified at the start of every engagement.
02 · Services
Kafka consulting services
Eight fixed-scope Kafka consulting services. Most engagements begin with the architecture and performance review and continue as the findings dictate; the 24×7 retainer keeps the same engineers on the platform afterwards.
Architecture design and design review
Topic taxonomy, partition and key strategy, replica placement across fault domains, retention and compaction policy, listener and broker-pool separation, and the controller quorum layout. Existing platforms receive a written review with prioritised, quantified findings rather than a generic best-practice list.
Performance engineering
Throughput and tail-latency work across the whole path: producer batching and compression, broker thread pools and file systems, replication fetcher parallelism, consumer poll-loop design and stream-processing state stores. Every change is benchmarked before and after so the improvement is a number.
KRaft and 4.x migration
ZooKeeper to KRaft through the 3.9 bridge, then the roll to 4.x with the client-protocol, Java and deprecation work that comes with it. Voter sizing, dual-write validation, ACL and dynamic-config verification, rollback checkpoints at every stage.
24×7 consultative support and incident response
Follow-the-sun coverage from engineers who operate distributed logs for a living. We join your incident channel, drive root-cause analysis and deliver a written post-incident report with the preventive change already scheduled.
Pipeline and integration engineering
Change data capture with Debezium from PostgreSQL, MySQL, Oracle and SQL Server, connector development and hardening, Apache Flink and Kafka Streams processing, and delivery into ClickHouse, PostgreSQL, S3, Iceberg and the warehouses. Idempotent sinks and reconciliation by design.
Managed Kafka: MSK, Confluent and repatriation
Self-managed to MSK or Confluent Cloud, Express versus Standard brokers, cloud to self-managed repatriation when the invoice justifies it, and cross-cloud replication cutovers with MirrorMaker 2 or MSK Replicator. We model both directions with your real numbers.
Cost optimisation
Broker right-sizing, tiered storage adoption, compression and retention policy, partition rationalisation, elimination of redundant replication traffic, and share groups where consumer fan-out was being paid for in partitions. Reductions of a third without touching durability guarantees are common.
Security, compliance and governance
mTLS and SASL design, OAuth client assertions (4.3), ACL rationalisation and drift audits, quotas, encryption at rest, audit logging, schema governance, and evidence packages for SOC 2, PCI DSS, HIPAA and GDPR reviews.
03 · Reference architecture
The production Kafka platform our Kafka consulting team builds and supports
A healthy Kafka deployment is layered. Publishers never write to storage directly, subscribers never depend on a single node, and cluster metadata lives in its own replicated quorum. The topology below is the shape we deploy, review and support most often in Kafka consulting work.
Figure 1. The Kafka consulting reference architecture: publishers, a rack-aware broker tier with the KRaft controller quorum and schema governance beside it, and the subscriber classes including share groups, which are new in 4.2.
Three design rules govern every cluster we build. The controller quorum is sized for an odd number of voters so a majority always exists: three voters tolerate one failure, five tolerate two. Replicas of the same partition are spread across fault domains with broker.rack, so a rack, availability zone or hypervisor loss never removes a majority of a partition’s replicas. The client tier is isolated from storage decisions by a schema contract, which lets payloads evolve without redeploying every consumer.
We also insist on separating workload classes onto different listeners and, where volume justifies it, different broker pools. Mixing a bursty clickstream with a low-latency payments topic on the same disks is the fastest route to unpredictable tail latency, and it is the single most common structural finding in our Kafka consulting reviews of platforms that grew organically.
04 · The log
Partitions, segments, offsets and hot partitions: the first hour of any Kafka consulting review
A Kafka topic is not a queue. It is a partitioned, append-only log; each partition is an ordered, immutable sequence of records addressed by a monotonically increasing offset, and ordering is guaranteed inside a partition and nowhere else. That single property is the root cause of most “events arrived out of order” tickets we investigate.
Segments and retention granularity
Each partition is stored as segments: a .log file with an offset index and a time index, a leader-epoch checkpoint that makes truncation deterministic, and partition.metadata. Segments roll when they exceed segment.bytes or segment.ms, and only closed segments are eligible for deletion or compaction. This is why a topic with seven-day retention and one-gigabyte segments can hold data far older than seven days: the active segment has not rolled. We tune segment size against retention granularity rather than leaving both at defaults.
Keys decide placement, and partition counts are a one-way door
The default partitioner hashes the record key with murmur2 and takes the modulus of the partition count. Records sharing a key always land on the same partition and preserve relative order, and increasing the partition count permanently breaks that mapping for existing keys. Partition counts are therefore planned at design review, not discovered during a migration. As a planning heuristic we keep partitions per broker in the low thousands, allow at least one partition per unit of intended consumer parallelism, and leave 30 to 40 percent headroom.
Every partition costs an open file-handle set, a replication fetch slot and a slice of controller metadata, so more is not free.
Hot-partition symptoms Kafka consulting engagements see, and the remedy we apply
| Symptom | Likely cause | Remedy |
|---|---|---|
| One partition 20× the size of its peers | Low-cardinality or skewed key, for example a tenant_id where one tenant dominates | Composite key, salted key, or a custom partitioner with explicit routing |
| Lag concentrated on one consumer | Hot partition pinned to one group member | Re-key the topic, or move the tenant to a dedicated topic; share groups where ordering is not required |
| Throughput plateau despite adding consumers | Consumers exceed partition count, so extras idle | Raise partitions with a planned re-key, fan out downstream, or switch that workload to a share group |
| Ordering violations after a scale-up | Partition count changed, so the hash modulus moved | Dual-write migration to a new topic with the final partition count |
Share groups (KIP-932, GA in Kafka 4.2) decouple consumer parallelism from partition count for workloads that need per-record acknowledgement rather than ordering. They are the first real answer to “we need 400 consumers on a 24-partition topic” that does not involve re-keying, and Kafka consulting reviews now ask which topics belong in that model.
05 · Durability
Replication, in-sync replicas and the durability guarantee Kafka consulting has to prove
Kafka durability is decided by the interaction of three settings: the replication factor, min.insync.replicas and the producer acknowledgement mode. Getting any one of them wrong silently converts a “no data loss” design into a best-effort one.
Figure 2. The durability model every Kafka consulting review starts from. Partition P0 with three replicas: the high watermark is the highest offset every in-sync replica holds, and consumers never read past it. The acknowledgement mode decides what “written” means.
With acks=all and min.insync.replicas=2 on a three-replica partition, the cluster tolerates one broker failure with no loss and no availability impact, and deliberately rejects writes with NOT_ENOUGH_REPLICAS if a second replica also fails. That rejection is a feature: it trades availability for correctness at exactly the moment the guarantee would otherwise be violated. A replica leaves the ISR when it has not fetched within replica.lag.time.max.ms, thirty seconds by default, and IsrShrinksPerSec flapping is almost always disk or network saturation rather than a Kafka bug.
We keep unclean.leader.election.enable=false on every business-critical topic. Enabling it allows an out-of-sync replica to become leader, which truncates committed records and produces the hardest class of data-loss incident to explain afterwards. Leader epochs and the leader-epoch-checkpoint file exist precisely to make truncation deterministic, and they only work if unclean elections stay disabled. Eligible Leader Replicas, in preview since 4.0 and hardened through 4.2, record which replicas are safe to lead after an ISR collapse; we enable it on new 4.x clusters and still leave unclean election off.
06 · Write path
The producer path and write-side Kafka consulting
Most throughput complaints we receive are not broker problems. They are client problems: unbatched sends, a single producer instance shared by hundreds of threads, or a compression codec chosen without measuring its CPU cost. Understanding the write path makes the fix obvious.
Figure 3. From send() to acknowledgement, with the settings that matter at each stage and the values Kafka consulting engagements usually end on.
Compression is the setting Kafka consulting changes most often. It is applied per batch, not per record, which is why linger.ms=0 combined with compression often delivers a worse ratio than no compression at all: there is nothing in the batch to compress against. We standardise on broker-side compression.type=producer so the cluster never spends CPU recompressing what the client already compressed, and we choose zstd when storage and network dominate and lz4 when the client CPU is the constraint.
Idempotence is on by default since 3.0 and should stay on; Kafka consulting reviews still find it switched off by a copied-in config from 2019. Sequence numbers per partition let the broker discard a retried duplicate without a transaction, and with idempotence enabled, five in-flight requests per connection still preserve ordering. The cost is a producer id and epoch per instance, which is one more reason a single shared producer per process beats a producer per thread.
07 · Read path
Consumer groups, rebalancing and lag: where Kafka consulting recovers availability
A consumer group is a coordination protocol, not a label. The group coordinator assigns partitions, tracks liveness through heartbeats and stores committed offsets in __consumer_offsets. Rebalances are where availability is usually lost, and lag is the only true measure of business freshness.
Figure 4. The consumer-side picture Kafka consulting works from: eager versus cooperative rebalancing, the two timeouts that get confused, and the four lag shapes that each map to a different action.
Since 4.0 the broker computes assignments under the KIP-848 protocol: members send heartbeats, the coordinator returns a target assignment, and only the partitions that move are revoked. The classic protocol’s stop-the-world revocation is gone for groups that have opted in with group.protocol=consumer, and 4.3 began logging deprecation warnings for those that have not.
Kafka Streams gained its own server-side protocol in 4.2, so a Streams topology no longer negotiates task assignment client side.
Two client timeouts are routinely confused. session.timeout.ms governs the background heartbeat and detects a dead process. max.poll.interval.ms governs how long the application may spend between polls before the coordinator assumes it is stuck. A slow database write inside the poll loop trips the second, not the first, and the resulting eviction looks like a network problem to the untrained eye. The fix is usually to lower max.poll.records and move blocking work off the poll thread.
-- lag and time-to-drain per partition, from the Admin API or kafka-consumer-groups lag(partition) = log_end_offset - committed_offset time_to_drain = lag / (consume_rate - produce_rate) when consume > produce -- what we page on: the slope, not the value healthy flat, small, sawtooth falling behind monotonic climb -> page stalled flat but large -> stuck consumer or hung transaction rebalance storm repeated spikes -> timeout tuning, cooperative protocol
We alert on the derivative because a batch job legitimately sits at high lag while a steadily climbing lag on a real-time topic is an emergency. Share groups add their own lag metric (KIP-1226) and a RENEW acknowledgement that extends a record’s processing deadline, both of which we wire into the same dashboards.
08 · Metadata
KRaft versus ZooKeeper, and the migration Kafka consulting runs most often
Modern clusters replace the external coordination service with a built-in Raft quorum that stores metadata in a dedicated internal log. For estates still on ZooKeeper the clock is running: 3.9 is the last line that supports it, and every release since 4.0 assumes it is gone.
Figure 5. The two metadata architectures and the migration path our Kafka consulting runbook follows: to 3.9 first, bridge mode with a rollback checkpoint, KRaft on 3.9, then the roll into the 4.x line.
The operational win is not merely fewer moving parts. Because every broker holds an up-to-date replica of the metadata log, controller failover no longer requires reloading state from an external store, so recovery time becomes largely independent of cluster size. The practical partition ceiling rises by an order of magnitude, there is one security model instead of two, and 4.3’s cordoning of brokers and log directories (KIP-1066) finally gives operators a clean way to drain a node before maintenance.
Our migration runbook covers voter sizing, the 3.9 bridge with metadata dual-write, rollback checkpoints that are kept until validation passes, and verification of every ACL, quota and dynamic configuration before ZooKeeper is decommissioned. The 4.x roll that follows brings its own checklist: Java 17 on brokers, client versions at 2.1 or higher before the broker upgrade, consumer groups moved to the new protocol, and the deprecations scheduled for 5.0 reviewed while there is still time.
09 · Semantics
Transactions and exactly-once: the Kafka consulting conversation we have most
Exactly-once is achievable, but only for the read-process-write pattern inside Kafka, and only when producers, consumers and the state store all participate. A good deal of Kafka consulting time goes on correcting the belief that one flag delivers it end to end.
-- consume -> transform -> produce, atomically
1. initTransactions() producer.id + epoch, fenced by the coordinator
2. beginTransaction()
3. send(outputTopic, ...) records land in the partitions
4. sendOffsetsToTransaction(consumedOffsets, groupId)
5. commitTransaction() coordinator writes COMMIT to __transaction_state
and control records to every touched partition
-- reader side: isolation.level = read_committed
|A A A|C| visible (C = commit marker)
|B B B|X| filtered (X = abort marker)
^ last stable offset (LSO)
The practical cost is throughput and latency: commit markers add round trips, and a hung transaction pins the last stable offset, which stalls every read_committed reader on that partition until transaction.timeout.ms expires. We keep transaction scope small, measure the overhead against idempotent writes with downstream deduplication, and recommend full transactional pipelines only where the business cannot tolerate duplicates.
Once data leaves Kafka for an external system, exactly-once requires either an idempotent sink or a transactional sink. Otherwise the correct design is at-least-once delivery with deduplication at the destination, which is how we build most ClickHouse and PostgreSQL sinks.
10 · Storage
Retention, compaction and tiered storage: where Kafka consulting cuts cost
Storage policy is where cost control lives. Two cleanup policies exist and they solve different problems: time or size based deletion for event streams, and key-based compaction for changelog and state topics.
Compaction
Compaction is the topic Kafka consulting reviews most often find misconfigured. It retains the latest value per key, never touches the active segment, runs on the cleaner threads governed by log.cleaner.threads and min.cleanable.dirty.ratio, and preserves original offsets, which is why a compacted topic legitimately shows offset gaps. Tombstones (a key with a null value) survive for delete.retention.ms so that consumers rebuilding state see the deletion. A cleaner that cannot keep up is a common hidden cause of disks filling on a topic that “should” be bounded, and it is invisible unless max-dirty-percent and cleaner lag are on the dashboard.
Tiered storage
Tiered storage, production-ready since 3.9, changes capacity planning fundamentally and is the largest single cost lever in Kafka consulting today. Local disk is sized for the hot working set under local.retention.ms, long retention moves to S3, GCS or Azure Blob at a fraction of the cost, and broker replacement stops being a multi-hour data-shuffling exercise; 4.3 lets a new follower bootstrap from the tiered offset (KIP-1023) rather than from the leader’s local log.
The trade-off is higher and less predictable latency for historical reads, so we validate replay performance against real backfill jobs before recommending it, and we set min.insync.replicas on the remote-log metadata topic (KIP-1235) because that topic is now part of the durability chain.
11 · Sizing
Capacity planning arithmetic Kafka consulting refuses to skip
We refuse to size Kafka clusters by intuition. The model below is deliberately simple and has survived contact with production many times; the numbers in the worked example are illustrative.
ingress_raw = events_per_sec * avg_event_bytes
ingress_on_wire = ingress_raw / compression_ratio
replication_load = ingress_on_wire * (replication_factor - 1)
broker_egress = ingress_on_wire * consumer_group_count + replication_load
daily_storage = ingress_on_wire * 86400 * replication_factor
retained_storage = daily_storage * retention_days * (1 + headroom)
partitions_min = max( target_throughput / per_partition_throughput,
required_consumer_parallelism )
-- worked example (illustrative)
200,000 events/s, 1.2 KB average, zstd 4:1, RF=3,
3 consumer groups, 7-day retention, 40% headroom
ingress_raw = 240 MB/s
ingress_on_wire = 60 MB/s
replication_load = 120 MB/s
broker_egress = 300 MB/s (fleet aggregate)
daily_storage = 60 MB/s * 86400 * 3 = 15.5 TB/day
retained_storage = 15.5 * 7 * 1.4 = 152 TB
-> 9 brokers with 20 TB usable NVMe each on 25 GbE,
or the same ingest with ~2 days local retention plus tiered storage
Notice that replication and fan-out, not raw ingest, dominate network sizing. A cluster that comfortably accepts 60 MB/s of writes can still saturate its network interfaces once three replicas and three consumer groups are accounted for. This is the calculation most often skipped in self-designed deployments, and the first one our Kafka consulting capacity studies run.
On managed platforms the same arithmetic drives the instance choice. MSK Express brokers advertise up to three times the throughput per broker of Standard brokers with storage handled by the service, which moves the constraint from disk to the per-broker network and partition limits; the model tells you which limit you will hit first.
12 · Tuning reference
Broker, JVM and operating-system tuning: the Kafka consulting reference table
Zero-copy transfer from the page cache is what makes the log fast, so memory handed to the JVM is memory taken from the mechanism that actually serves your consumers. Heap sizing is the single most valuable rule on this list.
| Layer | Parameter | Guidance |
|---|---|---|
| Broker | num.network.threads / num.io.threads |
Network threads near core count; I/O threads at roughly twice the number of data directories |
| Broker | num.replica.fetchers |
Raise to 4–8 on high-partition clusters so followers keep pace and the ISR stays stable |
| Broker | socket.send.buffer.bytes |
1 MB or more on high bandwidth-delay-product links, especially cross-region replication |
| Broker | log.flush.interval.messages |
Leave unset; rely on replication for durability rather than synchronous flush |
| JVM | Heap size | 6–12 GB is almost always correct; the page cache, not the heap, serves reads |
| JVM | Collector | G1 with a 20 ms pause target, or ZGC on very large heaps; alert on pauses above 100 ms |
| OS | vm.swappiness |
1, so the page cache is never swapped out under memory pressure |
| OS | vm.dirty_ratio / vm.dirty_background_ratio |
Lower the background ratio to smooth writeback spikes and avoid latency cliffs |
| OS | File descriptors | At least 100,000; every segment and connection consumes handles |
| Storage | Filesystem | XFS with noatime; separate data directories per physical device, never RAID 5; 4.3 cordoning to drain a directory before replacement |
| Network | MTU and offload | Consistent MTU end to end; verify segmentation offload is enabled on every host |
Every value above is a starting point that Kafka consulting engagements confirm against your own broker metrics before it is applied, staged one change at a time with a monitoring window, and recorded with its before-and-after evidence.
13 · Observability
The metrics Kafka consulting wires up first
Dashboards with two hundred panels are not observability. These are the signals we wire up first on every engagement, because each one maps directly to a decision. Note that 4.2 corrected metric names to the kafka.COMPONENT convention (KIP-1100), so dashboards built on 3.x names need a review during the upgrade.
| Signal | Source | Why it matters | Action threshold |
|---|---|---|---|
| UnderReplicatedPartitions | Broker JMX | Replicas are behind, so durability is degraded right now | Any non-zero value sustained beyond a minute |
| OfflinePartitionsCount | Controller JMX | Partitions have no leader and are unavailable | Page immediately on any non-zero value |
| ActiveControllerCount | Controller JMX | Must sum to exactly one across the cluster | Zero or more than one is a split-brain risk |
| IsrShrinksPerSec | Broker JMX | Flapping replicas usually mean disk or network saturation | Repeated shrink and expand cycles |
| RequestQueueTimeMs p99 | Broker JMX | Separates queueing delay from real work | Rising queue time with flat local time means thread starvation |
| RequestHandlerAvgIdlePercent | Broker JMX | Direct measure of I/O thread saturation | Below 30 percent means add threads or brokers |
| Consumer group lag and share-partition lag | Admin API | The only true measure of business freshness | Sustained positive slope, not absolute value |
| Log flush latency p99 | Broker JMX | Detects a failing or throttled disk before it takes a broker down | Deviation from the fleet baseline |
| Controller thread idle ratio (4.2) | Controller JMX | A saturated controller delays every metadata change | Idle ratio falling under load |
| GC pause duration | JVM | Long pauses cause spurious ISR shrinks and session timeouts | Any pause beyond 100 ms |
| Produce error rate by type | Client metrics | NOT_ENOUGH_REPLICAS and timeouts have different fixes | Any sustained non-zero rate |
Kafka consulting instruments clients as aggressively as brokers. Server-side dashboards that look perfect while a producer quietly buffers and drops records are one of the most dangerous blind spots in streaming operations, and client metrics are the only place that failure is visible.
14 · Multi-region
Multi-region topologies and disaster recovery in Kafka consulting
There is no single correct cross-region design, only trade-offs between recovery point objective, recovery time objective, write latency and cost. Kafka consulting engagements choose between three patterns and then rehearse the one chosen.
Figure 6. The three topologies Kafka consulting chooses between: active-passive, active-active and stretch cluster, with the RPO and RTO each can honestly promise.
Active-passive with MirrorMaker 2 or MSK Replicator is the cheapest and simplest, with a small but real loss window and an offset-translation step on failover that must be scripted, not improvised. Active-active with prefixed remote topics gives near-zero read RTO at the cost of consumers that understand remote topic naming and application-level conflict handling. A stretch cluster across three low-latency zones gives zero RPO and automatic failover, and it is a metro design only: sub-10 ms inter-zone latency is a hard requirement, so it never spans continents.
Whichever pattern is chosen, the recovery plan is only real if it is exercised. Our disaster-recovery engagements include scripted failover drills, consumer offset translation verification, and a documented decision tree for the human on call. We also verify that schema registry state, connector configurations and ACLs are replicated, because a cluster that comes up without them is not a recovered service.
15 · Security
Security hardening in Kafka consulting: mTLS, SASL and ACLs
Least privilege in Kafka is enforced with explicit allow rules per principal and resource. Wildcard grants are removed, quotas cap the damage any single misbehaving client can do, and the controller listener is strictly internal.
client broker | 1. TLS handshake, both sides present certificates |------------------------------------------>| | 2. SASL authentication: SCRAM-SHA-512, GSSAPI, | or OAUTHBEARER against your identity provider | (4.3 adds client assertions for client_credentials) |------------------------------------------>| | 3. principal extracted and mapped via | ssl.principal.mapping.rules | 4. authorizer checks ACLs | (principal, operation, resource, host)
Kafka consulting security reviews run ACL drift audits against the intended access model, keep credential rotation scripted so it is a routine change rather than a project, and apply the Connect client-override allowlist policy (KIP-1188) so a connector cannot quietly escalate its own permissions. For regulated estates the evidence package covers authentication configuration, authorisation matrices, encryption at rest, audit logging and retention, in the form the SOC 2, PCI DSS, HIPAA or GDPR reviewer will actually ask for.
16 · Diagnosis
Failure modes and how Kafka consulting diagnoses them
Our diagnostic method is consistent: establish whether the constraint is client-side, network, storage or coordination before changing anything. Configuration changes made without that classification are how a one-hour incident becomes a three-day one.
| What you observe | Underlying cause we usually find | Resolution |
|---|---|---|
| Producer timeouts under normal load | Request queue saturation, or a leader on a broker with a degraded disk | Add I/O threads, move leadership, cordon and replace the device |
| Endless rebalance loop | max.poll.interval.ms exceeded by slow downstream calls |
Reduce max.poll.records, make processing asynchronous, move to the KIP-848 protocol |
| Lag grows only on some partitions | Key skew creating a hot partition | Re-key, salt the key, isolate the noisy tenant, or use a share group |
| Disk fills despite short retention | Log cleaner starved, or a huge active segment that never rolled | Add cleaner threads, tune segment size and dirty ratio |
| Records disappear after a node restart | Unclean leader election combined with acks=1 |
Disable unclean election, move to acks=all with min.insync.replicas=2 |
| Consumers stall at a fixed offset | A hung transaction pinning the last stable offset | Lower transaction.timeout.ms, fence the zombie producer, shrink transaction scope |
| Cross-region replication falls behind | Undersized socket buffers on a high-latency link | Enlarge send and receive buffers, increase fetcher parallelism |
| Periodic latency spikes every few minutes | Garbage-collection pauses or page-cache writeback storms | Right-size the heap, tune dirty ratios, tune the collector |
| Schema evolution breaks a downstream job | No compatibility policy enforced in the registry | Enforce backward compatibility in CI and gate deployments on it |
| Cluster stalls after a controller change | Undersized or co-located controller quorum, or a lagging voter | Dedicated controller nodes, quorum health checks, 4.3 fetch-size controls (KIP-1219) |
18 · Questions
Kafka consulting: frequently asked questions
The questions we are asked most often on the first call, answered the way we answer them on the first call.
How many partitions should a Kafka topic have?
Enough to meet the throughput target and the consumer parallelism you need, plus 30 to 40 percent headroom, sized with the capacity arithmetic above at design time. Increasing the count later reshuffles key-to-partition mapping, so Kafka consulting reviews treat the number as a one-way door. Since Kafka 4.2, share groups remove the need to over-partition purely for consumer fan-out.
Do we still need ZooKeeper with Kafka?
No, and since Kafka 4.0 you cannot use it: 3.9 is the last release that supports ZooKeeper. New clusters are built on KRaft, and existing ZooKeeper clusters need a dated plan to reach 3.9, migrate through bridge mode and roll into the 4.x line before the 3.9 patch stream ends.
Can we get exactly-once delivery end to end?
Atomic read-process-write is achievable inside Kafka with transactions and read_committed isolation. Once data leaves for an external system you need an idempotent or transactional sink; otherwise the correct design is at-least-once delivery with downstream deduplication, which is how we build most ClickHouse and PostgreSQL sinks.
What causes most Kafka production incidents?
Client misconfiguration and partition skew cause more outages than broker failures. Defaults that are safe on a laptop are frequently wrong for a cluster handling hundreds of thousands of events per second, and the second most common cause is a durability setting, acks=1 or unclean election, that nobody remembers choosing.
Should we self-manage Kafka, use MSK, or use Confluent?
It depends on data gravity, latency requirements, compliance constraints and the true total cost at your volume. We model all three with your real numbers, including MSK Express brokers on Kafka 4.2 and Confluent Platform 8.3 on Kafka 4.3, and we support whichever you choose; our value is engineering judgement, not a licence resale margin.
Does Kafka consulting include the upgrade to 4.x?
Yes. The 4.x roll is its own project: Java 17 on brokers, clients at 2.1 or newer before the broker upgrade, consumer groups moved to the KIP-848 protocol, Kafka Streams onto the server-side protocol, dashboards updated for the corrected metric names, and the deprecations scheduled for 5.0 reviewed while there is time. Every step has a rollback checkpoint.
What do you need from us to start?
Read access to broker, controller and client metrics, the cluster and topic configuration export, and a week of query-free observation. No change is made until the findings report is agreed and change windows are booked. Most engagements are collaborative: we embed with your engineers and leave behind runbooks, dashboards and tuned configurations your team owns.
What does 24×7 Kafka support include?
Named engineers on a rota for broker and controller incidents, stalled consumers, replication and connector failures, disk and quota problems and cost spikes, plus a monthly performance review, a capacity forecast and a quarterly failover drill. S1 is acknowledged within 15 minutes around the clock, S2 within 12 hours, S3 within 24 hours and S4 within 48 hours, and every S1 closes with a written root-cause analysis.
Further reading
Related services
The systems Kafka consulting work usually finds on either side of the cluster, handled by the same practice.
Data engineering
The pipelines between sources, Kafka and the serving stores.
ClickHouse consulting
The most common real-time sink behind a Kafka topic.
PostgreSQL consulting
The most common CDC source in front of it.
Amazon RDS and Aurora support
Debezium and zero-ETL sources on AWS.
Redshift consulting
Streaming ingestion from Kinesis, MSK and Confluent.
Data strategy and analytics
Where streaming fits a wider platform.
Digital payments analytics
Payments pipelines on Kafka and ClickHouse.
Banking and FinTech
Regulated streaming platforms.
Talk to a Principal Architect about Kafka consulting
Bring your topology, your metrics and your worst-performing topic to the first Kafka consulting call. We will tell you what we would change, in what order, and what each change is worth.