Kafka Producer Knobs

From my experience working at GreyOrange. Refactored my article a bit with help of GPT.

I was working on a service with a senior engineer that did some light processing and produced events to Kafka in the background. This component was a part of broader pipeline that had to be exactly-once semantics. Hence as part of work, I came across interesting producer configs, which I will discuss.

Although before that, a quick brush up of Kafka. I am assuming below Kafka configs, based on which I will discuss few things, pretty similar to a setup I was working on:

Config Value
Kafka version 4.3.x (KRaft mode — no ZooKeeper)
Brokers 3
Instance type Spot VMs, 1 broker per zone (3 zones)
Replication factor 2
Retention 3 days
Storage 220–250 GiB pd-balanced per broker (zonal)
PodDisruptionBudget maxUnavailable: 1
Node pool Dedicated Spot pool, tainted, tolerations on Kafka pods
Storage class pd-balanced (not local SSD)

Three brokers, each pinned to a different GCP zone. Every partition has one leader and one follower — almost certainly in different zones.

Topic: warehouse.rack.size  (4 partitions)

Partition 0:  [msg1] [msg5] [msg9]  [msg13] ──► (append-only)
Partition 1:  [msg2] [msg6] [msg10] ──►
Partition 2:  [msg3] [msg7] [msg11] ──►
Partition 3:  [msg4] [msg8] [msg12] ──►
Consumer Group "analytics" reading topic with 4 partitions

  Consumer A ◄── Partition 0, Partition 1
  Consumer B ◄── Partition 2, Partition 3

  Add a third consumer → Kafka rebalances:
  Consumer A ◄── Partition 0
  Consumer B ◄── Partition 1
  Consumer C ◄── Partition 2, Partition 3

Partition 0  (replication factor = 3)

  Broker 1 (LEADER) ──► Broker 2 (follower) ──► Broker 3 (follower)
       │                      │                       │
   writes go here       copies from leader      copies from leader
   reads go here

  If Broker 1 dies:
  Broker 2 or 3 is elected new leader → traffic shifts automatically
3-broker cluster, 12 partitions:

  Broker 1: leads P0, P3, P6, P9    (also follows P1,P2,P4,P5,P7,P8,P10,P11)
  Broker 2: leads P1, P4, P7, P10
  Broker 3: leads P2, P5, P8, P11

In our 3-broker, RF=2 setup, each partition has 1 leader and 1 follower:
  Topic: orders  (4 partitions, replication factor: 2)
    Partition 0:  Leader → Broker 1 (zone-a)  |  Replica → Broker 2 (zone-b)
    Partition 1:  Leader → Broker 2 (zone-b)  |  Replica → Broker 3 (zone-c)
    Partition 2:  Leader → Broker 3 (zone-c)  |  Replica → Broker 1 (zone-a)
    Partition 3:  Leader → Broker 1 (zone-a)  |  Replica → Broker 3 (zone-c)
1. Producer asks any broker: "who leads partition 2 of topic warehouse.events?"
2. Broker responds: "Broker 3 leads that partition"
3. Producer connects directly to Broker 3, sends the compressed batch
4. Broker 3 writes the batch to its local log
5. Broker 1, Broker 2 (followers) pull the batch and write to their local logs
6. Once ISR quorum acks, Broker 3 replies to producer: "OK, offset 10042"

Producer ──► Broker 3 (leader)
                  │
                  ├──► Broker 1 (follower, acks)
                  └──► Broker 2 (follower, acks)
                  │
                  └──► "OK, offset 10042" ──► Producer

acks - Durability vs Throughput

batch.size and linger.ms — Throughput Tuning

linger.ms=0:
  Record arrives → batch sent immediately (low latency, low throughput)

linger.ms=10:
  Record arrives → wait up to 10ms → more records accumulate → larger batch sent
  (slightly higher latency, dramatically better throughput)

max.in.flight.requests.per.connection — Parallelism vs Ordering

# With idempotent producer (required):
max.in.flight.requests.per.connection=5   # max allowed; kafka enforces ordering via seq numbers

# Without idempotence (higher throughput, ordering risk):
max.in.flight.requests.per.connection=10+  # reordering possible on retry

ProduceRequestTimeout — Single RPC Timeout

request.timeout.ms=5000   # 5 seconds per RPC attempt

RecordDeliveryTimeout — Total Retry Budget per Record

t=0:00  Record enqueued, Kafka unreachable
t=0:05  Retry 1 → still unreachable (RPC timeout: 5s)
t=0:15  Retry 2 → still unreachable
t=0:30  Retry 3 → ...
...
t=2:00  delivery.timeout.ms expires → ErrRecordTimeout
        → record is dropped, callback fires with error
delivery.timeout.ms=120000   # 2 minutes total retry window

MaxBufferedRecords — In-Memory Cap

Bound Value Reason
Too low < 5,000 Buffer drains too fast under normal bursts
Reasonable default 50,000 Safe starting point for moderate write rates
Upper guard 500,000 Beyond this, memory pressure becomes real
HTTP request → produce() → [internal buffer, max N records]
                                    │
                                    └──► batcher → broker → ack
If buffer fills (Kafka unreachable):
  produce() → ErrMaxBuffered → caller handles the drop

ProducerBatchCompression — Compress Before Sending

Codec Ratio CPU cost When to use
gzip Highest Highest Legacy; rarely the right choice today
snappy Lower Very low CPU-bottlenecked producers
lz4 Good Low Strong default on constrained CPU budgets
zstd Often beats gzip Lower than gzip Good general-purpose default
compression.type=zstd   # on the producer

Idempotent Producer: Eliminating Duplicates on Retry

Guarantee What It Means Risk
At-most-once A message might be lost. Never delivered twice. Data loss
At-least-once Eventually delivered. Might be delivered more than once. Duplicates
Exactly-once Delivered exactly once. No loss, no duplicates. Requires deliberate config
Producer                    Broker
   │                            │
   ├──── Batch (seq=0) ────────►│  broker writes to log, prepares ack
   │                            │
   │         [network blip]     │
   │◄──── (ack never arrives)   │
   │                            │
   │  "I didn't get an ack,     │
   │   I'll retry"              │
   ├──── Batch (seq=0) ────────►│  ← broker writes it AGAIN (duplicate!)
   │◄──── "OK" ─────────────────│
Producer (PID = 42)              Broker
     │                               │
     ├── Batch (seq=0) ─────────────►│  written; seq=0 recorded for PID 42
     │          [ack lost]            │
     ├── Batch (seq=0) ─────────────►│  "already have seq=0 for PID 42 — discard"
     │◄────── "OK" ──────────────────│  producer thinks it succeeded (it did, the first time)
     │                               │
     ├── Batch (seq=1) ─────────────►│  written; seq=1 recorded
     │◄────── "OK" ──────────────────│

For those, you need Kafka Transactions. But I read online, that if related configs are not tuned properly then Kafka transactions hurt throughput.

Eg: Transactions: Atomic Multi-Partition Writes
Producer:
  beginTransaction()
    write(topic=A, partition=0, message=M1)
    write(topic=B, partition=2, message=M2)
  commitTransaction()   ← both M1 and M2 become visible atomically
  # or
  abortTransaction()    ← neither M1 nor M2 becomes visible

Other info: