←Back to Blogs

Handling de-duplication in Distributed Architecture

July 29, 2026

Distributed SystemsSystem DesignIdempotencyMessage Queues

Duplicate events are not a bug in modern distributed systems — they're the price you pay for reliability. Almost every message broker, API gateway, and stream processor in production today chooses at-least-once delivery over at-most-once, because losing data is worse than seeing it twice. That single design choice is why deduplication exists as a discipline at all.

The good news: In practice, nearly every real-world system leans on some combination of five patterns. This post walks through each one, when to reach for it, and what it costs you.

What is deduplication?

De-duplication is the architectural practice of ensuring that even if a system receives the same request, message, or event multiple times, it only processes the state change exactly once.

Where and How Deduplication can happen?

In distributed environments, network timeouts and automated retry mechanisms are standard. De-duplication was introduced to prevent these "retry storms" from causing catastrophic business logic failures—like charging a customer's credit card twice, sending duplicate emails, or corrupting aggregate data.

Duplicates can be introduced (and caught) at multiple points. Design decision #1 is: which layer should own it? Usually you want it as close to the source as possible, but you often need defense-in-depth across two or three layers.

Where and How Dedup Can Happen

Each numbered layer is a checkpoint. A resilient system typically implements at least two of these (e.g., idempotency key at the API layer and a unique DB constraint), because any single layer can fail (Redis eviction, race condition, key collision).

Core Vocabulary

  • Idempotency: An operation that produces the same end-state no matter how many times it's applied. SET x = 5 is idempotent. x = x + 5 is not.
  • Deduplication: Detecting that "this exact request/event/message has already been seen" and dropping/short-circuiting it.
  • Exactly-once processing: The effect of processing a message exactly once, even though the message may be delivered more than once. This is achieved by combining at-least-once delivery + idempotent processing (or dedup). True "exactly-once delivery" at the transport level essentially doesn't exist in distributed systems (see the two-generals problem) — what we really build is exactly-once semantics (EOS).

    Key insight:

    • Idempotency is a property of the operation.
    • Deduplication is a mechanism.

You often use dedup to compensate for an operation that isn't naturally idempotent.

The Five Patterns

1. Idempotency Keys (the API-layer standard)

This is the pattern Stripe, PayPal, and basically every serious payments/write API use. The client generates a unique token per logical operation (a UUID) and sends it with the request. The server stores key → result and, if it sees the same key again, returns the cached result instead of re-executing anything.

Idempotency Keys (the API-layer standard)

Pros

  • Client and server agree on intent explicitly — no guessing.
  • Works for any write API, not just event pipelines.
  • Simple mental model, easy to reason about in code review.

Cons

  • Only as good as the client's discipline — a client that generates a new key on every retry defeats the whole point.
  • Needs an atomic "reserve" step (SET NX), not just check-then-set, or two simultaneous retries both slip through.
  • Requires TTL management — keys can't live forever.

Use it for: any public or internal write API where a client might retry — payments, order creation, "send this webhook," anything money- or side-effect-adjacent.

2. Database Unique Constraints / Upsert

The least glamorous, most bulletproof option. Define a natural or composite unique key and let the database reject or merge the duplicate for you.

-- Postgres
INSERT INTO orders (order_id, ...) VALUES (...)
ON CONFLICT (order_id) DO NOTHING;

-- or, to merge instead of drop:
ON CONFLICT (order_id) DO UPDATE SET ...;

DynamoDB uses conditional writes (attribute_not_exists), Cassandra has IF NOT EXISTS, MySQL has ON DUPLICATE KEY UPDATE — same idea everywhere.

Pros

  • Unforgeable — this is the one mechanism that can't silently fail the way a cache can (eviction, crash, TTL expiry).
  • No extra infrastructure; you already have the database.
  • Great last line of defense even when something upstream (Redis, idempotency key) also exists.

Cons

  • Doesn't scale gracefully to billions of ephemeral keys — index bloat becomes real.
  • Adds write contention/index cost, especially under high write concurrency.
  • Only catches duplicates at the point of persistence — doesn't stop wasted work upstream (you still called the payment API twice, you just didn't record it twice).

Use it for: the correctness guarantee behind any dedup story. Most production systems treat this as ground truth and everything else (Redis, idempotency keys) as an optimization layer on top of it.

3. Redis-Based Dedup Store (the streaming workhorse)

Before a consumer executes a side effect (DB write, external API call, notification), it checks-and-sets an atomic marker.

Redis-Based Dedup Store

Pros

  • Extremely fast (single round trip), cheap to run at high throughput.
  • Decouples dedup logic from the broker or the DB schema.
  • Works uniformly across Kafka, SQS, Pub/Sub, Kinesis — anything with at-least-once delivery.

Cons

  • TTL sizing is the whole game. Too short and duplicates outside the window slip through; too long and memory grows unbounded. Size it to your real max redelivery/replay window, not "forever."
  • Not durable by default — a Redis crash without AOF/persistence can lose your dedup state and let duplicates through right when things are already unstable.
  • Naive GET then SET isn't atomic — always use SET NX, or you get a race window under concurrent consumers.

Use it for: stream consumers that need to avoid re-triggering expensive or non-idempotent side effects on message redelivery. This is the single most common pattern in Kafka/Kinesis/Pub-Sub consumer code.

4. Bloom Filters (dedup at massive scale)

When the universe of possible keys is enormous — billions of URLs, ad impressions, log lines — storing every key exactly becomes the bottleneck. A Bloom filter trades a small, tunable false-positive rate for huge memory savings: a bit array + k hash functions, O(k) membership check.

Bloom Filters (dedup at massive scale)

The key asymmetry: a Bloom filter never gives a false negative ("not seen" is always correct), but it can give a false positive ("seen" when it wasn't). That makes it safe as a fast pre-filter, risky as your only source of truth.

Pros

  • 10–100x smaller memory footprint than storing every key.
  • Constant-time lookups regardless of how many keys you've seen.
  • Battle-tested — this is literally the textbook motivating example for Bloom filters (web crawler "have I visited this URL" checks).

Cons

  • False positives silently drop legitimate events unless you back it with an exact check.
  • No deletion support (a Cuckoo filter solves this if you need a sliding/evicting window).
  • Needs periodic rebuild/snapshotting to survive restarts cleanly.

Use it for: clickstream analytics, ad impression dedup, crawler URL-seen-before checks, log ingestion — anywhere the key space is unbounded and an occasional false positive is an acceptable cost.

5. Kafka Idempotent Producer / Exactly-Once Semantics

Solves a specific, narrow problem: a producer retries a send (due to a timeout or transient error) and risks writing the same message twice. Kafka assigns each producer a Producer ID and a per-partition sequence number; the broker tracks the last few sequence numbers and silently discards a retried duplicate.

Kafka Idempotent Producer / Exactly-Once Semantics

This is enabled by default (enable.idempotence=true) and costs almost nothing — it should basically always be on. Full Kafka transactions (used by Kafka Streams for read-process-write topologies) extend this to atomically commit offsets and output together, giving true exactly-once across a topology — at the cost of extra coordinator round trips and noticeably more latency.

Pros

  • Solves producer-retry duplication essentially for free.
  • Transactions extend the guarantee across a full consume-transform-produce cycle.
  • No application code changes needed for the basic idempotent-producer case.

Cons

  • Only solves broker-level duplication from producer retries — it does nothing for two different producers emitting the same logical event, or a consumer's downstream side effect firing twice. This is the single most common confusion in interviews and design docs.
  • Full transactions add real latency and throughput cost — don't turn them on by default the way you would idempotence.

Use it for: any Kafka producer, always (idempotence). Reach for transactions specifically when you have a read-process-write pipeline where partial application would corrupt downstream state.

Trade-offs at a glance

PatternAccuracyCostBest for
Idempotency KeyExactLow, needs client cooperationWrite APIs (payments, orders)
DB Unique ConstraintExact, unforgeableMedium, index overhead at scaleGround-truth correctness layer
Redis Dedup StoreExact within TTL windowVery low latency, not durable by defaultStreaming consumer side effects
Bloom FilterApproximate (false positives only)Very low memoryBillions-scale, tolerant of rare drops
Kafka Idempotent ProducerExact (protocol-level)NegligibleAny Kafka producer, always on

Pitfalls worth remembering

  1. Layer confusion
    Kafka's idempotent producer stops that producer from double-writing on retry — it does not stop two producers, or a consumer's API call, from duplicating. Know which layer you're actually protecting.

  2. Unbounded dedup stores
    "We'll just keep every ID in Redis forever" always turns into an incident. Tie your TTL to your actual max redelivery/replay window.

  3. Check-then-set race conditions
    GET then SET is not atomic. Use SET NX or a DB conditional write so the check and reserve happen in one step.

  4. Treating a Bloom filter as ground truth
    It will drop real events at scale. Pair it with an exact check on the rare "possible hit" path, or accept the loss rate as an explicit product decision.

  5. Believing exactly-once delivery is achievable
    It isn't, in an asynchronous distributed system. What you can build — and what to say out loud in a design review — is exactly-once processing semantics on top of at-least-once delivery.

Cheat sheet

ScenarioReach for
Public write API (payments, orders)Idempotency-Key + atomic reserve-then-process store
Any Kafka producerenable.idempotence=true — default-on, no reason not to
Kafka read-process-write pipelineKafka Transactions
Streaming consumer with side effectsRedis dedup store (SET NX EX)
Billions of keys, some loss tolerableBloom filter + exact confirm on hit
Final correctness guaranteeDB unique constraint / upsert

Rule of thumb: pick one mechanism close to the client (idempotency key or producer idempotence) and one close to the database (unique constraint) — that combination alone covers the vast majority of real production duplicate scenarios. Add Redis or a Bloom filter only when volume or latency actually demands it.


Closing thought

None of this is about chasing "exactly-once delivery" as a transport guarantee — that ship sailed the moment you accepted at-least-once for reliability. What actually matters is picking, deliberately, which layer owns the guarantee: a client-generated key, a broker's retry protection, a consumer-side cache, or the database's own constraints.

The systems that get bitten by duplicates aren't usually missing a dedup mechanism entirely. They're missing an explicit answer to which layer is supposed to catch it — so when one layer fails silently, nobody notices until the customer does.