1. Give each event a stable identity
Require a unique event ID from the producer and carry it unchanged through retries. Include the event type and schema version so the consumer can validate what it received. A business key such as an order ID is not enough if multiple legitimate updates can occur for that order.
2. Make processing idempotent
Keep a record of processed event IDs with a uniqueness constraint. When updating a database, store the result and mark the event processed in the same transaction. If the ID already exists, acknowledge the duplicate without repeating the update. Retain IDs for at least as long as the broker can redeliver events, including replay windows.
PostgreSQL example: with processed_events.event_id and shipments.order_id unique, a duplicate ID inserts no new shipment. The shipment and deduplication record commit together:
BEGIN;
WITH first_seen AS (
INSERT INTO processed_events (event_id)
VALUES ('evt-1042')
ON CONFLICT (event_id) DO NOTHING
RETURNING event_id
)
INSERT INTO shipments (order_id, event_id)
SELECT 'order-1042', event_id FROM first_seen;
COMMIT;
External side effects need additional care. If a consumer calls a payment or email API, pass a stable idempotency key if that API supports one. A local processed-ID table alone cannot guarantee exactly-once execution across a separate service call and a database commit.
3. Retry temporary failures, not bad messages
Classify timeouts and temporary service failures as retryable. Use bounded exponential backoff with jitter to avoid overwhelming a recovering dependency. Validation failures and unsupported schema versions should not be retried indefinitely: send them to a dead-letter queue with a reason and an event ID. Acknowledge a message only after its effects are durable, following the broker's delivery semantics.
4. Make dead letters recoverable
Alert on a growing dead-letter queue and record the original payload, failure reason, and attempt count without leaking secrets. Define who investigates failures and how to replay after the cause is fixed. Replay should pass through the same idempotent consumer rather than bypassing its checks.
5. Test the failure boundaries
- Deliver the same event twice and verify only one business update occurs.
- Fail after the update but before acknowledgement; confirm redelivery does not repeat the effect.
- Simulate a temporary dependency outage and verify backoff stops at the configured limit.
- Send an invalid event and confirm it reaches the dead-letter queue with a useful reason.
- Replay a resolved dead letter and confirm the consumer processes it once.
A retry policy without idempotency moves the failure from a missed event to a duplicate side effect.
Use case: avoid duplicate shipments after a timeout
An order service emits a ReadyToShip event with a stable event ID. The warehouse consumer creates a shipment, but its acknowledgement to the broker times out. The broker delivers the event again. If the consumer blindly creates another shipment, the customer receives two parcels for one order.
Store the event ID and shipment record together in the warehouse database transaction, with a unique constraint on the event ID. On redelivery, look up the existing shipment and acknowledge the event without creating another. If a separate carrier API creates labels, send the same idempotency key on every attempt where supported; otherwise reconcile carrier state before retrying that external step. After bounded retries, send unresolved events to a dead-letter queue for investigation and controlled replay.