The hardest problem in distributed systems is uncertainty.
Imagine:
Order Service
|
| Create Order
v
Database
|
| Publish OrderCreated
v
Message BrokerWhat happens if:
- The database successfully saves the order.
- The service crashes.
- The event is never published.
Now the order exists, but downstream services never know.
Or the opposite:
- The event is published.
- The database transaction fails.
Now another service thinks an order exists when it does not.
This is the dual-write problem. The Transactional Outbox pattern addresses it by storing the business change and its event together in the same transaction, after which a separate process publishes the event. ([AWS Documentation][4])
Best structure for the article
Failure #1: The Request Timed Out — Did It Actually Fail?
Client
|
| Create Order
v
Order Service
|
| Request succeeded
X
Response lostThe client sees a timeout.
So it retries.
Now there may be two orders.
This is why retries alone are dangerous.
Microsoft's retry guidance explicitly warns that retries can repeat operations and cause unintended side effects unless the operation is idempotent. ([Microsoft Learn][5])
Idempotency: Make Repeated Requests Safe
Use:
POST /orders
Idempotency-Key: abc-123If the client sends the same request again:
abc-123 -> Already processedThe service returns the original result instead of creating another order.
AWS describes idempotent operations as having the same effect regardless of how many times they are executed. ([AWS Documentation][6])
Failure #2: Messages Can Be Delivered More Than Once
Many messaging systems use at-least-once delivery.
That means:
OrderCreated
|
v
Consumer processes message
|
X
Crash before acknowledgement
|
Message delivered againIf your consumer isn't idempotent:
OrderCreated -> Charge card
OrderCreated -> Charge card againAWS specifically recommends idempotent consumers when duplicate events are possible. ([AWS Documentation][4])
Failure #3: The Message Always Fails
Message
|
v
Consumer fails
|
v
Retry
|
v
Retry
|
v
Retry
|
X
Dead Letter QueueA Dead Letter Queue prevents permanently failing messages from blocking normal processing. AWS and Azure both recommend DLQ/error-handling approaches for messages that cannot be successfully processed. ([AWS Documentation][7])
Failure #4: Retry Storms
This is another strong section.
When a dependency starts failing:
1,000 clients
|
v
All retry immediately
|
v
Dependency receives even more traffic
|
v
Dependency gets worseUse:
- exponential backoff
- jitter
- retry limits
- circuit breakers
Retry is useful for transient failures, but uncontrolled retries can amplify an outage. ([Microsoft Learn][5])
The most impressive ending
I would end with something like:
A distributed system is not reliable because its services communicate successfully when everything is healthy. It is reliable because the system knows what to do when messages are duplicated, requests time out, services crash, and events arrive twice.
Retries require idempotency. Events require failure handling. Database writes require reliable publishing. The architecture only becomes interesting when something fails.
That would be a very strong ending.
My honest recommendation for your portfolio
These two blogs should have different personalities:
LLD
Calm, thoughtful, software-engineering focused.
Theme:
“How do you design code so future changes don't destroy it?”
Distributed Systems
Practical, failure-focused, systems engineering.
Theme:
“What happens when everything does not work perfectly?”
Together with your existing RAG article, they give your blog a strong progression:
RAG
↓
AI Engineering
LLD
↓
Software Engineering
Distributed Systems
↓
Backend Engineering
AI Agents
↓
Agentic Systems
On-Chain Agents
↓
AI + Blockchain
Solana Architecture
↓
Blockchain Engineering