Tag Fault Tolerance

When Events Became the Backbone of the System

The system was becoming increasingly difficult to change. Every time we added a new capability, another service needed to know about it. A payment completed. The ledger needed to update. Notifications needed to be sent. Fraud analytics needed the transaction.…

How We Handled a Payment System Incident at 2AM

It was 2AM. Most of the engineering team was asleep. The payment system was processing transactions normally. Then an alert fired. Duplicate charges were increasing. Not thousands. Not enough to immediately bring the entire system down. But enough to tell…

When Retries Made the Problem Worse

A downstream service started timing out. So we did what seemed like the responsible thing. We added retries. The thinking was straightforward. If a request fails because of a temporary network problem or a momentary downstream slowdown, try again. The…