Tag Observability

When One Region Wasn’t Enough

For a long time, one region was enough. The application ran there. The database ran there. The message infrastructure ran there. Backups existed. Monitoring was in place. We had redundancy inside the region. From an infrastructure perspective, the system looked…

When a Third-Party API Became Our Bottleneck

The application was healthy. CPU was normal. Memory was stable. The database was performing well. Internal services were responding within their expected latency. And yet one part of the system was getting slower. Payment requests. At first, we looked at…

How We Handled a Payment System Incident at 2AM

It was 2AM. Most of the engineering team was asleep. The payment system was processing transactions normally. Then an alert fired. Duplicate charges were increasing. Not thousands. Not enough to immediately bring the entire system down. But enough to tell…