Tag Retry Patterns

When a Third-Party API Became Our Bottleneck

The application was healthy. CPU was normal. Memory was stable. The database was performing well. Internal services were responding within their expected latency. And yet one part of the system was getting slower. Payment requests. At first, we looked at…

When Our APIs Became the Architecture

At first, the API was just an interface. A way for one component to call another. Nothing more. Then the system grew. The monolith became multiple services. Teams became independent. Deployments became more frequent. And suddenly almost every important business…

How We Handled a Payment System Incident at 2AM

It was 2AM. Most of the engineering team was asleep. The payment system was processing transactions normally. Then an alert fired. Duplicate charges were increasing. Not thousands. Not enough to immediately bring the entire system down. But enough to tell…

When Retries Made the Problem Worse

A downstream service started timing out. So we did what seemed like the responsible thing. We added retries. The thinking was straightforward. If a request fails because of a temporary network problem or a momentary downstream slowdown, try again. The…