The API call that never gave up (and took the site down with it)

Title card: The API call that never gave up (and took the site down with it)

The site wasn’t down. That was the confusing bit.

It answered. Slowly, grudgingly, forty-odd seconds for a page that normally paints in one — but it answered. No 500, no white screen, nothing in the error log worth a second look. Just a busy travel site that had quietly turned to treacle, right when the agents were trying to use the admin.

I want to write about this one because the root cause is a single number, and it’s a number that’s probably sitting in your codebase too.

What it looked like

The first reports came from the back office. Editors saving a post, waiting. The dashboard loading a widget that never showed up. Front-end pages were sluggish too, but the admin is where it hurt — people trying to do their jobs and getting a spinner.

The tell was the shape of it. It wasn’t slow the way a heavy query is slow, evenly, all the time. It came in waves. Fine for a while, then a stretch where everything jammed at once, then fine again. That pattern — good, then a cliff, then good — is worth remembering. It usually means something is being held, not something being computed.

Where the time was going

PHP-FPM has a slow log, and if you’re not already writing to one, this is your nudge to turn it on. It dumps a stack trace for any request that runs past a threshold you set. Half a day of “it’s slow, no idea why” collapses into one file.

Every stuck request pointed at the same line. curl_exec(), inside the function that talks to the travel supplier’s API. Same frame, over and over, across a dozen simultaneous requests that were all sitting there waiting on the network.

The supplier’s API was having a rough afternoon on their end. Not down — just slow to respond. And our code, it turned out, was willing to wait forever.

The one number

Here’s the cURL setup, more or less as it was:

$options = array(
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_ENCODING       => "",
    CURLOPT_MAXREDIRS      => 10,
    CURLOPT_TIMEOUT        => 0,
    CURLOPT_FOLLOWLOCATION => true,
    CURLOPT_HTTP_VERSION   => CURL_HTTP_VERSION_1_1,
    CURLOPT_CUSTOMREQUEST  => "GET",
);

CURLOPT_TIMEOUT => 0. Zero doesn’t mean “no time”. It means no limit. Wait as long as it takes. That’s the cURL default, and it’s the value that quietly rides along in half the snippets you’ll copy off the internet or export out of an API testing tool. Nobody sets it to zero on purpose. It’s just what you get if you never think about it.

On a good day it’s invisible. The API answers in 200ms and you never learn what that zero does. The bad day is the whole point of a timeout, and we didn’t have one.

Why one slow API stalls the whole box

This is the part that turns an annoyance into an outage.

PHP-FPM runs a fixed pool of worker processes. Say you’ve got twenty. Each request in flight holds one worker until it finishes. When the API goes slow and every call to it blocks on curl_exec(), those workers don’t come back. They sit there, waiting on a socket, doing nothing.

max_execution_time won’t save you the way you’d hope — cURL’s wait doesn’t always count against it cleanly, so a call can hang well past where you’d expect PHP to step in. And even a 60-second ceiling is an eternity when workers are the scarce resource. Enough slow calls arrive, the pool fills, and now every other request — pages that never touch the API at all — is queued behind a worker that won’t free up. One degraded third party, and the whole site is holding its breath.

The fix

Two lines.

CURLOPT_TIMEOUT        => 15,
CURLOPT_CONNECTTIMEOUT => 5,

CURLOPT_TIMEOUT caps the whole operation. Fifteen seconds and cURL gives up, hands you an error, and — crucially — hands the worker back. CURLOPT_CONNECTTIMEOUT caps just the connect phase at five, so if the supplier’s box is unreachable rather than merely slow, you fail fast instead of burning the full fifteen on a handshake that’s never coming.

Pick the numbers to suit the call. A user-facing lookup that needs to feel snappy might get five. A background sync you don’t mind waiting on can have thirty. The one value that’s never right is zero.

The other half — and don’t skip this — is deciding what happens when the timeout fires. A timeout that turns an infinite hang into an unhandled fatal hasn’t helped anyone. The function needs to catch the cURL error and return something the caller can cope with: an empty result, a cached response, a “temporarily unavailable” that the template knows how to render. Fail, but fail on your own terms.

What I’d take from it

Every outbound call in your app is a promise from someone else’s server, and other people’s servers have bad days. If you don’t set a timeout, you’ve handed a stranger the power to freeze your site by being slow. Not by attacking you. Just by being slow.

Two things I’d do today if I were you. Grep your codebase for CURLOPT_TIMEOUT — and while you’re there, look at your wp_remote_get and Guzzle calls too, they all default to their own flavour of “wait a while”. And turn on the PHP-FPM slow log before you need it, not after. When the pool’s jammed and the phone’s ringing, that file is the difference between a diagnosis and a guess.

The site never actually went down. It just forgot how to give up. Turns out those are close enough to the same thing when the workers run out.

Leave a Reply

Your email address will not be published. Required fields are marked *