A few weeks ago I wrote a whole post arguing that CURLOPT_TIMEOUT => 0 is never the right value. Then last week I opened the same class of file and typed CURLOPT_TIMEOUT => 0 with my own hands.
I stand by both. Here’s the bit I’d left out.
The feature that broke the assumption
We’ve got a Claude-powered concierge plugin on one of the travel sites. It’s been running happily for months — a visitor chats, the model answers, the answer streams back token by token over SSE. Turns take ten, twenty, maybe forty seconds when the model goes off and does a few web searches. Well inside anything you’d call a timeout.
The new bit is an advisor-only mode that builds a full itinerary. A dozen or so guided questions, then the model generates the whole trip as one structured JSON document — days, hotels, restaurants, transport, budget — which the server renders into a formatted spreadsheet the agent can hand to a client.
That turn is not like the other turns. It’s a large output, and it’s doing live searches along the way to check that the hotels it’s recommending actually exist and are actually bookable. Several minutes of legitimate work.
The first few builds came back as 502s.
It failed in three places and lied in all of them
The interesting part wasn’t that it failed. It’s that each place it failed looked like a different bug.
The blocking fallback path — the one we use when the SSE stream drops — was calling wp_remote_post with a 180 second timeout. A number I’d picked months earlier, with a smug little comment next to it about how that gave us plenty of headroom without being absurd. It absurded. Three attempts, three timeouts, a 502.
PHP’s own execution cap was at 600 via set_time_limit(). Generous for a chat turn. Not generous for this.
And the streaming path had this:
CURLOPT_TIMEOUT => 300, CURLOPT_CONNECTTIMEOUT => 15,
Which is exactly what my previous post told you to do, and it was the one doing the most damage. At 300 seconds, cURL hung up on a stream that was still delivering. Data flowing, model mid-sentence — and we killed it, because a stopwatch said so.
The layers matter here. Any of these can end your request, and the shortest one always wins. Raising one and not the others just moves which error message you get.
A total timeout is the wrong instrument
This is the thing I hadn’t properly thought through when I wrote the last post.
CURLOPT_TIMEOUT answers one question: how long has this been going? For a normal request/response that’s a perfectly good proxy for health, because the only two states are “waiting” and “done”. If you’ve been waiting half a minute for a hotel lookup, something is wrong. No ambiguity.
A stream has a third state. It can be working — bytes arriving steadily, just a lot of them. A total timeout can’t see the difference between that and a socket that died in silence. Both look like “still open, still not finished”.
So you get to pick your poison. Set it low and you kill healthy long generations. Set it high enough that a real build survives, and you’ve also given a genuinely dead connection permission to sit there holding an FPM worker for however long you picked. Which is the exact failure I wrote the last post about. There’s no value that gets you out of that — it’s the wrong question.
Ask the right question instead
cURL has a second timeout that almost nobody uses, and it’s the one that fits:
CURLOPT_TIMEOUT => 0, // no total cap CURLOPT_CONNECTTIMEOUT => 15, CURLOPT_LOW_SPEED_LIMIT => 1, // bytes per second floor CURLOPT_LOW_SPEED_TIME => 120, // sustained for this long, then give up
Read together, those last two say: if the transfer drops below one byte per second and stays there for two minutes, abort. Not “you’ve taken too long”. “You’ve gone quiet.”
That’s the actual thing I care about. A build that streams for six minutes never trips it, because it’s never silent for two of them. A connection that dies mid-generation trips it inside a couple of minutes and hands the worker back, which is what the total timeout was there to do in the first place. (It’s an averaged, polled check, so don’t read 120 as a precise deadline.)
CURLOPT_CONNECTTIMEOUT stays. That one’s still measuring a phase with only two states — connected or not — so a stopwatch is fine.
The catch, and it’s a real one
This only works if the far end is chatty when it’s busy.
Anthropic’s streaming API sends periodic ping events, including through long tool calls, so the socket is never actually idle while the model is thinking. That’s what makes the low-speed watchdog meaningful — silence genuinely means something is broken.
If your upstream goes properly quiet during long operations, this won’t work. You’ll pick 120 and discover the server’s legitimate think-time is 130. Check what the protocol does before you lean on it. Some APIs document their keepalive behaviour; for the rest, curl -N -v and a stopwatch will tell you in about a minute.
Getting the other layers out of the way
Fixing cURL alone gets you nothing if PHP kills the request instead. The stack had to be lifted to match:
// The stream is watchdogged by the low-speed abort in the HTTP client, // so PHP's cap is just a backstop. It should never be the thing that fires. @set_time_limit(900);
And the blocking fallback went from 180 to 600, with a comment that’s less pleased with itself than the last one.
The ordering I’d aim for: the mechanism that can actually distinguish healthy from dead should be the one that fires first, and everything above it is a backstop set high enough to stay out of the way. If your outermost layer is the strictest, you’ve built a system where the least-informed component makes the decision.
Worth checking whatever’s in front of PHP, too. An nginx proxy_read_timeout, a load balancer idle timeout, Cloudflare — any of those will cut you off at their own number and PHP will never know it happened.
The part users see
A six-minute wait is a support ticket even when it works.
Two things went in alongside the timeout changes. A progress banner, so the screen isn’t just a spinner and a growing suspicion — and an automatic retry for the specific case where the turn comes back empty or the model re-confirms instead of generating. One forced retry with an explicit instruction, then a clear message and the option to try again by hand.
The retry is a bit of an admission, honestly. It exists because a long stream through a chain of proxies will occasionally just not make it, and no amount of timeout tuning fixes that. Better to catch it and quietly go again than to hand someone a blank screen after six minutes.
What I’d actually write on the wall now
The old rule was “always set a timeout”. Still true, and I’m not walking it back.
The refinement: match the timeout to what you’re measuring. A request that returns one payload gets a total timeout, because elapsed time is a decent proxy for health. A stream gets an idle timeout, because for a stream the thing that means “broken” is silence, not duration.
Grep your codebase for CURLOPT_TIMEOUT. Anywhere it’s sitting on a streaming call, or on anything talking to a model, ask whether the number you picked is measuring the thing you actually care about. Mine wasn’t, and I’d written a post about it.
