Two JSON files cached side by side from the same feed, pulled by the same importer, using id and parentId in opposite senses. Nothing in either file tells you which one you’re holding. That single wrong assumption produced two failure modes: an admin dropdown that threw a TypeError the day it appeared, and a second one that had been silently collapsing twelve hundred ships into one row per cruise line for years. The loud one got reported in an afternoon. The quiet one just looked like a short list, and nobody counts the options in a dropdown. There’s a third copy of the same loop still cached on the front end.
The API call that never gave up (and took the site down with it)
The site wasn’t down, which was the confusing part. It answered, forty seconds at a time, in waves. The PHP-FPM slow log pointed every stuck request at the same curl_exec(), talking to a travel API that was having a bad afternoon. Our cURL options carried CURLOPT_TIMEOUT set to zero, which doesn’t mean no time — it means no limit. Slow calls held FPM workers until the pool filled, and then pages that never touched that API queued up behind them too. The fix is a total timeout, a connect timeout, and actually handling the error when it fires. Leave it at zero and you’ve handed a stranger the power to freeze your site by being slow.
Read more "The API call that never gave up (and took the site down with it)"
You can’t attach an IAM role to a Lightsail box
On EC2 you’d attach an IAM role and never see a key. Lightsail can’t do that, and during the migration both servers needed to reach S3. The setup: a dedicated IAM user with a tightly scoped policy, keys kept on the operator laptop only, and credentials handed over as environment variables for a single SSH command, never written to disk. Writing it up, I found a few things I’d got wrong, not least that the servers were getting far more permission than they needed. The fix is fifteen-minute, single-bucket credentials from STS.
The wrong canonical that CORS-blocked every font
Pages loaded, images loaded, but every font and icon was blocked by CORS. The page came from www and the fonts from the bare domain, because a canonical constant pinned in wp-config had forced the wrong hostname into every asset URL. Fonts are the assets browsers fetch in CORS mode, so they were the only ones to complain. Working out the right canonical automatically then went wrong twice, once by trusting a database value that had been silently overridden for years. What stuck is dull. A human writes it down, and a probe checks where the live redirect chain actually lands.
Read more "The wrong canonical that CORS-blocked every font"
A live static IP and a site nobody could reach
Phase seven of the migration moved the static IP to the new server. Phase eight couldn’t get a certificate. The site wasn’t reachable, and it hadn’t been before I touched it either: the domain’s DNS zone had no A records at all. Every preflight check had been about the server and none about the road to it. Preflight now resolves the domain through a public resolver before anything expensive happens, and for Cloudflare-fronted sites it checks Cloudflare can actually reach an origin. Check the assumptions your last step depends on, not just your first.
300 host-key prompts I refused to click
Done by hand, the fleet migration would have meant accepting an SSH host-key prompt about 300 times, and nobody reads the fingerprint on the fortieth. So the provisioning script fires a probe it expects to fail, scrapes the fingerprint out of plink’s refusal, and pins it for every connection after. It’s still trust-on-first-use, same as typing yes, but it happens seconds after creating the server through an authenticated API call. The catch turned up later. A server has more than one host key, and which one you get shown depends on what your own client has cached.
The 1 GB box that swap-thrashed itself to death
A backup plugin with a 400 MB appetite, a 1 GB server still serving live traffic, and a box that swap-thrashed until real visitors started timing out. Then the same plugin hung for six hours restoring a big multisite, doing row-by-row URL replacement across millions of rows. Too heavy at one end, too slow at the other. The fleet migration ended up with three ways to move a site, picked automatically: the plugin for ordinary sites, a streaming tar and mysqldump for the tiny boxes, and a direct database import for the giants.
Success: backup complete. So where’s the file?
All-in-One WP Migration printed “Success: Backup complete.” and wrote nothing. No archive, and no error on the command line. The admin screen told the real story: the plugin couldn’t create the guard files it drops into its backups folder, so it quietly gave up before writing a byte. On Bitnami that folder belongs to the bitnami user, but PHP runs as daemon, the group, with no write bit. The fix is the right owner plus chmod 2775, baked into the migration tooling so it never comes up again. A success message is a claim. Check the file is actually there.
Off Bitnami: migrating a WordPress fleet to managed Lightsail
Bitnami’s WordPress image on Lightsail is on its way out, which left me with roughly sixty live sites sitting on a foundation with a use-by date. This is the overview of moving all of them to AWS’s own managed blueprint. One site at a time, through a ten-phase pipeline that records every finished step to disk so a re-run picks up exactly where it broke. Old servers get stopped, not deleted, and sit there for a week in case something turns up. The constraints that shaped it, from no IAM roles on Lightsail to 300 host-key prompts, each get their own post in the series.
Read more "Off Bitnami: migrating a WordPress fleet to managed Lightsail"
I fixed this in August. It came back in September.
In August a forms plugin couldn’t write under uploads on a freshly migrated server, so I wrote a repair script for uploads. In September the migration plugin couldn’t “create” a file that already existed, nowhere near uploads. Same cause both times. Files unpacked by the admin user came out 755 and 644, and setgid on the parent carries the group across but not the permission bits. The real fix is setgid plus a umask of 002 at provision time, with a wider repair sweep for boxes already out there. When a repair script takes a path argument, be suspicious of the path.
Read more "I fixed this in August. It came back in September."
