the call is always the same, website not loading, is it hacked. nine times out of ten it is not hacked and the client is on hotel wifi, so the first move is to not touch anythign. i ask what they see, the exact words on the screen, and whether anyone else has reported it. then the site goes up on my phone with wifi off, becuase a mobile connection is a different resolver and a different route, and that one check splits the problem in half. if it loads on cellular but not on their machine it is local to them, a stale dns cache or an isp resolver still holding the old record or an office firewall being clever. ipconfig /flushdns has fixed more of these than i would like to admit. down everywhere and i go to dns first, dig the a record against 1.1.1.1 and against the authoritive nameserver, and if those two disagree somebody changed the zone recently and there is an awkward conversation coming. next is whether the box answers at all, curl -I against the domain and then against the raw ip with a host header, because if the ip responds and the name does not you are looking at dns or the edge, and if neither responds the server is off or the firewall ate you. a 502 or a 504 means the web server is alive and the upstream is dead, usually php-fpm sat at max_children after a leak or a bad plugin update, and restarting the pool buys you the hour you need to work out which one. a plain timeout with no status code at all is a different animal, that is network or host level. by that point i am on the host status page, and honestly that should be step two, i have burned forty minutes debugging a scheduled maintenance window more than once. cloudflare status too if they sit behind it, a 522 with a healthy origin is not yours to fix. last stop is always the client, did anyone install anything yesterday, did a card expire, did the domain lapse. a suprising number of dead sites are an unpaid invoice somewhere in the chain, and no amount of curl output will ever show you that.
Source: r/WebsiteHealth · by /u/Latter_B2