Updated
Site down at 2 a.m.: the first 15 minutes when you're the only one on call
A user message or an alert wakes you up: the site won't load. You have a laptop, maybe just a phone, and nobody to hand it to. The tempting move is to reboot the whole server. That can wipe out the evidence of what went wrong, and it can interrupt a database in the middle of a write. This guide gives you a fixed order for the first 15 minutes: confirm the impact, check the machine is alive, check the entry point and the app, read only the logs from the failure window, then write three lines before you change anything. The examples assume a typical Linux VPS with systemd, Nginx and Docker; swap in your own paths, service names and ports. This page is the first-pass triage; the deeper steps live in the linked guides.
1. Minutes 0–2: confirm the impact from outside
Run these on your own computer, not on the server. Replace example.com with your domain:
curl -sS -o /dev/null -w '%{http_code} %{time_total}s\n' --max-time 10 https://example.com/
dig +short example.com A
dig +short example.com AAAA
000or a timeout: no HTTP response came back at all. Think DNS, network, firewall or security group, TLS handshake, or an unreachable machine.502/504: the gateway (Nginx or a CDN) answered but got no valid response from upstream. Focus on the application; see Nginx 502 Bad Gateway.500: the request reached the application and the application failed. Go to its logs.200: it works from where you are. The problem may be limited to one region, one path or some users; ask which page and which network.digreturns an address you don't expect for your server or CDN: check whether DNS records changed or the domain expired.
What this does not tell you: one vantage point is not every user, and if a CDN sits in front, the status code may come from the CDN's error page rather than your origin. A second try over mobile data is a cheap cross-check.
2. Minutes 2–5: is the machine alive?
Try to SSH in. If SSH fails, go to your cloud provider's console: instance status, monitoring graphs, security group rules, and the web-based VNC or serial console if needed. In that situation OpsMate can't reach the machine either. If you can log in:
uptime
nproc
free -h
df -h
df -i
Compare the load averages from
uptimewith the core count fromnproc. Load that stays well above the core count means work is queueing for CPU or disk I/O.In
free -h, read theavailablecolumn, notused. Linux uses spare memory for cache, so highusedalone is not a memory shortage.A filesystem at 100% in
df -h, or inodes at 100% indf -i, means writes fail, and databases and apps often fall over as a result. Go to Linux disk usage suddenly spikes and don't start deleting files yet.
What this does not tell you: which process is responsible, or what the peak looked like when things broke. These are snapshots of right now.
3. Minutes 5–8: are the entry point and the app running?
systemctl --failed --no-pager
systemctl status nginx --no-pager
sudo ss -ltnp
docker ps -a --format 'table {{.Names}}\t{{.Status}}\t{{.Ports}}'
systemctl --failedlists units that failed. An empty list doesn't clear your app, because not every app runs under systemd.ss -ltnpshows what is listening on 80/443 and on your app's port (you need sudo to see process names). Nginx up but nothing on the app port is the classic 502 pattern.A container in
RestartingorExitedindocker ps -a: follow Docker container keeps restarting.
Then hit the app directly from the server, bypassing DNS, CDN and Nginx. The port and path are examples; use your real upstream and a safe health-check path:
curl -sS -o /dev/null -w '%{http_code}\n' --max-time 5 http://127.0.0.1:3000/
If this works locally but the site fails from outside, look at Nginx, the firewall, TLS certificates or DNS. If it fails locally too, start with the app. A listening port is not proof of a healthy app: it can accept connections and still hang or error on every request.
4. Minutes 8–12: read only the logs from the failure window
Ask yourself roughly when it started, then read around that time instead of scrolling from the top:
journalctl -p err --since "30 min ago" --no-pager | tail -n 100
sudo tail -n 100 /var/log/nginx/error.log
docker logs --since 30m --tail 200 --timestamps YOUR_CONTAINER
sudo dmesg -T | grep -i -E 'out of memory|killed process' | tail -n 20
journalctl -p errshows priority err and above. Many apps write errors to their own log files, so an empty result here proves little.Your Nginx
error.logpath depends on your config. Lines likeconnect() failed (111: Connection refused)point at an upstream that isn't listening.In
docker logs, look for the first error before the exit, not just the last line. ReplaceYOUR_CONTAINERwith the real name.Out of memory/Killed processindmesgmeans the kernel killed something at some point. It only counts as evidence if the timestamp lines up with the outage. If logs mention a full disk, Docker logs filling your disk may also apply.
Before you share or submit any of this output, redact passwords, tokens, connection strings, customer emails and internal addresses.
5. Minutes 12–15: write three lines before you touch anything
Before any restart or edit, write this down, in your notes or your team chat:
Symptom: when it started, who is affected, what the outside
curlreturned.Evidence: which command and which output line support your theory, and what is still unconfirmed.
Smallest action + rollback: the one thing you plan to do, how you'll verify it worked, and how you'll undo it if it doesn't.
Common smallest actions. Each one is a change you decide on, not a diagnostic step:
Restart only the one failing service or container, not the whole machine. A restart clears in-memory state, so copy the output from the steps above somewhere safe first.
Before editing Nginx config, back up the original file, run
sudo nginx -tto check syntax, then plan the reload. If it goes wrong, put the backup back and reload again.If the disk is full, remove only what you've confirmed is disposable; snapshot or back up business data first.
If you deployed recently and the evidence points at the new code, rolling back to the previous release is usually safer than a 2 a.m. hotfix, as long as you know whether the database migration can be reversed.
Leave a full reboot for last. Before rebooting, check whether
/var/log/journalexists. If it doesn't, the journal may live only in memory and this boot's logs will be gone after the reboot.
After the action, repeat the outside curl from step 1 and the resource checks from step 2, confirm recovery, and note the time.
Doing this in OpsMate
OpsMate puts an SSH terminal and an AI assistant on the same server page. You can type every command above straight into the terminal. When you can't remember the right command, describe the problem to the AI in plain language. The AI proposes troubleshooting commands, and beyond logs it can use ps, df, ss, docker and journalctl for checks. After a command has run, click Analyze to get a conclusion from its output. The commands and output stay in the terminal so you can check them and review them later. When you're away from your desk, patrol checks and Telegram alerts can tell you about a problem first. Receiving Telegram alerts doesn't require storing credentials in the cloud; only running SSH on the server remotely from Telegram does. See Server alerts in Telegram. For what that looks like on a phone, see Troubleshoot your server from Telegram.
The limits, plainly:
If you can't SSH into the machine, neither can OpsMate. Start with your cloud provider's console.
The AI is mainly for troubleshooting: dangerous commands are blocked. Low-risk fixes such as restarting a service or rotating logs run automatically by default. The rollbacks and config edits in this guide are yours to decide on. For what OpsMate does and doesn't do on its own, see the FAQ.
An AI summary is a starting point. Check it against the command output; it is not proof of a root cause.
When you click
Analyze, the command output is sent to cloud AI for analysis. What the desktop app keeps on your own machine by default is your SSH credentials. If the output contains customer data or secrets, don't clickAnalyzeon it; redact it yourself and give the AI the redacted excerpt.
Illustrative example
Illustrative example (not a real customer case): at 2 a.m. the outside curl returns 502. SSH works and load is low. systemctl status nginx shows Nginx running, but ss -ltnp shows nothing listening on the app's port 3000. docker ps -a shows the app container stuck in Restarting, and the first error in docker logs is a failed database write. Back in df -h, the root filesystem is at 100%. The 502 is only the symptom; the disk is the problem. Follow the disk guide to find out whether logs or business data filled it, remove only what's confirmed disposable, then watch whether the container recovers. Rebooting the whole machine in minute one would have left the disk just as full, the outage would have come back a few minutes later, and there would have been less evidence to go on.
AI diagnostic prompt
The site stopped loading about 30 minutes ago. Without changing anything on this server, check load, available memory, disk space and inodes, failed systemd units, listening ports, Docker container status, and system, Nginx and application errors from the last 30 minutes. Separate confirmed facts from hypotheses and list what information is still missing. Do not restart services, delete files or change configuration. If a change is needed, propose the smallest step with its risks, how to verify it and how to roll it back, and wait for my approval.
Try OpsMate
500 free AI calls per month and unlimited servers. The desktop app keeps your SSH credentials on your own machine by default.
Need help interpreting the evidence?
OpsMate helps developers and operators investigate with AI. Review the evidence. After you click Analyze, the command output is sent to cloud AI for analysis; redact sensitive information first.