Updated

Site down at 2 a.m.: the first 15 minutes when you're the only one on call

A user message or an alert wakes you up: the site won't load. You have a laptop, maybe just a phone, and nobody to hand it to. The tempting move is to reboot the whole server. That can wipe out the evidence of what went wrong, and it can interrupt a database in the middle of a write. This guide gives you a fixed order for the first 15 minutes: confirm the impact, check the machine is alive, check the entry point and the app, read only the logs from the failure window, then write three lines before you change anything. The examples assume a typical Linux VPS with systemd, Nginx and Docker; swap in your own paths, service names and ports. This page is the first-pass triage; the deeper steps live in the linked guides.

1. Minutes 0–2: confirm the impact from outside

Run these on your own computer, not on the server. Replace example.com with your domain:

curl -sS -o /dev/null -w '%{http_code} %{time_total}s\n' --max-time 10 https://example.com/
dig +short example.com A
dig +short example.com AAAA

What this does not tell you: one vantage point is not every user, and if a CDN sits in front, the status code may come from the CDN's error page rather than your origin. A second try over mobile data is a cheap cross-check.

2. Minutes 2–5: is the machine alive?

Try to SSH in. If SSH fails, go to your cloud provider's console: instance status, monitoring graphs, security group rules, and the web-based VNC or serial console if needed. In that situation OpsMate can't reach the machine either. If you can log in:

uptime
nproc
free -h
df -h
df -i

What this does not tell you: which process is responsible, or what the peak looked like when things broke. These are snapshots of right now.

3. Minutes 5–8: are the entry point and the app running?

systemctl --failed --no-pager
systemctl status nginx --no-pager
sudo ss -ltnp
docker ps -a --format 'table {{.Names}}\t{{.Status}}\t{{.Ports}}'

Then hit the app directly from the server, bypassing DNS, CDN and Nginx. The port and path are examples; use your real upstream and a safe health-check path:

curl -sS -o /dev/null -w '%{http_code}\n' --max-time 5 http://127.0.0.1:3000/

If this works locally but the site fails from outside, look at Nginx, the firewall, TLS certificates or DNS. If it fails locally too, start with the app. A listening port is not proof of a healthy app: it can accept connections and still hang or error on every request.

4. Minutes 8–12: read only the logs from the failure window

Ask yourself roughly when it started, then read around that time instead of scrolling from the top:

journalctl -p err --since "30 min ago" --no-pager | tail -n 100
sudo tail -n 100 /var/log/nginx/error.log
docker logs --since 30m --tail 200 --timestamps YOUR_CONTAINER
sudo dmesg -T | grep -i -E 'out of memory|killed process' | tail -n 20

Before you share or submit any of this output, redact passwords, tokens, connection strings, customer emails and internal addresses.

5. Minutes 12–15: write three lines before you touch anything

Before any restart or edit, write this down, in your notes or your team chat:

  1. Symptom: when it started, who is affected, what the outside curl returned.

  2. Evidence: which command and which output line support your theory, and what is still unconfirmed.

  3. Smallest action + rollback: the one thing you plan to do, how you'll verify it worked, and how you'll undo it if it doesn't.

Common smallest actions. Each one is a change you decide on, not a diagnostic step:

After the action, repeat the outside curl from step 1 and the resource checks from step 2, confirm recovery, and note the time.

Doing this in OpsMate

OpsMate puts an SSH terminal and an AI assistant on the same server page. You can type every command above straight into the terminal. When you can't remember the right command, describe the problem to the AI in plain language. The AI proposes troubleshooting commands, and beyond logs it can use ps, df, ss, docker and journalctl for checks. After a command has run, click Analyze to get a conclusion from its output. The commands and output stay in the terminal so you can check them and review them later. When you're away from your desk, patrol checks and Telegram alerts can tell you about a problem first. Receiving Telegram alerts doesn't require storing credentials in the cloud; only running SSH on the server remotely from Telegram does. See Server alerts in Telegram. For what that looks like on a phone, see Troubleshoot your server from Telegram.

The limits, plainly:

Illustrative example

Illustrative example (not a real customer case): at 2 a.m. the outside curl returns 502. SSH works and load is low. systemctl status nginx shows Nginx running, but ss -ltnp shows nothing listening on the app's port 3000. docker ps -a shows the app container stuck in Restarting, and the first error in docker logs is a failed database write. Back in df -h, the root filesystem is at 100%. The 502 is only the symptom; the disk is the problem. Follow the disk guide to find out whether logs or business data filled it, remove only what's confirmed disposable, then watch whether the container recovers. Rebooting the whole machine in minute one would have left the disk just as full, the outage would have come back a few minutes later, and there would have been less evidence to go on.

AI diagnostic prompt

The site stopped loading about 30 minutes ago. Without changing anything on this server, check load, available memory, disk space and inodes, failed systemd units, listening ports, Docker container status, and system, Nginx and application errors from the last 30 minutes. Separate confirmed facts from hypotheses and list what information is still missing. Do not restart services, delete files or change configuration. If a change is needed, propose the smallest step with its risks, how to verify it and how to roll it back, and wait for my approval.

Try OpsMate

500 free AI calls per month and unlimited servers. The desktop app keeps your SSH credentials on your own machine by default.

Need help interpreting the evidence?

OpsMate helps developers and operators investigate with AI. Review the evidence. After you click Analyze, the command output is sent to cloud AI for analysis; redact sensitive information first.

Start free Desktop with local credentials

Troubleshooting guides