Updated

Inherited a server: what to check first when you have no ops hire

You've inherited a server someone else built, and the first job is finding out what's running on it. You have your first paying users. The contractor or former colleague who set up the server has moved on, and the handover was an IP address and a password. You aren't sure where the logs are, when the certificate expires, or who to call when something breaks, and a full-time ops hire isn't in the budget. This guide is for founders, product managers and the developer who "also looks after the server": spend an hour taking inventory of the server you inherited and write it down on one page, set up a 10-minute weekly read-only check, and decide in advance who you'll call when things go wrong.

This is about surviving a server you inherited. If you want to know which metrics to monitor after launch, see the small-team VPS monitoring checklist. Every command below is read-only and works on a typical Linux VPS. If the output makes no sense yet, that's fine: save it as-is and read it alongside the explanations here.

1. Find out who owns what before you touch the server

Small teams rarely get stuck on a mistyped command. They get stuck because nobody can log in to the account that matters. Fill in this table first. Record where things are and who owns them, never the passwords or keys themselves:

ItemWhereWhose loginWho can change itExpiry / renewal
Cloud account (including billing)
Domain registrar and DNS
HTTPS certificate
Code repository and deploy method
Database and where backups live
Third-party API keys (email, payments, SMS…)

2. Take inventory: what is actually running here?

System and resources:

cat /etc/os-release
uname -r
uptime
nproc
free -h
df -h

This gives you the OS version, kernel, uptime, CPU core count, memory and disk. It does not tell you whether those resources are enough; that takes trends over time.

Running services and listening ports:

systemctl list-units --type=service --state=running --no-pager
sudo ss -ltnp

Containers and other process managers:

docker ps --format 'table {{.Names}}\t{{.Image}}\t{{.Status}}\t{{.Ports}}'
docker compose ls
pm2 list

The front door (Nginx):

ls -l /etc/nginx/sites-enabled/ /etc/nginx/conf.d/ 2>/dev/null
sudo nginx -T 2>/dev/null | grep -E '^\s*(server_name|listen|proxy_pass|root)\s'

nginx -T prints the full active configuration. The filter keeps only domain names, ports, proxy targets and static roots, so you can connect "this domain → this port → this service". Check a full config for internal addresses or credentials before sharing it with anyone.

Scheduled jobs and backups:

crontab -l
sudo crontab -l
ls -l /etc/cron.d/
systemctl list-timers --no-pager
sudo grep -rIl -E 'backup|dump|rsync|restic|borg' /etc/cron* /var/spool/cron 2>/dev/null

Certificates:

sudo certbot certificates

From your own computer you can also check the live certificate's expiry date (replace example.com with your domain):

echo | openssl s_client -connect example.com:443 -servername example.com 2>/dev/null | openssl x509 -noout -enddate -issuer

certbot certificates only returns something if the server uses certbot. Your certificate may be managed by a CDN or load balancer instead, in which case check that provider's console.

Who can log in:

awk -F: '$3 >= 1000 && $3 < 65534 {print $1}' /etc/passwd
getent group sudo wheel
sudo find /root /home -name authorized_keys -exec wc -l {} \;

This lists regular users, who has sudo, and roughly how many public keys each authorized_keys file holds (usually one per line). Write down accounts and keys belonging to people who have left; they are the first things to decide about revoking. For the order that takes access back without locking yourself out, see Contractor or ex-employee left? How to revoke their server access.

3. Write it down on one page

One row per service is enough:

ServiceRuns via (systemd / Docker / PM2)PortConfig locationLog locationHow to restartWho knows it best

The cells you can't fill in are where you'll get stuck during the next incident. Fill those first, or ask whoever built the setup.

4. A 10-minute weekly read-only check

uptime
free -h
df -h
df -i
systemctl --failed --no-pager
docker ps -a --format 'table {{.Names}}\t{{.Status}}'

Also glance at when the last backup ran and how long until the certificate expires. For a fuller list of checks and how to route alerts, see the small-team VPS monitoring checklist. For the times you're away from your desk, patrol checks and Telegram alerts can notify you first. Receiving Telegram alerts doesn't require storing credentials in the cloud; only running SSH on the server remotely from Telegram does. See Server alerts in Telegram.

What the numbers in an alert mean

NumberWhat it isWhen it mattersFirst command to run
LoadTasks running or waiting for CPU/disk. The three numbers in uptime are 1-, 5- and 15-minute averagesStaying well above the core count while users notice slowness. Short spikes are usually fineuptime, nproc
Memory %Share of memory in use. Linux uses spare memory for cache, so this number often looks highWhen available stays very low, swap keeps growing, or the kernel log shows processes being killedfree -h
Disk %Used capacity on a filesystemNear 100%, writes fail and databases and logs break first. The growth rate matters more than the number itselfdf -h
Inode %The budget for the *number* of files, separate from capacityAt 100% you can't create files even with free space, typically from huge numbers of small filesdf -i

There is no universal threshold. Compare against what's normal for this machine. If a disk or inode alert fires, follow Linux disk usage suddenly spikes before deleting anything.

5. Your escalation path

  1. Gather evidence first. Run the read-only checks in the order from Site down at 2 a.m. and save the output. Helpers move much faster when you hand them evidence.

  2. Then call someone. Keep contact details ready for whoever built the setup, a contractor or consultant you can pay per incident, and your cloud provider's support portal.

  3. Make temporary access revocable. Give outside help their own account and key, never a shared root password. When the work is done, remove their key, disable the account and rotate any secrets they handled. When someone leaves, see Contractor or ex-employee left? How to revoke their server access.

  4. If you suspect a break-in, don't start deleting things yourself; see Was my Linux server hacked?

Where OpsMate fits

OpsMate puts an SSH terminal and an AI assistant on the same page. For the inventory, you can type each command above into the terminal yourself. If you can't remember them, ask the AI in plain language, for example "check which services are running on this machine, which ports are listening and what scheduled jobs exist, without changing anything". The AI proposes troubleshooting commands and can use ps, df, ss, docker and journalctl for checks. After a command has run, click Analyze to get its reading of the output. The commands and output appear in the terminal for you to read and verify. The inventory page is still yours to write and keep up to date; the AI saves you from memorizing commands and helps you make sense of the output.

The limits:

Illustrative example

Illustrative example (not a real customer case): a founder inherits a VPS with 4 GB of memory. The inventory shows Nginx forwarding two domains to ports 3000 and 3001; docker compose ls shows one Compose project with web, worker and postgres containers; ss -ltnp shows postgres listening on 0.0.0.0:5432; root's crontab runs pg_dump at 3 a.m. every day and writes the file to the same disk, with no copy anywhere else; certbot's systemd timer renews the certificate, which has 40 days left; and authorized_keys holds three keys, two of which nobody can identify. The priorities: a backup with no off-site copy that has never been test-restored, a database port that may be exposed to the internet, and unexplained keys. All three fixes (copying backups elsewhere, tightening the security group or database listen address, revoking keys) are changes to plan: assess the impact, prepare a rollback, then act, with professional help if needed.

AI diagnostic prompt

I've just inherited this server and don't know what runs on it. Without changing anything, check the OS version and resource overview, running services, listening ports and the processes behind them, Docker containers and Compose projects, Nginx domains and proxy targets, cron jobs and systemd timers, likely backup jobs, certificate expiry, and which users and SSH keys exist. Turn this into a draft inventory, marking what is confirmed, what is a guess and what is missing. Do not modify, delete or restart anything. If you find a risk that needs action, explain the impact and a rollback plan first and wait for my approval.

Try OpsMate

500 free AI calls per month and unlimited servers. The desktop app keeps your SSH credentials on your own machine by default.

Need help interpreting the evidence?

OpsMate helps developers and operators investigate with AI. Review the evidence. After you click Analyze, the command output is sent to cloud AI for analysis; redact sensitive information first.

Start free Desktop with local credentials

Troubleshooting guides