Updated
Inherited a server: what to check first when you have no ops hire
You've inherited a server someone else built, and the first job is finding out what's running on it. You have your first paying users. The contractor or former colleague who set up the server has moved on, and the handover was an IP address and a password. You aren't sure where the logs are, when the certificate expires, or who to call when something breaks, and a full-time ops hire isn't in the budget. This guide is for founders, product managers and the developer who "also looks after the server": spend an hour taking inventory of the server you inherited and write it down on one page, set up a 10-minute weekly read-only check, and decide in advance who you'll call when things go wrong.
This is about surviving a server you inherited. If you want to know which metrics to monitor after launch, see the small-team VPS monitoring checklist. Every command below is read-only and works on a typical Linux VPS. If the output makes no sense yet, that's fine: save it as-is and read it alongside the explanations here.
1. Find out who owns what before you touch the server
Small teams rarely get stuck on a mistyped command. They get stuck because nobody can log in to the account that matters. Fill in this table first. Record where things are and who owns them, never the passwords or keys themselves:
| Item | Where | Whose login | Who can change it | Expiry / renewal |
|---|---|---|---|---|
| Cloud account (including billing) | ||||
| Domain registrar and DNS | ||||
| HTTPS certificate | ||||
| Code repository and deploy method | ||||
| Database and where backups live | ||||
| Third-party API keys (email, payments, SMS…) |
The cloud and domain accounts should belong to the company, not to someone who has left. Turn on two-factor authentication.
Stop sharing a single root password. Give everyone who needs access their own account or key, so you can revoke one person without locking out everyone.
2. Take inventory: what is actually running here?
System and resources:
cat /etc/os-release
uname -r
uptime
nproc
free -h
df -h
This gives you the OS version, kernel, uptime, CPU core count, memory and disk. It does not tell you whether those resources are enough; that takes trends over time.
Running services and listening ports:
systemctl list-units --type=service --state=running --no-pager
sudo ss -ltnp
Most of the list is the operating system's own services. Look for names you recognize:
nginx,docker,mysql/mariadb,postgresql,redis, and anything named after your product.ss -ltnpshows which process listens on which port. An address of0.0.0.0or[::]means every network interface; unless your cloud security group blocks it, the internet can reach that port.127.0.0.1means local only. A database port open to the world is usually a risk to get someone to look at soon.
Containers and other process managers:
docker ps --format 'table {{.Names}}\t{{.Image}}\t{{.Status}}\t{{.Ports}}'
docker compose ls
pm2 list
docker compose lslists running Compose projects and the path to each config file, which often answers "how was this app deployed?" Add-ato include stopped projects.If a Node.js app is managed by PM2, run
pm2 listas the user who started it. If the command doesn't exist, PM2 isn't installed; move on.
The front door (Nginx):
ls -l /etc/nginx/sites-enabled/ /etc/nginx/conf.d/ 2>/dev/null
sudo nginx -T 2>/dev/null | grep -E '^\s*(server_name|listen|proxy_pass|root)\s'
nginx -T prints the full active configuration. The filter keeps only domain names, ports, proxy targets and static roots, so you can connect "this domain → this port → this service". Check a full config for internal addresses or credentials before sharing it with anyone.
Scheduled jobs and backups:
crontab -l
sudo crontab -l
ls -l /etc/cron.d/
systemctl list-timers --no-pager
sudo grep -rIl -E 'backup|dump|rsync|restic|borg' /etc/cron* /var/spool/cron 2>/dev/null
This is where backup scripts, certificate renewals and log cleanup usually live. The last
grepis a rough keyword search; some matches will be system jobs rather than your backups, so open each one.Once you find a backup job, confirm three things: where the backup files actually end up, when the last one ran, and whether anyone has ever restored one on another machine. A backup on the same disk as the database disappears with it if the disk fails or the server is deleted.
Certificates:
sudo certbot certificates
From your own computer you can also check the live certificate's expiry date (replace example.com with your domain):
echo | openssl s_client -connect example.com:443 -servername example.com 2>/dev/null | openssl x509 -noout -enddate -issuer
certbot certificates only returns something if the server uses certbot. Your certificate may be managed by a CDN or load balancer instead, in which case check that provider's console.
Who can log in:
awk -F: '$3 >= 1000 && $3 < 65534 {print $1}' /etc/passwd
getent group sudo wheel
sudo find /root /home -name authorized_keys -exec wc -l {} \;
This lists regular users, who has sudo, and roughly how many public keys each authorized_keys file holds (usually one per line). Write down accounts and keys belonging to people who have left; they are the first things to decide about revoking. For the order that takes access back without locking yourself out, see Contractor or ex-employee left? How to revoke their server access.
3. Write it down on one page
One row per service is enough:
| Service | Runs via (systemd / Docker / PM2) | Port | Config location | Log location | How to restart | Who knows it best |
|---|
The cells you can't fill in are where you'll get stuck during the next incident. Fill those first, or ask whoever built the setup.
4. A 10-minute weekly read-only check
uptime
free -h
df -h
df -i
systemctl --failed --no-pager
docker ps -a --format 'table {{.Names}}\t{{.Status}}'
Also glance at when the last backup ran and how long until the certificate expires. For a fuller list of checks and how to route alerts, see the small-team VPS monitoring checklist. For the times you're away from your desk, patrol checks and Telegram alerts can notify you first. Receiving Telegram alerts doesn't require storing credentials in the cloud; only running SSH on the server remotely from Telegram does. See Server alerts in Telegram.
What the numbers in an alert mean
| Number | What it is | When it matters | First command to run |
|---|---|---|---|
| Load | Tasks running or waiting for CPU/disk. The three numbers in uptime are 1-, 5- and 15-minute averages | Staying well above the core count while users notice slowness. Short spikes are usually fine | uptime, nproc |
| Memory % | Share of memory in use. Linux uses spare memory for cache, so this number often looks high | When available stays very low, swap keeps growing, or the kernel log shows processes being killed | free -h |
| Disk % | Used capacity on a filesystem | Near 100%, writes fail and databases and logs break first. The growth rate matters more than the number itself | df -h |
| Inode % | The budget for the *number* of files, separate from capacity | At 100% you can't create files even with free space, typically from huge numbers of small files | df -i |
There is no universal threshold. Compare against what's normal for this machine. If a disk or inode alert fires, follow Linux disk usage suddenly spikes before deleting anything.
5. Your escalation path
Gather evidence first. Run the read-only checks in the order from Site down at 2 a.m. and save the output. Helpers move much faster when you hand them evidence.
Then call someone. Keep contact details ready for whoever built the setup, a contractor or consultant you can pay per incident, and your cloud provider's support portal.
Make temporary access revocable. Give outside help their own account and key, never a shared root password. When the work is done, remove their key, disable the account and rotate any secrets they handled. When someone leaves, see Contractor or ex-employee left? How to revoke their server access.
If you suspect a break-in, don't start deleting things yourself; see Was my Linux server hacked?
Where OpsMate fits
OpsMate puts an SSH terminal and an AI assistant on the same page. For the inventory, you can type each command above into the terminal yourself. If you can't remember them, ask the AI in plain language, for example "check which services are running on this machine, which ports are listening and what scheduled jobs exist, without changing anything". The AI proposes troubleshooting commands and can use ps, df, ss, docker and journalctl for checks. After a command has run, click Analyze to get its reading of the output. The commands and output appear in the terminal for you to read and verify. The inventory page is still yours to write and keep up to date; the AI saves you from memorizing commands and helps you make sense of the output.
The limits:
It doesn't replace a backup strategy, security hardening or architecture decisions. Complex changes still need someone qualified.
The AI is mainly for troubleshooting: dangerous commands are blocked. Low-risk fixes such as restarting a service or rotating logs run automatically by default. Deletions and config edits are your decision. For what OpsMate does and doesn't do on its own, see the FAQ.
The desktop app keeps SSH credentials on your own machine by default, but that doesn't mean all data stays there: when you click
Analyze, the command output is sent to cloud AI for analysis. If the output contains secrets or customer data, redact it first and give the AI the redacted excerpt.
Illustrative example
Illustrative example (not a real customer case): a founder inherits a VPS with 4 GB of memory. The inventory shows Nginx forwarding two domains to ports 3000 and 3001; docker compose ls shows one Compose project with web, worker and postgres containers; ss -ltnp shows postgres listening on 0.0.0.0:5432; root's crontab runs pg_dump at 3 a.m. every day and writes the file to the same disk, with no copy anywhere else; certbot's systemd timer renews the certificate, which has 40 days left; and authorized_keys holds three keys, two of which nobody can identify. The priorities: a backup with no off-site copy that has never been test-restored, a database port that may be exposed to the internet, and unexplained keys. All three fixes (copying backups elsewhere, tightening the security group or database listen address, revoking keys) are changes to plan: assess the impact, prepare a rollback, then act, with professional help if needed.
AI diagnostic prompt
I've just inherited this server and don't know what runs on it. Without changing anything, check the OS version and resource overview, running services, listening ports and the processes behind them, Docker containers and Compose projects, Nginx domains and proxy targets, cron jobs and systemd timers, likely backup jobs, certificate expiry, and which users and SSH keys exist. Turn this into a draft inventory, marking what is confirmed, what is a guess and what is missing. Do not modify, delete or restart anything. If you find a risk that needs action, explain the impact and a rollback plan first and wait for my approval.
Try OpsMate
500 free AI calls per month and unlimited servers. The desktop app keeps your SSH credentials on your own machine by default.
Need help interpreting the evidence?
OpsMate helps developers and operators investigate with AI. Review the evidence. After you click Analyze, the command output is sent to cloud AI for analysis; redact sensitive information first.