Updated
Exit code 137 and OOMKilled: was your container out of memory, or killed by something else?
docker ps -a shows Exited (137) or Restarting (137), the app's logs stop mid-sentence with no error, and the server is a small 1–2 GB VPS running several services. It looks like "out of memory", and often it is, but exit code 137 only says the process received SIGKILL. This guide is for developers who run their own Docker hosts and want to find out who sent that SIGKILL: the kernel's OOM killer, Docker itself, or something else on the host. It goes deeper on 137 than the table in Docker exit codes explained. If the container is stuck in a restart loop for other reasons, start with Docker container keeps restarting.
Every command here is read-only. Run them only on a server you are authorized to access, and replace YOUR_CONTAINER with the real container name or ID. The kernel log checks need sudo.
What exit code 137 means, and why the logs are silent
137 is 128 + 9: the main process was ended by signal 9, SIGKILL. A process cannot catch or delay SIGKILL, so it gets no chance to log a shutdown message. Logs that simply stop, with no error, fit a 137 and do not clear the app.
Several senders produce the same 137:
| Who sent SIGKILL | Where the evidence usually is | Docker's OOMKilled |
|---|---|---|
| Kernel OOM killer, because the container's memory limit was reached | Kernel log: Memory cgroup out of memory: Killed process …; an oom event in docker events | Usually true |
| Kernel OOM killer, because the whole host ran out of memory | Kernel log: Out of memory: Killed process … | May be true or false, depending on the cgroup version and the Docker/containerd versions (step 4) |
| A userspace OOM daemon such as systemd-oomd, or earlyoom (which may try SIGTERM first) | That daemon's own journal lines; nothing from the kernel | false |
docker kill or docker rm -f | A kill event with signal=9 in docker events | false |
docker stop, docker compose down or a redeploy, after the app ignored SIGTERM for the whole grace period (10 s by default) | kill with signal=15, then signal=9 about 10 s later | false |
A script, agent or person on the host running kill -9 on the process | Only a die event in Docker; maybe the tool's own logs | false |
So memory is the first suspect, not the answer. The steps below narrow it down.
1. Capture the exit state before a restart overwrites it
docker ps -a --format 'table {{.Names}}\t{{.Status}}\t{{.Image}}'
docker inspect --format 'status={{.State.Status}} exit={{.State.ExitCode}} oom={{.State.OOMKilled}} started={{.State.StartedAt}} finished={{.State.FinishedAt}} restarts={{.RestartCount}} policy={{.HostConfig.RestartPolicy.Name}}' YOUR_CONTAINER
When the container starts again, Docker resets
ExitCodeto 0 andOOMKilledtofalse. With a restart policy, you'll often inspect a container that is already running again (status=running exit=0 oom=false restarts=5) and looks clean. The previous exit is then only visible indocker events(step 2) and the kernel log (step 4).Restarting (137)indocker psonly shows while Docker waits between restarts.status=exited exit=137 oom=true: the lead hypothesis is the container's own memory limit. Still confirm it with step 3 and step 4.status=exited exit=137 oom=false: memory is not ruled out (see step 4), but check Docker's own events first.oom=trueon a container that is still running, or with an exit code other than 137: in recent Docker versions the flag is set when the kernel kills any process in the container, not only the main one. A killed worker or child process may leave the container running, or the main process may exit afterwards with its own code, such as 1.Memory problems don't always end in 137. When a runtime hits its own heap limit first, it usually logs an error and exits another way: Node.js prints
JavaScript heap out of memoryand often exits with 134 (SIGABRT); a JVM logsjava.lang.OutOfMemoryError.finishedis in UTC (it ends inZ). Convert it before comparing with host logs in local time.
What this does not tell you: who sent the kill. OOMKilled is a record Docker keeps, not a full account of what happened on the host.
2. Ask Docker whether it sent the signal
docker events --since 24h --until "$(date +%s)" --filter container=YOUR_CONTAINER --filter event=oom --filter event=kill --filter event=die --filter event=start
docker inspect --format 'stop_signal={{.Config.StopSignal}} stop_timeout={{.Config.StopTimeout}}' YOUR_CONTAINER
Repeated event= filters are combined with OR, so this lists only those four event types. Common patterns:
oom, thendiewithexitCode=137: the kernel killed a process under the container's memory limit and the main process died.oomwith nodieafter it: a process inside was killed, but the main process kept running. Look in the app logs for a worker that died around that time.killwithsignal=15, about 10 s laterkillwithsignal=9, thendiewithexitCode=137: a stop request that ran out of grace period. This is not memory.killwithsignal=9, thendie: someone or something randocker killordocker rm -f.diewithexitCode=137and nokillfrom Docker before it: the signal came from outside Docker, most often the kernel. Go to step 4.A
startright after eachdieis the restart policy at work. Newer Docker versions also addexecDuration(seconds the process ran) todie; if the exits are memory kills, similar values each time suggest a process that grows until it's killed.
docker events only returns the recent events the daemon keeps in memory (the last 256), and they are gone after the daemon restarts. Run it soon after the incident.
For the stop pattern: an empty stop_signal and stop_timeout=<nil> mean the defaults (SIGTERM, 10 seconds). If PID 1 is a shell wrapper that doesn't pass SIGTERM on to your app, or an app without a SIGTERM handler, every stop ends in SIGKILL. The entrypoint check in the exit codes guide shows what PID 1 is.
3. Check the memory limit and who is using memory
docker inspect --format 'memory_limit_bytes={{.HostConfig.Memory}} memory_swap_bytes={{.HostConfig.MemorySwap}}' YOUR_CONTAINER
docker stats --no-stream --format 'table {{.Name}}\t{{.MemUsage}}\t{{.MemPerc}}'
free -h
ps -eo pid,user,rss,comm --sort=-rss | head -n 10
memory_limit_bytes=0: no limit was set, so the container can use host memory until the whole host runs short. A host-wide OOM kill is then the more likely path.memory_swap_bytesis memory plus swap in total. The same value as the limit means no swap for this container;-1means unlimited swap.docker statscovers running containers.MemUsageshows usage and the limit; without a limit, the limit column shows the host's total memory. Add up the big consumers and compare them with the host's RAM: several containers without limits on a 2 GB host compete for the same memory.In
free -h, read theavailablecolumn rather thanfree, and check whether there is any swap at all.pslists the largest processes on the host by resident memory, including processes inside containers.
In a host-wide OOM, the kernel picks its victim mainly by memory size and oom_score_adj, so the process that got killed is not always the one that caused the pressure. A container can be killed because a neighbour grew.
These are snapshots. They show the present, not the moment of the kill, and a stopped container has no stats. To tell a one-off spike from steady growth, take a few snapshots after the container starts again:
for i in 1 2 3 4 5 6 7 8 9 10; do date '+%H:%M:%S'; docker stats --no-stream --format '{{.Name}} {{.MemUsage}}' YOUR_CONTAINER; sleep 60; done
A value that climbs steadily and never levels off points to a leak or an unbounded cache or queue. A value that jumps during specific requests or jobs points to a peak workload. Ten minutes is only a hint; a real trend needs monitoring over hours or days.
4. Read the kernel log, and know when OOMKilled can be false
sudo dmesg -T | grep -i -E 'oom-kill|out of memory|killed process' | tail -n 20
sudo journalctl -k --since "2026-09-29 14:00" --until "2026-09-29 14:30" --no-pager | grep -i -E 'oom-kill|out of memory|killed process'
docker inspect --format '{{.Id}}' YOUR_CONTAINER
Use a window around
finishedfrom step 1 (converted to local time) or adietime from step 2.journalctlrecords when each line arrived, which is usually more reliable than the converted time indmesg -T.dmesgonly holds the current boot, and older lines scroll out. If the host rebooted,sudo journalctl -k -b -1shows the previous boot, but only when the journal is persistent.Memory cgroup out of memory: Killed process 2314 (node)means a cgroup limit was reached, normally the container's own limit.Out of memory: Killed process …without "Memory cgroup" means the whole host ran out.Newer kernels print an
oom-kill:line just before it:constraint=CONSTRAINT_MEMCGfor a cgroup limit,CONSTRAINT_NONEtogether withglobal_oomfor host-wide.task_memcg=shows the killed process's cgroup, which includes the container ID on most setups (for example/system.slice/docker-<id>.scopeor/docker/<id>). Search for the first 12 characters of the ID from the third command.The PID in the kernel line is the host PID, not the PID inside the container.
anon-rssis roughly the memory the process held when it was killed.
No kernel line at a matching time? Check for a userspace OOM daemon, which sends SIGKILL without the kernel logging an OOM:
systemctl is-active systemd-oomd earlyoom
sudo journalctl -u systemd-oomd -u earlyoom --since "2026-09-29 14:00" --no-pager
active means that daemon is running on this host.
When memory was the cause but OOMKilled is false. How Docker learns about an OOM kill depends on the kernel, the cgroup version and the containerd and Docker versions, so treat the following as things to check, not rules:
The container started again before you inspected it (step 1).
A userspace OOM daemon did the killing, not the kernel.
The whole host ran out of memory. On cgroup v1, the event Docker listens for is tied to the container's own limit, so a host-wide kill often leaves the flag
false. On cgroup v2, the kernel counts OOM kills of any kind per cgroup, so the flag is more likely to be set, but older containerd and Docker versions had gaps.On Docker Desktop, memory runs out inside Docker's VM, so tools on the Mac or Windows host won't show it.
To see which cgroup version the host uses:
docker info --format 'cgroup={{.CgroupVersion}} driver={{.CgroupDriver}}'
For a container that is still running, its own cgroup counters show kills since it last started (they usually reset when it restarts). On cgroup v2, if the image includes cat:
docker exec YOUR_CONTAINER cat /sys/fs/cgroup/memory.events
A nonzero oom_kill means the kernel killed a process in this container since it started. On cgroup v1 the files have different names; memory.oom_control under /sys/fs/cgroup/memory/ has an oom_kill counter on most current kernels.
Kubernetes aside
On Kubernetes, kubectl describe pod POD_NAME shows Last State: Terminated, Reason: OOMKilled, Exit Code: 137 when a container went over its memory limit. Evicted is a different thing: the kubelet removed the pod because the node was under pressure. A failed liveness probe or a pod deletion that outlasts the termination grace period can also end in 137 without OOMKilled. The limit being enforced is the container's resources.limits.memory.
5. Choose a next step: options you decide on
Match the option to the evidence. Each one is a change for you to plan, with a way back.
The container's limit was hit and the workload really needs more. Raise the limit (
mem_limitordeploy.resources.limits.memoryin Compose,--memoryondocker run), after checking that the host has room: the sum of limits plus the OS and anything outside Docker should fit in RAM.docker update --memorychanges a running container, but the next recreate from Compose goes back to what the Compose file says.No limit, and the host ran out. Setting limits keeps one container's growth inside that container, so it gets killed instead of the database next to it. If the services together need more memory than the host has, the realistic options are a larger instance or fewer services on this box.
Memory climbs steadily after every start. That looks like a leak or an unbounded cache or queue, and a higher limit only delays the kill. Compare with the last deploy, and use the runtime's own tools (heap snapshots, memory profilers) to find what grows. Recycling workers after a number of requests, where your server supports it, is a stopgap; it doesn't remove the cause.
Runtime heap settings. Keep the runtime's own limit below the container limit so a failure shows up as a clear error in the logs rather than a silent 137. JVM: current JDKs (10 and later, 8u191 and later) read the container limit, and by default typically cap the heap at a quarter of it;
-XX:MaxRAMPercentageor-Xmxsets it explicitly. Leave headroom, because metaspace, thread stacks and direct buffers live outside the heap. Node.js:--max-old-space-size=<MB>(for example throughNODE_OPTIONS); depending on the version, the default may not follow the container limit. Multi-process servers (several web workers, PHP-FPM children) multiply per-worker memory, so the worker count matters as much as a single process.Swap, with caveats. Host swap can absorb short peaks, but under sustained pressure it makes everything slow instead of failing fast, and it hides a leak for longer. Whether a container may use swap depends on
--memory-swap, and some kernels don't support swap limits for containers (docker infothen prints a warning). Test on a quiet host first.It was a stop, not memory. Make the app handle SIGTERM, use the exec form of
CMD/ENTRYPOINTorexecin entrypoint scripts so the app is PID 1, or add a minimal init (--init,init: truein Compose). Only raise the grace period (docker stop -t,stop_grace_period) if a clean shutdown genuinely takes longer.Something else on the host sent it. Find the script, agent or person with the timestamp from step 2. Memory settings won't change this one.
Before changing anything in production, write down the evidence, the change, how you'll check it worked, and how you'll roll it back.
Doing this in OpsMate
OpsMate puts an SSH terminal and an AI assistant on the same server page. You can type the docker inspect, docker events and docker stats commands above yourself, and the kernel log checks need sudo, so running those yourself is the straightforward path. You can also describe the problem in plain language, for example "find out why YOUR_CONTAINER keeps exiting with 137 and summarize its last log lines before each restart today". The AI proposes troubleshooting commands, and beyond logs it can use docker, journalctl, ps, df and ss for checks. After a command has run, click Analyze to get a conclusion; the commands and output stay in the terminal so you can check it against the raw lines. When you click Analyze, the command output is sent to cloud AI for analysis; what the desktop app keeps on your own machine by default is your SSH credentials. If the logs contain customer data or secrets, run the commands yourself and share only a redacted excerpt.
Boundaries
137 is a clue, and
OOMKilledis evidence, not proof. Base a conclusion on the exit state, Docker events and kernel log lines with matching timestamps.A snapshot can't show a trend. Deciding between a leak and an undersized limit needs memory history over time, not one
docker statsreading.The commands in this guide don't change your containers. The AI is mainly for troubleshooting: dangerous commands are blocked. Memory limits, heap flags, swap, restart policies and stop settings are changes you decide on. For what OpsMate does and doesn't do on its own, see the FAQ.
An AI summary is a starting point. Check it against the inspect output, the events and the kernel log.
Under heavy memory pressure, SSH itself can hang. If you can't SSH into the host, OpsMate can't reach it either; use your cloud provider's console.
Illustrative example
Illustrative example (not a real customer case): a 2 GB VPS runs api (Node.js), postgres and nginx with Compose. Users say the API drops out for a few seconds several times a day. docker ps -a shows api as Up 12 minutes, and inspect prints status=running exit=0 oom=false restarts=5, which looks clean only because the last start reset both values. docker events --since 24h shows five die events with exitCode=137, each followed by start, with no kill event before them, so neither docker kill nor docker stop sent the signal. memory_limit_bytes=0. One die is at 06:05:41 UTC, 14:05:41 on this host (UTC+8). sudo journalctl -k at that minute shows an oom-kill: line with global_oom and a task_memcg path containing the api container's ID, followed by Out of memory: Killed process 2314 (node) … anon-rss:1398212kB. Ten docker stats snapshots a minute apart after the restart show api rising by about 15 MB per minute while postgres stays flat. Together: the host ran out of memory, the kernel chose the Node process in api, and api grows steadily after every start, which fits a leak or an unbounded cache better than a one-off peak. Not proven: which code path grows. The next steps are decisions for a person: give api a memory limit so the pressure stays inside it, set the Node heap below that limit so the failure is logged, and look for the growth in the app, starting from the last deploy.
AI diagnostic prompt
A container on this server exited with code 137 or keeps restarting. Without changing anything, get its State.Status, ExitCode, OOMKilled, FinishedAt, RestartCount and memory limit, list its Docker events from the last 24 hours (oom, kill, die, start), and summarize its last log lines before each exit, with timestamps. Tell me which sender of the SIGKILL the evidence supports and what it does not rule out. Kernel log checks need sudo: list the exact commands for me to run rather than running them. Separate confirmed facts from hypotheses and list what is missing. Do not restart, remove or update containers, and do not change memory limits, swap, restart policies or stop settings. If a change is needed, propose the smallest step with its risks, how to verify it and how to roll it back, and wait for my approval.
Next checks
Need the other exit codes (1, 125–127, 139, 143)? Docker exit codes explained.
Restart loop with a different cause? Docker container keeps restarting covers restart policies and log windows.
What memory evidence looks like in practice: an anonymized case where a server reached 88% memory, and what that evidence could not prove.
Want memory history instead of snapshots? Small-team VPS monitoring checklist.
Try OpsMate
500 free AI calls per month and unlimited servers. The desktop app keeps your SSH credentials on your own machine by default.
Need help interpreting the evidence?
OpsMate helps developers and operators investigate with AI. Review the evidence. After you click Analyze, the command output is sent to cloud AI for analysis; redact sensitive information first.