Updated

Exit code 137 and OOMKilled: was your container out of memory, or killed by something else?

docker ps -a shows Exited (137) or Restarting (137), the app's logs stop mid-sentence with no error, and the server is a small 1–2 GB VPS running several services. It looks like "out of memory", and often it is, but exit code 137 only says the process received SIGKILL. This guide is for developers who run their own Docker hosts and want to find out who sent that SIGKILL: the kernel's OOM killer, Docker itself, or something else on the host. It goes deeper on 137 than the table in Docker exit codes explained. If the container is stuck in a restart loop for other reasons, start with Docker container keeps restarting.

Every command here is read-only. Run them only on a server you are authorized to access, and replace YOUR_CONTAINER with the real container name or ID. The kernel log checks need sudo.

What exit code 137 means, and why the logs are silent

137 is 128 + 9: the main process was ended by signal 9, SIGKILL. A process cannot catch or delay SIGKILL, so it gets no chance to log a shutdown message. Logs that simply stop, with no error, fit a 137 and do not clear the app.

Several senders produce the same 137:

Who sent SIGKILLWhere the evidence usually isDocker's OOMKilled
Kernel OOM killer, because the container's memory limit was reachedKernel log: Memory cgroup out of memory: Killed process …; an oom event in docker eventsUsually true
Kernel OOM killer, because the whole host ran out of memoryKernel log: Out of memory: Killed process …May be true or false, depending on the cgroup version and the Docker/containerd versions (step 4)
A userspace OOM daemon such as systemd-oomd, or earlyoom (which may try SIGTERM first)That daemon's own journal lines; nothing from the kernelfalse
docker kill or docker rm -fA kill event with signal=9 in docker eventsfalse
docker stop, docker compose down or a redeploy, after the app ignored SIGTERM for the whole grace period (10 s by default)kill with signal=15, then signal=9 about 10 s laterfalse
A script, agent or person on the host running kill -9 on the processOnly a die event in Docker; maybe the tool's own logsfalse

So memory is the first suspect, not the answer. The steps below narrow it down.

1. Capture the exit state before a restart overwrites it

docker ps -a --format 'table {{.Names}}\t{{.Status}}\t{{.Image}}'
docker inspect --format 'status={{.State.Status}} exit={{.State.ExitCode}} oom={{.State.OOMKilled}} started={{.State.StartedAt}} finished={{.State.FinishedAt}} restarts={{.RestartCount}} policy={{.HostConfig.RestartPolicy.Name}}' YOUR_CONTAINER

What this does not tell you: who sent the kill. OOMKilled is a record Docker keeps, not a full account of what happened on the host.

2. Ask Docker whether it sent the signal

docker events --since 24h --until "$(date +%s)" --filter container=YOUR_CONTAINER --filter event=oom --filter event=kill --filter event=die --filter event=start
docker inspect --format 'stop_signal={{.Config.StopSignal}} stop_timeout={{.Config.StopTimeout}}' YOUR_CONTAINER

Repeated event= filters are combined with OR, so this lists only those four event types. Common patterns:

docker events only returns the recent events the daemon keeps in memory (the last 256), and they are gone after the daemon restarts. Run it soon after the incident.

For the stop pattern: an empty stop_signal and stop_timeout=<nil> mean the defaults (SIGTERM, 10 seconds). If PID 1 is a shell wrapper that doesn't pass SIGTERM on to your app, or an app without a SIGTERM handler, every stop ends in SIGKILL. The entrypoint check in the exit codes guide shows what PID 1 is.

3. Check the memory limit and who is using memory

docker inspect --format 'memory_limit_bytes={{.HostConfig.Memory}} memory_swap_bytes={{.HostConfig.MemorySwap}}' YOUR_CONTAINER
docker stats --no-stream --format 'table {{.Name}}\t{{.MemUsage}}\t{{.MemPerc}}'
free -h
ps -eo pid,user,rss,comm --sort=-rss | head -n 10

In a host-wide OOM, the kernel picks its victim mainly by memory size and oom_score_adj, so the process that got killed is not always the one that caused the pressure. A container can be killed because a neighbour grew.

These are snapshots. They show the present, not the moment of the kill, and a stopped container has no stats. To tell a one-off spike from steady growth, take a few snapshots after the container starts again:

for i in 1 2 3 4 5 6 7 8 9 10; do date '+%H:%M:%S'; docker stats --no-stream --format '{{.Name}} {{.MemUsage}}' YOUR_CONTAINER; sleep 60; done

A value that climbs steadily and never levels off points to a leak or an unbounded cache or queue. A value that jumps during specific requests or jobs points to a peak workload. Ten minutes is only a hint; a real trend needs monitoring over hours or days.

4. Read the kernel log, and know when OOMKilled can be false

sudo dmesg -T | grep -i -E 'oom-kill|out of memory|killed process' | tail -n 20
sudo journalctl -k --since "2026-09-29 14:00" --until "2026-09-29 14:30" --no-pager | grep -i -E 'oom-kill|out of memory|killed process'
docker inspect --format '{{.Id}}' YOUR_CONTAINER

No kernel line at a matching time? Check for a userspace OOM daemon, which sends SIGKILL without the kernel logging an OOM:

systemctl is-active systemd-oomd earlyoom
sudo journalctl -u systemd-oomd -u earlyoom --since "2026-09-29 14:00" --no-pager

active means that daemon is running on this host.

When memory was the cause but OOMKilled is false. How Docker learns about an OOM kill depends on the kernel, the cgroup version and the containerd and Docker versions, so treat the following as things to check, not rules:

To see which cgroup version the host uses:

docker info --format 'cgroup={{.CgroupVersion}} driver={{.CgroupDriver}}'

For a container that is still running, its own cgroup counters show kills since it last started (they usually reset when it restarts). On cgroup v2, if the image includes cat:

docker exec YOUR_CONTAINER cat /sys/fs/cgroup/memory.events

A nonzero oom_kill means the kernel killed a process in this container since it started. On cgroup v1 the files have different names; memory.oom_control under /sys/fs/cgroup/memory/ has an oom_kill counter on most current kernels.

Kubernetes aside

On Kubernetes, kubectl describe pod POD_NAME shows Last State: Terminated, Reason: OOMKilled, Exit Code: 137 when a container went over its memory limit. Evicted is a different thing: the kubelet removed the pod because the node was under pressure. A failed liveness probe or a pod deletion that outlasts the termination grace period can also end in 137 without OOMKilled. The limit being enforced is the container's resources.limits.memory.

5. Choose a next step: options you decide on

Match the option to the evidence. Each one is a change for you to plan, with a way back.

Before changing anything in production, write down the evidence, the change, how you'll check it worked, and how you'll roll it back.

Doing this in OpsMate

OpsMate puts an SSH terminal and an AI assistant on the same server page. You can type the docker inspect, docker events and docker stats commands above yourself, and the kernel log checks need sudo, so running those yourself is the straightforward path. You can also describe the problem in plain language, for example "find out why YOUR_CONTAINER keeps exiting with 137 and summarize its last log lines before each restart today". The AI proposes troubleshooting commands, and beyond logs it can use docker, journalctl, ps, df and ss for checks. After a command has run, click Analyze to get a conclusion; the commands and output stay in the terminal so you can check it against the raw lines. When you click Analyze, the command output is sent to cloud AI for analysis; what the desktop app keeps on your own machine by default is your SSH credentials. If the logs contain customer data or secrets, run the commands yourself and share only a redacted excerpt.

Boundaries

Illustrative example

Illustrative example (not a real customer case): a 2 GB VPS runs api (Node.js), postgres and nginx with Compose. Users say the API drops out for a few seconds several times a day. docker ps -a shows api as Up 12 minutes, and inspect prints status=running exit=0 oom=false restarts=5, which looks clean only because the last start reset both values. docker events --since 24h shows five die events with exitCode=137, each followed by start, with no kill event before them, so neither docker kill nor docker stop sent the signal. memory_limit_bytes=0. One die is at 06:05:41 UTC, 14:05:41 on this host (UTC+8). sudo journalctl -k at that minute shows an oom-kill: line with global_oom and a task_memcg path containing the api container's ID, followed by Out of memory: Killed process 2314 (node) … anon-rss:1398212kB. Ten docker stats snapshots a minute apart after the restart show api rising by about 15 MB per minute while postgres stays flat. Together: the host ran out of memory, the kernel chose the Node process in api, and api grows steadily after every start, which fits a leak or an unbounded cache better than a one-off peak. Not proven: which code path grows. The next steps are decisions for a person: give api a memory limit so the pressure stays inside it, set the Node heap below that limit so the failure is logged, and look for the growth in the app, starting from the last deploy.

AI diagnostic prompt

A container on this server exited with code 137 or keeps restarting. Without changing anything, get its State.Status, ExitCode, OOMKilled, FinishedAt, RestartCount and memory limit, list its Docker events from the last 24 hours (oom, kill, die, start), and summarize its last log lines before each exit, with timestamps. Tell me which sender of the SIGKILL the evidence supports and what it does not rule out. Kernel log checks need sudo: list the exact commands for me to run rather than running them. Separate confirmed facts from hypotheses and list what is missing. Do not restart, remove or update containers, and do not change memory limits, swap, restart policies or stop settings. If a change is needed, propose the smallest step with its risks, how to verify it and how to roll it back, and wait for my approval.

Next checks

Try OpsMate

500 free AI calls per month and unlimited servers. The desktop app keeps your SSH credentials on your own machine by default.

Need help interpreting the evidence?

OpsMate helps developers and operators investigate with AI. Review the evidence. After you click Analyze, the command output is sent to cloud AI for analysis; redact sensitive information first.

Start free Desktop with local credentials

Troubleshooting guides