From 1228deaf330f0a70ae6559d24702ee2aba097c15 Mon Sep 17 00:00:00 2001 From: admin9webs Date: Fri, 10 Jul 2026 16:11:23 +0800 Subject: [PATCH] Add devops/oom-diagnosis --- skills/devops/oom-diagnosis/SKILL.md | 118 +++++++++++++++++++++++++++ 1 file changed, 118 insertions(+) create mode 100644 skills/devops/oom-diagnosis/SKILL.md diff --git a/skills/devops/oom-diagnosis/SKILL.md b/skills/devops/oom-diagnosis/SKILL.md new file mode 100644 index 0000000..5b007b8 --- /dev/null +++ b/skills/devops/oom-diagnosis/SKILL.md @@ -0,0 +1,118 @@ +--- +name: oom-diagnosis +description: Diagnose and fix OOM (Out of Memory) kills on Linux servers — memory profiling, identifying memory hogs, adjusting OOM scores, and reducing memory pressure. +category: devops +triggers: + - Process keeps getting killed / disappearing + - User reports OOM, out of memory, or process crashes under load + - free -h shows low available memory or heavy swap usage + - dmesg/journalctl shows oom-killer invocations +--- + +# OOM Diagnosis & Memory Optimization + +Systematic diagnosis and remediation when processes get killed by the Linux OOM killer. + +## Step 1: Assess current state + +```bash +free -h # overall memory picture +ps aux --sort=-%mem | head -20 # top consumers by RSS +``` + +## Step 2: Find OOM killer evidence + +```bash +# Recent OOM events (most reliable sources first) +grep -i "oom-kill\|Killed process\|out of memory" /var/log/kern.log.1 | tail -20 +zgrep -i "oom-kill\|Killed process" /var/log/kern.log.*.gz | tail -20 +sudo dmesg | grep -i "oom\|killed process" +journalctl -k --no-pager | grep -i "oom\|killed" +``` + +The kernel log table with `rss` and `oom_score_adj` columns reveals which process was killed and why. + +## Step 3: Check OOM scores + +```bash +# For each suspect process +cat /proc//oom_score # higher = more likely to be killed +cat /proc//oom_score_adj # adjustment (0=default, 200+ means target) +``` + +A process with `oom_score_adj` ≥ 200 is deliberately marked as a preferred victim. Common sources: +- Chromium/Playwright self-adjusts to 300 (wants to be killed to avoid data loss) +- Some daemons set it explicitly +- Hermes gateway may set it for self-protection in very-low-memory environments + +## Step 4: Identify common memory hogs + +On a typical 2GB VPS running Hermes + web + automation: + +| Process | Typical RSS | Fix | +|---------|------------|-----| +| warp-svc (Cloudflare) | 200-300MB | `sudo systemctl stop warp-svc && sudo systemctl disable warp-svc` if not needed | +| Apache prefork workers | 40-60MB each | Reduce workers in mpm_prefork.conf | +| Chromium/Playwright leftovers | 80-150MB | Kill stale instances after use | +| Hermes LSP (TypeScript) | 350-650MB | `hermes config set lsp.enabled false` + kill the spawned node/tsserver processes | +| Hermes gateway | 120-240MB | Accept; it's load-bearing | +| Hermes session | 150-200MB+ | Accept; normal per-session cost | + +## Step 5: Fix Apache workers (Docker) + +When Apache runs inside a Docker container: + +```bash +# Find and read current config +sudo docker exec cat /etc/apache2/mods-enabled/mpm_prefork.conf + +# Adjust: StartServers→2, MinSpareServers→2, MaxSpareServers→4, MaxRequestWorkers→10 +sudo docker exec bash -c "sed -i \ + -e 's/StartServers.*[0-9]/StartServers 2/' \ + -e 's/MinSpareServers.*[0-9]/MinSpareServers 2/' \ + -e 's/MaxSpareServers.*[0-9]/MaxSpareServers 4/' \ + -e 's/MaxRequestWorkers.*[0-9]/MaxRequestWorkers 10/' \ + /etc/apache2/mods-enabled/mpm_prefork.conf" + +# Reload +sudo docker exec apache2ctl graceful +``` + +## Step 6: Adjust OOM score (last resort) + +```bash +# Remove a process from the OOM victim list +echo 0 | sudo tee /proc//oom_score_adj + +# Protect a critical process entirely +echo -1000 | sudo tee /proc//oom_score_adj +``` + +WARNING: Lowering OOM score just shifts the burden. If memory runs out, a *different* process dies. Only do this combined with actual memory reduction. + +## Step 7: Verify + +```bash +free -h +ps aux --sort=-%mem | head -15 +``` + +## Pitfalls + +- `dmesg` requires root on many systems; fall back to kernel logs +- OOM events may be rotated out quickly on busy systems +- Killing a process doesn't free its swapped pages immediately +- Apache in Docker has no `systemctl` — use `docker exec apache2ctl graceful` +- `oom_score_adj=200` on Hermes gateway is sometimes intentional self-protection +- **CRITICAL: oom_score_adj inherits to child processes.** If the gateway has adj=200, every spawned child (LSP, tsserver, typingsInstaller) inherits it. This can mark 600MB+ of TypeScript tooling as preferred OOM victims alongside the gateway itself. Check `for pid in $(pgrep -P ); do cat /proc/$pid/oom_score_adj; done` — if children inherit 200, lower the gateway's adj first: `echo 0 | sudo tee /proc//oom_score_adj`, then kill and restart the children. + +## Post-fix + +After resolving the immediate OOM, check if the machine needs an upgrade. 2GB is marginal for Hermes + browser automation + web server + database. 4GB is the practical minimum for a multi-service setup. + +### Quick wins (biggest impact first) + +1. **Disable LSP if unused:** `hermes config set lsp.enabled false` then kill spawned tsserver processes. Can free 350-650MB instantly. +2. **Lower gateway oom_score_adj:** `echo 0 | sudo tee /proc//oom_score_adj` — prevents children inheriting the 200 target score. +3. **Reduce Apache workers** (see Step 5). +4. **Stop warp-svc** if Cloudflare WARP is not needed.