Skip to content
Back to skills

Opencode Bash Background Process Survival

ASecurity

Diagnose and fix backgrounded processes (dev servers, nohup'd jobs) that get KILLED when the bash tool's command terminates, completes, or times out — even when started with nohup & disown. Use when: a backgrounded server died by the next turn or when the launching command timed out; the dev server started last turn returns connection refused / curl 000; 'shell tool terminated command after exceeding timeout' appears and the PID is gone; you need a process to survive across multiple bash-tool...

  • 39 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 30, 2026
toolspythongoshellbashnextjsgit

Security analysis

A100/100

Scanned September 30, 2026

npx -y skills add ericmjl/skills --skill opencode-bash-background-process-survival --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Opencode Bash Background Process Survival?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Opencode Bash Background Process Survival
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/ericmjl-opencode-bash-background-process-survival/badge)](https://www.skillsdirectory.com/skills/ericmjl-opencode-bash-background-process-survival)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: opencode-bash-background-process-survival
description: >-
  Diagnose and fix backgrounded processes (dev servers, nohup'd jobs) that get KILLED when the bash tool's command terminates, completes, or times out — even when started with nohup & disown. Use when: a backgrounded server died by the next turn or when the launching command timed out; the dev server started last turn returns connection refused / curl 000; 'shell tool terminated command after exceeding timeout' appears and the PID is gone; you need a process to survive across multiple bash-tool turns on macOS; or about to add more disown/nohup flags hoping to make a job survive (insufficient alone). ROOT CAUSE: disown detaches from the job table but NOT the process group; the bash tool signals the whole group on termination. FIX: the subshell double-fork '( nohup cmd >log 2>&1 & )' — the subshell exits immediately, orphaning the child to launchd (PID 1); verify with 'ps -p <pid>' next turn. Lookalike distinctions live in the body.
created_by: autolearn
created_at: "2026-08-04"
---

# Opencode Bash Background Process Survival

Diagnose and fix backgrounded processes (dev servers, nohup'd jobs, long-running commands) that get KILLED when the opencode bash tool's command terminates, completes, or times out — even though they were started with 'nohup ... & disown'. Use when: you backgrounded a server/process with 'nohup cmd >log 2>&1 & disown' and it died by the next turn or when the launching command timed out; the dev server you started last turn returns connection refused / curl 000; 'shell tool terminated command after exceeding timeout' appears and the backgrounded PID is gone; you need a process to survive across multiple bash-tool turns on macOS; or you are about to add more 'disown' / 'nohup' flags hoping to make a backgrounded job survive (disown alone is insufficient). ROOT CAUSE: 'disown' detaches a job from the shell's JOB TABLE but does NOT move it to a new process group/session; when the bash tool terminates its command (on normal completion OR timeout), SIGTERM/SIGKILL is delivered to the entire PROCESS GROUP, and the disowned child shares that group and dies with it. macOS does NOT ship 'setsid' (no easy 'setsid cmd &' like Linux). FIX (robust macOS pattern): the subshell double-fork — '( nohup cmd >log 2>&1 & )' — the parentheses spawn a subshell that backgrounds the child then EXITS immediately, orphaning the child which gets reparented to launchd (PID 1); combined with nohup (ignores SIGHUP) and redirected stdio, the child survives the parent shell's death AND the bash tool's process-group signal. Verify with 'ps -p <pid>' on the next turn. Distinct from Python stdout block-buffering (memory: visibility into a surviving process, not survival itself), from pixi-run process-group signaling (memory #832: 'setsid' workaround for signaling a pixi job, not surviving bash-tool termination), and from the marimo/Lektor preview-server lifecycle rules (those say WHEN to background; THIS says HOW to background so it actually survives). Generalizes to any long-running process backgrounded from the opencode bash tool on macOS (Lektor preview, marimo edit, vLLM/Ollama serve, python http.server, uvicorn, dev servers).

## The failure, step by step

The opencode bash tool runs each command in a shell. When that command finishes (or the tool kills it on timeout), the tool terminates the process **group**, not just the lead process:

1. You run `nohup my_server > /tmp/srv.log 2>&1 & disown` and capture PID `$!`.
2. `disown` removes the job from the shell's **job table** (so the shell won't send it SIGHUP on exit, and `jobs` won't list it).
3. **But `disown` does NOT create a new process group or session.** The backgrounded child still shares the shell's process group (PGID).
4. The bash tool terminates the command → SIGTERM/SIGKILL is delivered to the **whole process group** → your "detached" child dies with it.
5. Next turn: `curl localhost:PORT` returns `000` (connection refused), `ps -p <pid>` shows nothing. The `disown` gave a false sense of safety.

The string `shell tool terminated command after exceeding timeout` in the tool output is the smoking gun: the timeout kill took the process group, and your server with it.

## The fix: subshell double-fork

macOS does **not** ship `setsid` (the Linux one-liner `setsid cmd &`). The macOS-equivalent robust pattern is the **subshell double-fork**:

```bash
mkdir -p /tmp/srv-logs
( nohup my_server > /tmp/srv-logs/srv.log 2>&1 & echo $! > /tmp/srv-logs/srv.pid )
```

Why this works:
- The `( ... )` runs the body in a **subshell**.
- Inside, `nohup cmd &` backgrounds the child in the subshell, and the subshell writes the PID to a file then **exits immediately**.
- When the subshell exits, the child is **orphaned** and **reparented to launchd (PID 1)**.
- `nohup` makes it ignore SIGHUP; the redirected stdio detach it from the tool's pipe.
- Because the child is now a child of launchd in its own reparented state, it is **outside** the bash tool's process group and survives the tool's group-wide signal.

This is the pattern that makes a server survive across many turns. (`setsid` would also work if installed; `brew install util-linux` provides it — but the subshell pattern needs no extra install and is the portable macOS default.)

## Verification (do this every time, on the NEXT turn)

```bash
PID=$(cat /tmp/srv-logs/srv.pid); ps -p "$PID" && echo "ALIVE" || echo "DEAD"
curl -sS -o /dev/null -w "%{http_code}\n" --max-time 5 http://localhost:PORT
```

If `ps -p` reports DEAD or curl returns `000`, the detachment failed — re-launch with the subshell pattern.

## Common variants

**Lektor preview server (website repo, two worktrees, two ports):**
```bash
mkdir -p /tmp/lektor-dev
( cd /Users/ericmjl/github/website/worktree-a && nohup /Users/ericmjl/github/website/.pixi/envs/default/bin/lektor server -p 5001 -h 0.0.0.0 > /tmp/lektor-dev/a.log 2>&1 & echo $! > /tmp/lektor-dev/a.pid )
```
Note `/tmp/` is cleared on reboot — always `mkdir -p` the log dir first, or the redirect silently fails with "no such file". (See memory #1291.)

**marimo edit server:**
```bash
( nohup uvx marimo edit --sandbox --no-token /path/to/nb > /tmp/marimo.log 2>&1 & echo $! > /tmp/marimo.pid )
```

**python http.server (static expose over Tailscale):**
```bash
( nohup python3 -m http.server 8080 > /tmp/http.log 2>&1 & echo $! > /tmp/http.pid )
```

## Gotchas that compound this

- **Python stdout is block-buffered** when redirected to a file: even a *surviving* process can look dead because its log stays empty for minutes. Use `python -u` or `PYTHONUNBUFFERED=1` so you can actually see it working. This is a VISIBILITY problem, distinct from the SURVIVAL problem this skill fixes. (Memory #775.)
- **pgrep/pkill self-match**: when checking liveness, `pgrep -af my_server` matches its own command line and returns a phantom PID every call. Capture `$!` at launch (as above) and check with `ps -p <PID>` instead. (Memory #781.)
- **macOS port 5000 is occupied** by ControlCenter (AirPlay Receiver) — use 5001/5002/etc. for dev servers. (Memory #1197/#1291.)
- **The bash tool's default 120s timeout** is separate from this: even a correctly-detached process whose *launch command itself* runs >120s will be killed mid-launch. Detach fast (the subshell returns immediately) and verify on the next turn.
- **DO NOT add `disown` inside the subshell** — `( nohup cmd >log 2>&1 & disown )` is a trap. The `&` already detached the job in the subshell's context, so `disown` has no job-table entry to act on and errors (`disown: not a job control shell` / `no such job`). Worse, that error makes the subshell return NON-ZERO, which aborts any `&&` chain that follows the launch (e.g. `( ... & disown ) && sleep 4 && cat log` never runs the sleep/cat because the subshell exited non-zero) AND can prevent the server from starting cleanly at all. The correct pattern is `( nohup cmd >log 2>&1 & )` with NOTHING after the `&` — the subshell exits 0 immediately, orphaning the child to launchd. The `disown` is not merely redundant here, it is actively harmful. (Discovered 2026-08-08 brain42 vault-browser vite dev server launch: `disown` errored, the dev server never came up, and the verification `cat log` was skipped because the `&&` chain aborted on the subshell's non-zero exit.)

## When NOT to use this

- **Interactive creative tools** (the website video editor at localhost:4096, a GUI app the user will drive): per memory #62, hand the user the launch command instead of backgrounding — the user owns the lifecycle of an interactive tool, the agent owns the lifecycle of a preview/serve process.
- **A command whose output you need inline** (a one-shot build, a test run): don't background at all; let it complete in the foreground and raise the bash-tool `timeout` if needed.
- **You need the process to die when the session ends**: then a plain `&` (no detach) is correct — the survival pattern is for processes you want to PERSIST.

## Related

- Memory #832 — pixi `run dev` puts itself + grandchildren in one process group; `setsid` isolates a backgrounded pixi job for *signaling*. Sibling technique, different goal (signal a job vs. survive bash-tool termination).
- Memory #1197 / #1291 — Lektor preview-server lifecycle: WHEN to background + dual-port dual-worktree setup. THIS skill supplies the HOW that makes the backgrounding actually survive.
- Memory #775 — Python stdout block-buffering (visibility into a surviving process).
- Memory #781 — pgrep/pkill self-match when checking liveness.

## After backgrounding a dev server: verify which port YOUR instance bound

- A curl 200 on the default port (3000) may come from a PRE-EXISTING server — often the MAIN checkout's dev server when you started yours from a git worktree, serving DIFFERENT code. Next.js auto-increments ports (3000→3001) and logs the actual port in the server's own log. PROCEDURE: after backgrounding the server, read THE LOG YOU REDIRECTED (not a curl probe of :3000) for the Ready/port line, then run browser QA / curls against THAT port. Never treat a successful curl on the default port as evidence your instance is serving — you would QA the wrong app (stale code from another branch). Discovered 2026-08-16 (learn-anything portal-gantt worktree: main checkout's server held :3000; worktree instance bound :3001).

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…