diff --git a/docs/superpowers/specs/2026-08-07-xy-readiness-gate-design.md b/docs/superpowers/specs/2026-08-07-xy-readiness-gate-design.md new file mode 100644 index 0000000..9c37dd7 --- /dev/null +++ b/docs/superpowers/specs/2026-08-07-xy-readiness-gate-design.md @@ -0,0 +1,218 @@ +# xy — Readiness Gate (`wait-for`) + +**Date:** 2026-08-07 +**Status:** Approved — ready for implementation planning + +## Problem + +Since start-on-login landed, `xy daemon` runs from a launchd agent at login. +Docker-backed servers (`gitea`, `signoz`) race OrbStack and lose. + +Measured on this machine 2026-08-07: OrbStack's process started at 12:47:56 +and created `~/.orbstack/run/docker.sock` at **12:48:14** — 18 seconds later. +OrbStack is a BTM login item, not a LaunchAgent, so it is not even guaranteed +to begin before the agent does. + +Against that, `docker run` fails instantly while the socket is absent, and the +default restart budget is `backoff-initial 1s`, `backoff-max 30s`, +`max-retries-per-minute 5`. Attempts therefore land at t=0, 1, 3, 7, 15s; the +fifth trips the cap and `policy::decide` returns `MarkFailed`, which is +terminal. **A docker-backed server gives up ~15 seconds after the daemon +starts** and stays down until a manual `xy start`. + +The 2026-08-07 log shows this nearly firing during a reload: socket error at +10:47:58, recovery at 10:48:15 — 17 seconds, right at the edge. + +Widening the retry budget per-server would paper over it. The real defect is +that a server with an unmet *infrastructure dependency* consumes its +crash-retry budget while waiting for something that was never its own fault. + +## Goals + +- Let a server declare a precondition that must hold before it is spawned. +- Waiting must never consume the crash-retry budget. +- Waiting must be visible and self-explaining, not indistinguishable from a hang. +- Waiting must stay responsive to `stop` / `shutdown`. + +## Non-goals (deferred) + +- Ordering between supervised servers (`depends-on `). The conditions + here can express it (`tcp 127.0.0.1:3928`) without a dependency graph. +- Liveness probes on an already-running server. This gate is start-time only; + HTTP health probes remain a non-goal from the MVP spec. +- Re-checking the condition while a server runs. + +## Decisions + +### Three condition kinds + +```kdl +wait-for { + path "~/.orbstack/run/docker.sock" + timeout "120s" + interval "1s" +} +``` + +`path`, `tcp`, and `command` are all supported: + +| Kind | Ready when | +|---|---| +| `path "

"` | the path exists | +| `tcp ""` | a TCP connection succeeds | +| `command "" args "…"` | the process exits 0 | + +Exactly one condition per `wait-for` block. Zero or more than one is a parse +error, so `xy reload` exits 3 and changes nothing — matching how invalid +config already behaves. + +`path` values get `~` expanded at parse time via `etcetera::home_dir()`, since +the motivating case is literally `~/.orbstack/run/docker.sock`. The expanded +path is what `xy status` displays, so there is no ambiguity about what was +checked. + +The `command` kind is bounded by the poll `interval`: each invocation is run +under a timeout of `interval` and a non-exit is treated as not-ready, so a hung +`docker info` cannot wedge the poll loop. + +### Timeout marks the server Failed without spawning + +On timeout, the server goes to `Failed` and nothing is spawned. `xy status` +names the condition that timed out. + +Rejected: spawning anyway on timeout, which reintroduces exactly the +retry-burn this feature removes; and waiting forever, which turns a typo in a +path into a server that sits in `waiting` with nothing ever surfacing. + +### Waiting never touches the retry window + +`RetryWindow::record` continues to be called only for real process exits. This +is the property that fixes the reported bug: a cold Docker costs patience, not +retry budget. + +### Every start is gated + +The gate runs for the initial start, the explicit `restart` command, and the +automatic post-crash respawn. If Docker dies and takes the container with it, +the respawn should wait for Docker again rather than burn retries. + +### New `Waiting` state, reason on `status` only + +`ServerState` gains `Waiting` — an additive enum variant, backward compatible +on the wire. `xy list` shows a plain `waiting` to keep the table narrow; +`xy status ` renders what it is blocked on and for how long: + + $ xy status gitea + state: waiting + wait-for: path /Users/olsson/.orbstack/run/docker.sock + 43s elapsed, timeout 120s + +Rejected: reusing `Starting`, which is indistinguishable from a slow-starting +process — the exact ambiguity this feature exists to remove. + +Elapsed time is computed at **read time** from a carried `Instant`, never +precomputed at publish time. This is the rule established by the 2026-08-07 +uptime bug, where a value snapshotted into the watch channel froze at 0. + +## Schema + +`xy-protocol/src/config.rs`: + +```rust +pub enum WaitCondition { + Path(PathBuf), + Tcp(String), + Command { command: PathBuf, args: Vec }, +} + +pub struct WaitForConfig { + pub condition: WaitCondition, + pub timeout: Duration, // default 120s + pub interval: Duration, // default 1s +} +``` + +`ServerConfig` gains `#[serde(default)] pub wait_for: Option`. +Absent `wait-for` means no gate and behaviour identical to today. + +## Supervisor + +The gate lives inside `do_start`, before the spawn. This is the one structural +change worth calling out: `do_start` currently returns +`std::io::Result`, which cannot express the gate's outcomes. It +becomes: + +```rust +enum StartFailure { + Spawn(std::io::Error), + WaitTimedOut, + Cancelled, // Stop arrived while waiting + Shutdown, // Shutdown arrived while waiting +} + +async fn do_start(&mut self, cause: StartCause) -> Result +``` + +All three call sites are updated. `Shutdown` must propagate a `return` from +`run`, exactly as the backoff sleep already does. `Cancelled` settles the +server in `Stopped` — a `stop` issued during a wait leaves it stopped, not +failed, since the user asked for it. `WaitTimedOut` settles in `Failed`. + +**The poll loop is interruptible.** It selects over the interval sleep and +`self.cmd_rx`, the same shape as the existing backoff sleep (commit `b366df0`, +"make backoff sleep interruptible by Stop/Shutdown"). Without this, `xy stop` +on a waiting server would block for up to the full timeout. + +`Status` carries the wait progress rather than a rendered string: + +```rust +pub struct WaitProgress { + pub description: String, + pub started_at: Instant, + pub timeout: Duration, +} +``` + +with `Status.waiting: Option`. Handlers derive `elapsed_secs` +when building the RPC response. `StatusDetail` gains +`wait_for: Option`. + +## Testing + +TDD throughout. + +- Config parse tests per condition kind; defaults applied; `~` expanded; zero + conditions and two conditions both rejected. +- Condition checks in isolation: path present/absent; TCP against a listener + bound to an ephemeral port, and against a closed port; command exiting 0, + exiting non-zero, and one that hangs past `interval`. +- Supervisor: condition already met spawns immediately; timeout yields `Failed` + with no spawn **and an untouched retry window** (the regression assertion for + the reported bug); `Stop` during a wait cancels promptly rather than after + the timeout; `Shutdown` during a wait returns. +- Formatting: `waiting` in the list table; `xy status` renders the description + and a *growing* elapsed value. +- Integration: a server whose `wait-for` path can never exist reaches `Failed` + after a short configured timeout without ever spawning. + +Timeouts and intervals in tests are milliseconds, so no test waits on wall time. + +## Documentation + +`README.md` gains a `wait-for` section with the three condition kinds and the +OrbStack example, plus a note that waiting does not consume the restart budget. + +## Rollout + +Once shipped, `gitea.kdl` and `signoz.kdl` each gain: + +```kdl +wait-for { + path "~/.orbstack/run/docker.sock" + timeout "180s" +} +``` + +Until then those two servers remain vulnerable at every login. The interim +mitigation, if wanted, is `backoff-max "15s"` and `max-retries-per-minute 20` +in both files.