Files
xy/docs/superpowers/specs/2026-08-07-xy-readiness-gate-design.md
T
logaritmiskandClaude Opus 5 73e17d51f3 docs(plan): implementation plan for the wait-for readiness gate
Six TDD tasks: schema + KDL parsing, the Waiting state and WaitInfo, a
ready module that evaluates one condition, the interruptible gate inside
do_start, CLI surfacing, and the README.

Also corrects the spec: xy-protocol has no etcetera dependency, so ~ is
expanded from $HOME; and in KDL `command`/`args` are sibling nodes inside
the wait-for block, which changes how the parser counts conditions.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EGntTHCW3sEPy1VBRopNNp
2026-08-07 15:10:48 +02:00

234 lines
8.5 KiB
Markdown

# xy — Readiness Gate (`wait-for`)
**Date:** 2026-08-07
**Status:** Approved — ready for implementation planning
## Problem
Since start-on-login landed, `xy daemon` runs from a launchd agent at login.
Docker-backed servers (`gitea`, `signoz`) race OrbStack and lose.
Measured on this machine 2026-08-07: OrbStack's process started at 12:47:56
and created `~/.orbstack/run/docker.sock` at **12:48:14** — 18 seconds later.
OrbStack is a BTM login item, not a LaunchAgent, so it is not even guaranteed
to begin before the agent does.
Against that, `docker run` fails instantly while the socket is absent, and the
default restart budget is `backoff-initial 1s`, `backoff-max 30s`,
`max-retries-per-minute 5`. Attempts therefore land at t=0, 1, 3, 7, 15s; the
fifth trips the cap and `policy::decide` returns `MarkFailed`, which is
terminal. **A docker-backed server gives up ~15 seconds after the daemon
starts** and stays down until a manual `xy start`.
The 2026-08-07 log shows this nearly firing during a reload: socket error at
10:47:58, recovery at 10:48:15 — 17 seconds, right at the edge.
Widening the retry budget per-server would paper over it. The real defect is
that a server with an unmet *infrastructure dependency* consumes its
crash-retry budget while waiting for something that was never its own fault.
## Goals
- Let a server declare a precondition that must hold before it is spawned.
- Waiting must never consume the crash-retry budget.
- Waiting must be visible and self-explaining, not indistinguishable from a hang.
- Waiting must stay responsive to `stop` / `shutdown`.
## Non-goals (deferred)
- Ordering between supervised servers (`depends-on <name>`). The conditions
here can express it (`tcp 127.0.0.1:3928`) without a dependency graph.
- Liveness probes on an already-running server. This gate is start-time only;
HTTP health probes remain a non-goal from the MVP spec.
- Re-checking the condition while a server runs.
## Decisions
### Three condition kinds
```kdl
wait-for {
path "~/.orbstack/run/docker.sock"
timeout "120s"
interval "1s"
}
```
`path`, `tcp`, and `command` are all supported:
| Kind | Ready when |
|---|---|
| `path "<p>"` | the path exists |
| `tcp "<host:port>"` | a TCP connection succeeds |
| `command "<c>" args "…"` | the process exits 0 |
Exactly one condition per `wait-for` block. Zero or more than one is a parse
error, so `xy reload` exits 3 and changes nothing — matching how invalid
config already behaves.
`path` values get a leading `~/` expanded at parse time from `$HOME`, since the
motivating case is literally `~/.orbstack/run/docker.sock`. `$HOME` rather than
`etcetera::home_dir()` deliberately: `xy-protocol` is a leaf crate with no
directory dependencies, and adding one for a two-line prefix substitution is
not worth it. If `$HOME` is unset the literal path is kept, which simply never
exists and surfaces as a normal timeout. The expanded path is what `xy status`
displays, so there is no ambiguity about what was checked.
Note the KDL shape of the `command` kind: `command` and `args` are **sibling
nodes** inside the block, mirroring the top level, not one node with trailing
arguments:
wait-for {
command "docker"
args "info"
}
so the parser accumulates them separately and combines them after the loop.
`args` without `command` is a parse error.
The `command` kind is bounded by the poll `interval`: each invocation is run
under a timeout of `interval` and a non-exit is treated as not-ready, so a hung
`docker info` cannot wedge the poll loop.
### Timeout marks the server Failed without spawning
On timeout, the server goes to `Failed` and nothing is spawned. `xy status`
names the condition that timed out.
Rejected: spawning anyway on timeout, which reintroduces exactly the
retry-burn this feature removes; and waiting forever, which turns a typo in a
path into a server that sits in `waiting` with nothing ever surfacing.
### Waiting never touches the retry window
`RetryWindow::record` continues to be called only for real process exits. This
is the property that fixes the reported bug: a cold Docker costs patience, not
retry budget.
### Every start is gated
The gate runs for the initial start, the explicit `restart` command, and the
automatic post-crash respawn. If Docker dies and takes the container with it,
the respawn should wait for Docker again rather than burn retries.
### New `Waiting` state, reason on `status` only
`ServerState` gains `Waiting` — an additive enum variant, backward compatible
on the wire. `xy list` shows a plain `waiting` to keep the table narrow;
`xy status <name>` renders what it is blocked on and for how long:
$ xy status gitea
state: waiting
wait-for: path /Users/olsson/.orbstack/run/docker.sock
43s elapsed, timeout 120s
Rejected: reusing `Starting`, which is indistinguishable from a slow-starting
process — the exact ambiguity this feature exists to remove.
Elapsed time is computed at **read time** from a carried `Instant`, never
precomputed at publish time. This is the rule established by the 2026-08-07
uptime bug, where a value snapshotted into the watch channel froze at 0.
## Schema
`xy-protocol/src/config.rs`:
```rust
pub enum WaitCondition {
Path(PathBuf),
Tcp(String),
Command { command: PathBuf, args: Vec<String> },
}
pub struct WaitForConfig {
pub condition: WaitCondition,
pub timeout: Duration, // default 120s
pub interval: Duration, // default 1s
}
```
`ServerConfig` gains `#[serde(default)] pub wait_for: Option<WaitForConfig>`.
Absent `wait-for` means no gate and behaviour identical to today.
## Supervisor
The gate lives inside `do_start`, before the spawn. This is the one structural
change worth calling out: `do_start` currently returns
`std::io::Result<S::Child>`, which cannot express the gate's outcomes. It
becomes:
```rust
enum StartFailure {
Spawn(std::io::Error),
WaitTimedOut,
Cancelled, // Stop arrived while waiting
Shutdown, // Shutdown arrived while waiting
}
async fn do_start(&mut self, cause: StartCause) -> Result<S::Child, StartFailure>
```
All three call sites are updated. `Shutdown` must propagate a `return` from
`run`, exactly as the backoff sleep already does. `Cancelled` settles the
server in `Stopped` — a `stop` issued during a wait leaves it stopped, not
failed, since the user asked for it. `WaitTimedOut` settles in `Failed`.
**The poll loop is interruptible.** It selects over the interval sleep and
`self.cmd_rx`, the same shape as the existing backoff sleep (commit `b366df0`,
"make backoff sleep interruptible by Stop/Shutdown"). Without this, `xy stop`
on a waiting server would block for up to the full timeout.
`Status` carries the wait progress rather than a rendered string:
```rust
pub struct WaitProgress {
pub description: String,
pub started_at: Instant,
pub timeout: Duration,
}
```
with `Status.waiting: Option<WaitProgress>`. Handlers derive `elapsed_secs`
when building the RPC response. `StatusDetail` gains
`wait_for: Option<WaitInfo { description, elapsed_secs, timeout_secs }>`.
## Testing
TDD throughout.
- Config parse tests per condition kind; defaults applied; `~` expanded; zero
conditions and two conditions both rejected.
- Condition checks in isolation: path present/absent; TCP against a listener
bound to an ephemeral port, and against a closed port; command exiting 0,
exiting non-zero, and one that hangs past `interval`.
- Supervisor: condition already met spawns immediately; timeout yields `Failed`
with no spawn **and an untouched retry window** (the regression assertion for
the reported bug); `Stop` during a wait cancels promptly rather than after
the timeout; `Shutdown` during a wait returns.
- Formatting: `waiting` in the list table; `xy status` renders the description
and a *growing* elapsed value.
- Integration: a server whose `wait-for` path can never exist reaches `Failed`
after a short configured timeout without ever spawning.
Timeouts and intervals in tests are milliseconds, so no test waits on wall time.
## Documentation
`README.md` gains a `wait-for` section with the three condition kinds and the
OrbStack example, plus a note that waiting does not consume the restart budget.
## Rollout
Once shipped, `gitea.kdl` and `signoz.kdl` each gain:
```kdl
wait-for {
path "~/.orbstack/run/docker.sock"
timeout "180s"
}
```
Until then those two servers remain vulnerable at every login. The interim
mitigation, if wanted, is `backoff-max "15s"` and `max-retries-per-minute 20`
in both files.