Phase 3a rendered a command line and refused to run it. This phase forks — and the interesting part is not the forking, it is what happens when a bridge will not start.
Every retry is a real API call. A bridge refuses to start when its gate self-test fails, and that self-test invokes the model, so a restart loop is a billing event rather than just noise. But give-up alone would turn a twenty-minute upstream outage into an overnight bridge outage. The failure space splits cleanly: transient failures — upstream unreachable, server restarting, a network blip — resolve unaided; permanent ones — a wrong binary path, a bound port, bad config — never do.
So the policy is a generous threshold rather than a twitchy one (roughly five attempts over minutes, not three in thirty seconds), the failure counter resets when the supervisor restarts so a bounce clears every stuck bridge at once, the last error text is stored with the failure state because “the gate never fired” and “address already in use” call for opposite responses, and the retry is a button in the UI rather than a terminal command.
A bound gate port used to cost 28 cents and log a lie. The broker calls ListenAndServe inside a goroutine and reports approval broker listening whether or not the bind succeeded, then proceeds straight into the self-test. So address already in use — the textbook permanent failure — arrives asynchronously, after the money is already committed, while the captured log’s most prominent line says everything is fine. The fatal can even fire mid-self-test, killing the bridge and orphaning its worker.
The fix is small and worth its weight: before every spawn the supervisor binds the port itself and closes it immediately. A bound port becomes a zero-cost, correctly classified, correctly attributed failure instead of a 28-cent race with a misleading log. ⚠ Stated honestly, check-then-spawn leaves a TOCTOU window of milliseconds — this is a cost reduction and a better error message, not a guarantee.