Phase 3b — Spawn and supervise: the supervisor actually forks
31 AUG AT 09:48 AM

Phase 3b — Spawn and supervise: the supervisor actually forks

0 LOVES 4 VIEWS
Forking is the easy half. This phase is about what happens when a bridge will not start — and about a production hang that only a mutation could find.

What happens when a bridge will not start

Phase 3a rendered a command line and refused to run it. This phase forks — and the interesting part is not the forking, it is what happens when a bridge will not start.

Every retry is a real API call. A bridge refuses to start when its gate self-test fails, and that self-test invokes the model, so a restart loop is a billing event rather than just noise. But give-up alone would turn a twenty-minute upstream outage into an overnight bridge outage. The failure space splits cleanly: transient failures — upstream unreachable, server restarting, a network blip — resolve unaided; permanent ones — a wrong binary path, a bound port, bad config — never do.

So the policy is a generous threshold rather than a twitchy one (roughly five attempts over minutes, not three in thirty seconds), the failure counter resets when the supervisor restarts so a bounce clears every stuck bridge at once, the last error text is stored with the failure state because “the gate never fired” and “address already in use” call for opposite responses, and the retry is a button in the UI rather than a terminal command.

A bound gate port used to cost 28 cents and log a lie. The broker calls ListenAndServe inside a goroutine and reports approval broker listening whether or not the bind succeeded, then proceeds straight into the self-test. So address already in use — the textbook permanent failure — arrives asynchronously, after the money is already committed, while the captured log’s most prominent line says everything is fine. The fatal can even fire mid-self-test, killing the bridge and orphaning its worker.

The fix is small and worth its weight: before every spawn the supervisor binds the port itself and closes it immediately. A bound port becomes a zero-cost, correctly classified, correctly attributed failure instead of a 28-cent race with a misleading log. ⚠ Stated honestly, check-then-spawn leaves a TOCTOU window of milliseconds — this is a cost reduction and a better error message, not a guarantee.

The defect the mutation run found

This is the phase’s strongest argument for mutation testing, and it is worth stating bluntly: review did not catch this, and no test written against the intended behaviour would have. It surfaced only because one mutation — signal the pid instead of the process group — made the whole test binary hang rather than fail.

The bug. Spawning assigned a plain writer to the child’s stdout and stderr. When os/exec is handed anything that is not an *os.File, it creates its own pipe and a copying goroutine, and cmd.Wait() blocks until that goroutine finishes — which requires every write end of the pipe to be closed.

Why that is not hypothetical for a bridge. The bridge starts its worker in its own process group, so a group kill never reaches it, and hands it the supervisor’s pipe as stderr. Stopping a bridge mid-turn therefore leaves the worker holding the write end: Wait never returns, the done channel never closes, and stopping the child hangs forever. A supervisor shutdown would then hang too, because it stops every child it started. That is a production hang, reachable by the ordinary act of stopping a bridge while someone is talking to it.

The fix. An explicit pipe. The write end is passed to the child as a real file, so os/exec starts no goroutine and Wait returns the moment the process exits, regardless of who else holds the pipe. The supervisor runs its own copier from the read end, which ends when the last writer closes — so an orphan’s output is still captured rather than lost.

It is pinned by a test whose grandchild deliberately lands in its own process group, exactly matching the worker’s shape. And the pin was verified the only way that means anything: by reverting to the buggy form and watching what happened. The test reports a named fifteen-second failure rather than hanging the suite — which is the difference between a check that reports and one that merely stops.

Pushing Tin — managing a bridge fleet from inside the product
Pushing Tin — managing a bridge fleet from inside the product
Aug 29, 2026 Pushing Tin
← Back to Pushing Tin