Phase 3c — Re-adoption and cutover: the first phase that reaches production
31 AUG AT 10:17 AM

Phase 3c — Re-adoption and cutover: the first phase that reaches production

0 LOVES 4 VIEWS
A supervisor restart that neither orphans nor kills a bridge — and a decision from the previous phase that had to be reversed to make adoption real rather than a fiction.

Neither orphaned nor killed

A supervisor that kills its bridges when it restarts is not a supervisor anyone will run. Re-adoption had to be solved on day one — and it needed no new state file, because the previous phase already recorded a pid on every successful start, precisely so this one would have it.

A pid alone is not an identity. Pid reuse is real, and a naive pid check eventually signals something innocent. So adoption compares the kernel’s process start time, read at spawn and stored beside the pid. The wall-clock timestamp the supervisor already wrote could not be used: it differs from the kernel value by milliseconds, so equality would never match.

On boot, for each bridge with a recorded pid: if the pid is dead, clear the record. If the pid is alive but the start time disagrees, refuse to adopt and refuse to start, loudly. That second refusal is the important one — starting a replacement would put two processes on one chat channel, and a channel with two bridges answers everything twice. If both match, adopt it, with no output capture, because the pipe belonged to a supervisor that no longer exists.

Reading the start time is platform-specific, and the unsupported case is deliberately harsh: on a platform it cannot read, the supervisor refuses to start at all — not “adopt anyway”, not “spawn anyway”. An unverifiable identity must never be allowed to become a duplicate bridge.

The proof, and what it cost. A bridge was started under the new supervisor, the supervisor was interrupted — and the bridge survived, where the previous phase would have killed it. Restarting adopted the same pid, and the gate self-test count did not move: a supervisor bounce costs nothing, where a restart-everything approach would have spent real money per bridge. A tampered start-time token was then blocked and nothing was spawned; restoring it allowed adoption again, which proves the refusal was a guard rather than a dead end. Throughout, the two live production bridges were untouched.

The reversal that made adoption real

The previous phase gave the child an explicit pipe, and that was the right call there — it fixed a production hang where stopping a bridge mid-turn could block forever. This phase had to take it back, and the reason is worth keeping.

Measured with a standalone experiment: when the parent exits, the read end of that pipe dies, and a write to stdout or stderr raises SIGPIPE — killing the child. The previous phase could tolerate that, because it stopped its children on shutdown anyway. This phase cannot, because “leave them running” is the entire point. A pipe would have made adoption a fiction: the bridge would die moments before it could be re-adopted, and the logs would show a supervisor confidently adopting a corpse.

So the child now gets the log file directly. That keeps the earlier constraint intact — a real file means os/exec starts no copier goroutine, so waiting still cannot hang on a surviving descendant — and costs only that tailing reads back off disk. Which an adopted bridge needs regardless, since the current supervisor never saw its output go past in the first place.

⚠ One more thing measurement caught: a bridge started by hand through make sits inside make’s process group rather than leading its own, so a group signal aimed at it would hit make’s group instead. Supervisor-spawned bridges are always group leaders — but stopping by pid now checks rather than assumes.

What this phase shipped is the mechanism, not the migration. The cutover itself waited on two operator actions and followed the next day. The part that stands alone is the one that matters: a supervisor restart no longer orphans a bridge, and no longer kills one either.

Pushing Tin — managing a bridge fleet from inside the product
Pushing Tin — managing a bridge fleet from inside the product
Aug 29, 2026 Pushing Tin
← Back to Pushing Tin