Phase 1 — Liveness: the bridge says it is alive
29 AUG AT 11:55 AM

Phase 1 — Liveness: the bridge says it is alive

0 LOVES 1 VIEWS
Answer "is the bridge up?" from the database, without ps. The one phase that stands alone — and the only one whose real defects could not have been caught by any unit test.

Is the bridge up?

Before: two bridges ran as anonymous OS processes. The only way to know whether one was alive was ps aux | grep bobit on the host, and the only way to know whether it was working was to watch the chat. If one died overnight, nothing anywhere recorded that it had.

After: each bridge is a row in the agent registry with a last_heartbeat that advances every 60 seconds, and a dead one is flipped to offline by the orchestrator within 180 seconds, on its own. Studio’s existing agents list renders it with no template change at all — the read path really was as type-agnostic as the audit claimed, which is the kind of prediction worth checking rather than assuming.

Liveness was moved ahead of identity deliberately: it is the one phase that delivers something before a supervisor exists. It adds no supervisor, starts and stops nothing, builds no UI, records no cost, and moves no flag into config — those are Phases 2 to 5. It writes rows a later phase will read, plus one that a running system reads immediately.

It needed three rows, not two. The plan said a bridge needs an Agent row and an API key. Reading the middleware found a third: without an active AgentSiteAssignment, every call returns 403 agent is not assigned to this site. That third row is also why this phase closed the untyped task-assignment paths in the same slice — the assignment is precisely what makes a bridge look assignable to code that never filters on type.

And that was not theoretical. sites.agents_enabled gates the agent UI but not the agent API. Measured: roughindustries has it off, so its bridge was shielded by accident; willartley has it on, so that bridge would have appeared in a live assign-agent dropdown. Guards that work by accident are not guards, and that is why the two exclusions shipped here rather than being deferred as belt-and-braces.

What only a live run found

Three defects surfaced on 2026-08-15 that no unit test would have caught, because each one lives in the gap between the binary and a real server. They are the justification for two flags that would otherwise look like gold-plating.

The Studio host answers the agent API with a 307 redirect to the site host. Go’s HTTP client drops the Authorization header on a cross-host redirect — so a bridge pointed at the Studio host follows the redirect, arrives without its bearer token, and gets a 401.

What makes this the significant one is how it fails: the request looks completely correct at the call site, and the key never appears on the wire at all. Nothing in a unit test reaches it either — an httptest server has one host and issues no redirect.

Measured across four hosts: the Studio host 307s and loses the header; the site host, the kontent host and the loopback prod port all return a clean 401, which is the correct answer for a live route that wants a key. ⛔ A bridge must be pointed at its own site’s host — and that is the strongest argument for -api-base being a full URL. The host becomes something an operator states explicitly and can see, rather than something assembled behind their back from a flag named after a different connection.

connectAndJoin hardcoded ws://, and the dev server is TLS-only. So the binary was structurally untestable against the real dev environment — which is also why every one of its defaults pointed at the dogfood tree. That is a quiet admission worth reading twice: a component only ever exercised against production is not a component anyone can safely change.

The obvious workaround was to open a plain-HTTP port on dev so the existing binary could connect. It was rejected, and the reasoning is the durable part: it would have made the development environment less secure in order to let a test pass. The fix went into the binary instead, as -ws-scheme.

The dev agent-worker died on startup: the production instance already held the wildcard bind for the skill API, so the dev one could not bind and exited. The fix was a one-line port override.

It is recorded here because of what it is rather than what it cost. This is the plan’s own thesis, one level up. The supervisor that Phase 3 goes on to build must allocate ports rather than let two instances race for a hardcoded one — and the thing that monitors the fleet turned out to have exactly the unmanaged-daemon problem the fleet itself has. The monitor was hand-launched too. That observation is what eventually became Phase 14.

Proof, not assertion

Six smoke steps, executed on dev on 2026-08-15, all six passing. Five of them prove reporting. Only the sixth proves the thing the phase actually exists for.

  • Bridge rows seeded — agent 6, type bridge, assigned to its site.
  • Key minted through the existing account-admin UI — no provisioning code was written, because that path already existed.
  • Heartbeat on connect: offline → idle, live last_heartbeat, current_task_id still null.
  • The heartbeat ticker: 17:43:22 → 17:44:22, exactly 60 seconds — which proves the timer rather than just the connect-time call.
  • Renders in the agent registry with a bridge badge and zero template changes, confirming the type-agnostic read path.
  • Absent from the assign-agent dropdown: five agents listed, zero occurrences of the bridge anywhere on the page.

The sixth step is the one that matters. Kill the bridge, wait past the threshold with the orchestrator running, and confirm it flips to offline on its own. Steps one to five prove that a bridge can report; only this proves that the absence of reporting is detected. Without it, the feature is a status that can only ever say “alive”.

The evidence, from the orchestrator’s own heartbeat events — three healthy check-ins, then the flip on the first tick past the threshold, which is exactly the designed behaviour for a 180-second threshold sampled every 60 seconds:

  • 17:46:13 — healthy
  • 17:47:13 — healthy
  • 17:48:13 — healthy (171s stale, correctly still under threshold)
  • 17:49:13 — marked_offline / stale_heartbeat (231s)

⚠ Measure staleness from last_heartbeat, not from when you killed the process. A first pass reported “~140s” by timing from the kill; the real interval was 231s. The wrong number makes a threshold that is behaving perfectly look broken — and a smoke test that reports the wrong number is worse than one that does not run, because it gets believed.

Rollback was free, and that is a property worth designing for. The enum value is additive and unused by any other row. The two exclusions only ever remove bridges from lists that contained none until this phase created them, so they were no-ops against the database as it stood. And the live bridges were unaffected until a key appeared in their environment — removing one line reverts them, with no redeploy.

Pushing Tin — managing a bridge fleet from inside the product
Pushing Tin — managing a bridge fleet from inside the product
Aug 29, 2026 Pushing Tin
← Back to Pushing Tin