Pushing Tin — managing a bridge fleet from inside the product
29 AUG AT 11:17 AM

Pushing Tin — managing a bridge fleet from inside the product

0 LOVES 1 VIEWS
BoBit worked; nothing managed it. Nineteen phases that moved a fleet of chat bridges out of the terminal and into the product — declared, supervised, seen and metered.

What it was, measured

BoBit worked. Nothing managed it. Two bridges were running when this program started — one for Rough Industries, one for Will Artley — both hand-launched, both from a developer’s own build. Everything about them lived in filesystem and flag conventions rather than in the product.

  • Config was CLI flags. Nineteen of them. Nothing was in the database.
  • Lifecycle was make and nohup. make stop-bobit fell back to pkill -f "./bobit" when the pidfile was missing — which would have killed both bridges.
  • One bridge had no make target and no pidfile at all. Its exact flags existed only in the process table; had it died, the launch command was unrecoverable. Written down is not managed.
  • Cost was invisible. Every message spawns claude -p, which spends real money, and none of it was recorded anywhere.
  • Port allocation was manual. A third bridge meant remembering that 8766 and 8767 were taken.

The key finding was that the agent system already fitted, and safely. Agent already carried everything a process manager needs — status, heartbeats, config revisions, cost events, budgets — while the dispatcher’s existing agent_type = worker filter kept a bridge out of task execution by construction. That property is what made the whole approach safe.

The verification pass that followed then narrowed it honestly, which is the more useful half: “by construction” covered dispatch, not assignment. Two paths handed a bridge a task with no type filter at all, and Phase 1 closed the gap. The enum audit that came with it was a list rather than a feeling — 12 agent_type sites, 6 templates, 4 seeds, and 27 Agent.Query() call sites of which 6 lacked a type filter — and every later phase applied that list instead of re-deriving it.

The rulings that shaped it

Decisions that were settled out loud, before the code that depended on them was written.

An agent’s status was answering two different questions at once: is this agent enabled (governance) and should a process be running (desired state). The collision was real, not theoretical — a budget pause would have stopped nothing while making a perfectly healthy bridge report as dead.

Resolved by splitting the field, not by filtering it. A new desired_state: running|stopped sits beside status, and the supervisor runs a bridge only when desired_state = running AND status = active. The operator’s intent now survives an unpause instead of being overwritten by it. A type filter was rejected because it left status conflated and removed the spend brake entirely; a paused_reason column was rejected as one field still doing two jobs.

A bridge’s configuration is operator-editable, and the supervisor turns it into a command line. So the rule was written before any of it was built: nothing in config_json may resolve to an executable.

Measuring it found the rule needed three buckets, not two. A value may be forbidden because it is (a) an executable, (b) a path the supervisor reads or writes, or (c) a destination that receives a credential. That third one is the subtle one: a URL is never an executable, but -api-base is exactly where the bridge sends its API key, so an editable one redirects the credential. Of the 19 flags, 6 are operator-editable and 13 are not.

The same rule is why the launchd plist is hand-written and static. A ProgramArguments entry is an executable path, so generating bridge plists from config would have re-opened the exact hole this closed — same technology, opposite exposure.

The adversarial pass found that cmd/bobit imports no ent client and opens no database. Two planned phases had quietly assumed a capability the bridge does not have.

The resolution split responsibility, and is better than database access would have been. The supervisor reads config_revisions and renders CLI flags; the bridge stays flag-driven and unchanged. There is no config-read endpoint on the agent API, so configuration cannot flow from the database into a bridge — it flows database → supervisor → flags. Reporting goes the other way over the existing agent API, authenticated with an agent key.

The unplanned dividend: multi-machine works for free. Because a bridge reports over HTTP rather than by writing to a database, one on another host reports fine. What stays genuinely single-host is supervision — a supervisor can only fork processes on its own machine — so the open question became one supervisor per host, not whether reporting works at all.

Adoption-in-place was the tempting option: leave the running bridges alone and have the supervisor simply take them over. It was rejected, and the reasoning is the useful part.

Adopting in place would permanently manage processes whose real flags could never be checked against the configuration they supposedly came from. It skips the proof. So instead both live bridges were stopped and brought back up from their declarations, with the operator present, at a chosen moment. Seconds of downtime buys the only proof that the declarations are correct and complete — a mis-bucketed flag surfaces immediately rather than at the next unplanned restart, months later, with nobody watching.

The cutover then cost nothing anyway: stopping the supervisor left both bridges running, and the new one adopted them — 2 adopted, 0 blocked, pids unchanged, no restart, and none of the $0.12 gate self-tests a restart would have spent.

A bridge account needs a three-link access chain: site membership on the bridge’s site, a channel ACL entry, and access to the post the channel hangs off. All three refusals carry the identical access_denied envelope, with no id and no channel id to tell them apart — so fixing one and retrying looks exactly like the fix not working.

⚠ And the reason the retries failed is sharper than “the errors look alike”. Measured against the chat service while drafting the access-truth phase: only post visibility refuses a join at all. The join gate consults post access and explicitly discards the channel-ACL result; membership and the ACL govern the bridge’s reply, not its arrival. The production rows date it exactly — membership granted at 08:21, channel ACL at 08:22, post access at 08:36, and the join succeeded at 08:38. Only the 08:36 grant was ever going to work; the first two were never consulted by the thing that was failing.

It had been invisible for months because the old shared bot was a global admin, which bypasses every check. Least-privilege identities are what surfaced it — the chain always existed, and nobody wrote it down because the first account through it made every link a no-op.

The operator’s ruling was not to automate the grants. A machine acquiring reach into private content must stay a deliberate human act; automating it would let a bridge silently gain access to a private chat and a private post as a side effect of submitting a form. The fix is legibility at the touch points instead. One consequence worth stating: each grant is a licence to wake another agent’s worker, which makes it a spend decision as well as an access one.

A run order of blocks was layered on top of the phases, so every piece of work had two names at once — block 5 and phase 13 — and then slices were subdivided in chat prompts into a third layer that existed in no document at all.

Every status error since lived at a seam between those layers. One report silently dropped a slice and omitted three whole phases, because it was assembled from memory of conversations rather than read from the plan. Two of those phases had been filed as “a separate track”, which is precisely how they vanished: a phase either holds a position in the table or it does not exist.

The rule, and it is the reason this cluster looks the way it does: if work needs subdividing, the phase doc changes first, then the prompt is written. Status is generated by reading the table and the git log — never from recollection. One vocabulary: phases, with slices where a phase doc defines them, and nothing below that.

The board

Nineteen phases. The numbers are creation order, not run order — nothing said that out loud until Phase 8 was drafted, so it is said here. Each one has its own post in this cluster.

  • 1 · Liveness — the bridge says it is alive, and Studio can answer “is it up?” without a ps. Shipped.
  • 2 · Identity and config — a bridge is declared, not remembered. Shipped.
  • 3a · Render and reconcile — the supervisor says what it would do. Shipped.
  • 3b · Spawn and supervise — it actually forks. Shipped.
  • 3c · Re-adoption and cutover — a supervisor restart neither orphans nor kills what is running. Shipped.
  • 4 · The UI — one writer, one viewer, and a surface that cannot lie. Shipped (4a–4f).
  • 5 · Agent identity — one user per bridge, on real evaluated access with no admin bypass anywhere in the fleet. Shipped.
  • 6 · Cost — the meter that makes the spend brake able to fire at all. Shipped.
  • 7 · Sleeping bridges — parked. Sketched, never scheduled; the economics were redone and the conclusion changed.
  • 8 · Access truth — a bridge that cannot hear must not read as healthy. Shipped.
  • 9 · Attached machines — made visible where humans work. Shipped.
  • 10 · The audit backlog — 32 findings across 14 slices. Shipped.
  • 11 · Archiving a bridge — completely, including what only a bridge has. Shipped.
  • 12 · The deploy boundary — production stopped running its binaries out of the development tree. Shipped.
  • 13 · Bridges are agents — and the UI never joined them. Shipped.
  • 14 · The orchestrator — which had served production for fifteen days from a closed session’s scratch directory. Shipped.
  • 15 · An agent speaks as itself — attribution, and the termination rule that goes with it. Shipped.
  • 16 · A watcher that refuses — parked by operator call: “if it isn’t broken I don’t want to mess with it right now.”
  • 17 · What the audit filed — four slices shipped, one measured and dropped because its remedy was the defect. Shipped.

Parked is not blocked, and not unfinished. Phases 7 and 16 hold positions here on purpose, and so does launchd — whose gate was released, whose plist is correct, and whose remaining work is one launchctl load the operator has chosen not to run. Two phases once vanished from every status report by being filed as a separate track; parking by deletion is how that happens, so nothing is parked by deletion here.

← Back to Pushing Tin