Phase 5 — Agent identity: one user per bridge
31 AUG AT 10:52 AM

Phase 5 — Agent identity: one user per bridge

0 LOVES 1 VIEWS
Two bridges sharing one user are structurally deaf to each other. Giving each its own least-privilege identity fixed that — and surfaced an access chain that had been invisible for months.

Why one user per bridge

Four facts, all verified against the running system, and together they made this phase unavoidable.

  • Two bridges sharing one user are structurally deaf to each other. Each bridge skips any message whose sender is itself — and with one shared user, bridge A cannot distinguish bridge B’s message from its own. This is the requirement; everything else is why it took two pieces.
  • The user was a single global value, read from config, not a property of a bridge.
  • There was no way to mark an account as a bot. A user carried a role and a status and nothing else.
  • The correct id was per-deploy and unguessable — and the wrong one was the owner. In production, id 1 is the operator and id 5 is the bot. On dev they are swapped.

Which is why the two pieces had to ship together or not at all. Adding a per-bridge user picker without an agent-account flag would have re-opened exactly the hole the UI phase closed when it removed the bot-user field: a form listing every human on the platform, one of whom is the owner with superadmin. That earlier removal was described in its own doc as “a real reduction, not a fix”, and it named this phase as the fix. Shipping only the relationship would have been a regression wearing the clothes of progress.

What landed: each bridge gets its own user, its own key, its own rate-limit bucket and its own audit trail — and, critically, real evaluated access with no admin bypass anywhere in the fleet. Four gates evaluated, not zero.

The chain a global admin had been hiding

Replacing a global-admin bot with a least-privilege account did something nobody planned: every check the old identity had been silently skipping became mandatory at once.

The bridge connected, reported healthy, heartbeated every sixty seconds for nine and a half hours, and received nothing. The operator sent a test message and got no reply, no error, no timeout, no clue.

The chain has three links — site membership, a channel ACL entry, and access to the post the channel hangs off — and all three refusals carry the identical envelope, with nothing in it to say which one refused. So fixing one and retrying looks exactly like the fix not working.

⚠ But the mechanism is sharper than “the errors look alike”, and the first version of this write-up got it wrong. Measured against the chat service: only post visibility refuses a join at all. The join gate consults post access and explicitly discards the channel-ACL result; membership and the ACL govern the bridge’s reply, not its arrival. The production rows date it exactly — membership at 08:21, channel ACL at 08:22, post access at 08:36, join succeeded at 08:38. Only the 08:36 grant was ever going to work. The first two were never consulted by the thing that was failing.

⭐ The chain always existed. It was invisible because every bridge ran as a global admin — a provisioning chain nobody wrote down, because the first account through it had permissions that made every link a no-op. Doing the right thing is what surfaced it.

And the fix was ruled out. The obvious move — have bridge creation grant all three — was rejected deliberately: granting a machine access to a private chat and a private post is exactly the act that should require a human to mean it. Automating it would let a bridge acquire reach into private content as a side effect of a form submission. The answer is legibility at the touch points, not automation.

Pushing Tin — managing a bridge fleet from inside the product
Pushing Tin — managing a bridge fleet from inside the product
Aug 29, 2026 Pushing Tin
← Back to Pushing Tin