Reference 13 — Network formation

In one paragraph

A libtracer network is formed by ordinary vertex writes. A third party — typically a web UI — joins as an ephemeral peer with delegated admin, then on other devices it (1) creates controllers and transport connections through one in-band mechanism (ADR-0017 — Vertex creation is an in-band, ACL-gated field-write, ADR-0027 — A transport, and each connection within it, is a first-class / vertex), and (2) binds data flows by issuing consumer-initiated subscribe-writes into producers’ :subscribers[] (ADR-0026 — Subscription is consumer-initiated). It then disconnects, leaving the wired devices talking to each other. A node is one path tree — data endpoints, controllers, and transports — all addressed, created, ACL’d, await’d, and reconciled uniformly. There are no privileged roles: “orchestrator” and “router” below are transient situations any peer can be in. The network is pure-decentralized and self-healing — it depends on no central authority, and bindings re-establish themselves on reconnect.

Formation uses no orchestration-specific wire behaviour. It composes the three primitives — read, write, await — with the ACL and subscriber-edge surfaces specified elsewhere in this suite: 04 — Communication flows describes the data plane, and this section describes the formation plane built from the same operations.

Transient hats, not fixed roles

None of the following is a fixed role or a privileged node — they are hats any peer wears transiently. The same peer is a producer on one edge and a consumer on another; “orchestrating” means a peer holding admin and issuing formation writes; “forwarding” means a peer that has ≥2 transports. The network has no central authority.

Hat

What it means (transient)

Owner

A peer holding the provisioned root token that bootstraps a device’s ACL and delegates admin (CONTEXT.md ACL / subject-token).

Orchestrating

A peer the owner granted WRITE_ACL (admin) that is issuing formation writes. Usually a web UI, joining temporarily. Not architecturally special — a peer doing vertex writes, then leaving.

Producing

A vertex that holds an edge and fans out (e.g. /A/sensor).

Consuming

A vertex that receives delivery (e.g. /B/in). Control-passive, data-rich.

A peer that is orchestrating is an edge that exists temporarily → modifies bindings → departs, leaving producer and consumer wired. Because formation is just vertex writes, the cables it patches outlive the hand that plugged them — and because nothing privileged holds the graph together, a rebooted or reconnected peer re-forms its own bindings (§Self-healing without a coordinator) with no coordinator present.

A peer is anything that speaks the wire format: an MCU, a host process, a container, a browser tab. Peers are symmetric, and the transport does not change a peer’s standing: the same graph forms over CAN, WebSocket, TCP or UDP. A co-processor relationship (one MCU front-ending another on the same board) is a property of a product’s board layout, not of the protocol; the two parts are still ordinary peers on their shared bus.

The formation flow

        sequenceDiagram
    participant O as Orchestrator (web UI, temp admin)
    participant B as Consumer device B
    participant A as Producer device A
    Note over O,A: 0. discover peers (mDNS / static)
    Note over O,B: 1. owner delegates admin → O
    O->>B: 2. write /B/net/quic-client/conn SPEC{name=a, config{addr=A_addr}}
    B->>A: QUIC dial (consumer dials)
    O->>A: 3. write /A/sensor:subscribers[] += SUBSCRIBER{target=/B/in}
    Note over A: A:acl authorizes the subscriber (fan-out gate)
    A-->>B: 4. fan-out: delivery = ordinary write to /B/in
    Note over B: B:acl on /B/in authorizes the writer (fan-in gate)
    O--xO: 5. orchestrator disconnects — A↔B persist
    

Discovery

Peers are found by a discovery module emitting (transport_label, transport_address) tuples — discovery_static (pre-configured) or discovery_mdns (dynamic announce); see 10 — Module catalog, which also records that both are v1 scope with no reference module, and what the committed pieces are. Version compatibility is settled here, not per-frame: a distinct service name, port or CAN-ID prefix per protocol version (ADR-0013 — Protocol-v1 scope boundaries); for mDNS that name is _libtracer._tcp (ADR-0002). See also 07 — Host embedding.

Erratum (2026-08-24) — the tuple carries no identity

This tuple used to be written (peer_id, transport_label, transport_address). That contradicts a decision already taken: carrying identity at the discovery layer was considered and rejected (RFC-0011 §Alternatives), because discovery is IP-peer bootstrap — “what can I dial” and nothing else (ADR-0044 point 4) — while CAN peers and nodes behind a single IP entry never appear in a discovery record at all, and the consumers that need a node’s key need it through any path. A node’s identity is read in-band, from read <vertex>:identity (RFC-0011), after the dial. If a discovery record does carry a peer_id-shaped token, treat it as an unauthenticated dial hint and never as a subject: nothing authenticates it, and the ACL plane must not be keyed on it. Text-only — the wire surface does not move.

A browser peer discovers nothing. A JS/WASM binding cannot do mDNS at all, and the formation flow does not ask it to: a browser joins by dialing a URL it was given (§The formation flow — it is the orchestrator, the transient hat, not a discovered peer), and it learns about other peers in-band, by reading the graph it just reached — :children[] on a net root, the peer facets of a bus link. Discovery therefore happens on a native peer, and the browser’s knowledge of the network arrives the same way every other fact does: as a read. This is why the tuple above is a module’s output rather than a step every peer performs.

Admin delegation

A device persists identity only — a stable peer_id, and a PKI key as a stronger subject-token where one is provisioned. It does not persist graph wiring. The owner peer grants the orchestrator WRITE_ACL on the subtree it may manage (NFSv4-style ACE with INHERIT; ADR-0020 — Access control uses NFSv4-style ACEs with inheritance). The orchestrator holds delegated admin for the duration of its session.

Creation — controllers and transport connections through one mechanism

Creation is an in-band write of a SPEC to a device-designated creator-endpoint vertex (ADR-0059 — Creation and removal are writes to a creator endpoint vertex, superseding the :children[] creation-field spelling of ADR-0017). The same mechanism brings up a transport link, because a transport — and each connection within it — is itself a vertex (ADR-0027). For a controller the SPEC names a device-catalog type; for a transport connection the endpoint’s location supplies the type, as below.

A controller exposes its own input-port and output-port vertices and subscribes to nothing at creation — the patch-cable model: creation exposes ports, binding is separate.

The creator endpoint for transport connections

RFC-0014 — Creator endpoint: connection lifecycle and link liveness (accepted) specifies one creator endpoint per (transport, role) module rather than one global catalog. A creatable pair is a self-contained module mounted flat under the net root — conventionally /net, which is a recommendation (the constructor default, overridable per node), never a library rule — with names like ws-client, ws-server, quic-client, can, …, each declared by the application through register_module (modules are declared-only, ADR-0073 §4; the spellings here are the built-ins’ suggested names). Each module exposes a creator-endpoint child named conn:

/net/<module>/conn                          ; the creator endpoint (a / vertex, not a : field)
   write SPEC{ name, config }   → create /net/<module>/<name>, atomically
   write NAME{ <name> }         → retire /net/<module>/<name>
   read  :schema                → the module's config catalog
/net:children[]                             ; enumerate the modules
/net/<module>:children[]                    ; enumerate that module's connection vertices
  • Create and remove collapse onto one control, distinguished by the TLV type of the written value: SPEC creates, NAME removes. Any other payload is ERROR{tr::schema::type_mismatch}; the endpoint never falls through to an ordinary assign. A NAME naming an unresolvable connection is a no-op success; NAME{conn} is rejected, so the endpoint cannot self-destruct.

  • The role is positional because it is the module. ws-client is DIAL, ws-server is LISTEN, can is a multi-peer bus. SPEC therefore carries { name, config } with no type and no role field, and each module’s :schema is its own catalog — a dial target and a bind port are different catalogs, which is the reason to split them.

  • This endpoint is the ONE door. The superseded global spelling write /net:children[] += SPEC{type, name, config}, with its client / listener child types and its role config override, was retired at #492 S7: the reference implementation registers those two child types no longer, so that write now answers SCHEMA_NOT_FOUND. There is no second way in and no migration window. What survives under the same spelling is :children[] as an enumeration — the reads in the block above — and :children[] creation for stored_value and any type an application registers itself with graph_t::register_child_type; only the two connection types died.

  • The path carries a name, never an address. The created connection is addressed /net/<module>/<name>; addr, port, backoff and connect_timeout are creation-time config — they travel in the SPEC’s config SETTINGS and are parsed into the transport-private tr::net::conn_settings_t. They do not live in the vertex :settings namespace: that core namespace was emptied outright by RFC-0022 §3.B, and conn_settings_t is explicitly not part of any vertex’s protocol :settings surface — a device-private facet (ADR-0021 §Decision 3) reached through the transport’s own config door. Nor is there a post-creation reconfiguration door: the only accessor, transport_vertex_t::settings_of, hands out a const view. So a peer’s IP changing is not a :settings edit — today it means retiring the connection (NAME) and re-creating it (SPEC), which does tear down the routes under it, because remove_connection un-routes and retires the identity vertex. Routing that makes /net/<module>/<name> addressable is ADR-0061 — Per-module mount routing.

  • All four of those keys are consumed today. (A keepalive pair is still accepted but ignored — it had no consumer, so #1666 removed the field that stored it.) backoff / connect_timeout feed the §4 liveness engine (#492 S5, self_heal_link_t) on a kind registered self_heal_dial — since #1548 every built-in point-to-point DIAL kind (udp, tcp, ws). On a kind not opted in (a bus kind such as can, or a provided link) they are parsed with no consumer. The per-key record, including which kinds honour max_frame, is the connection-config module page.

  • SPEC naming an existing name is PATH_IN_USE — a re-SPEC is never a reconfiguration. A retrying orchestrator can treat that rejection as “already exists”, so the create is idempotent-safe.

  • Creation is atomic. One write yields a fully configured connection vertex; there is no live-but-unconfigured window.

  • Gating reuses the existing access-mask bits (see 05 — Protocol TLVs): SPEC (create) gates on CREATE (0x08) on the endpoint’s own ACL, so the create right is delegable without any right on the parent transport; NAME (remove) gates on WRITE (0x02), not on DELETE — DELETE (0x10) is reserved-and-unused for protocol v1 (RFC-0009 — Vertex removal and subscriber eviction §A.2). A peer can hold create-but-not-remove, or the reverse.

  • An absent creator endpoint answers PATH_NOT_FOUND — the missing-/-vertex answer, and the sanctioned creatability probe, since conn is hidden from /net/<module>:children[]. SCHEMA_NOT_FOUND is reserved for the distinct “endpoint present, config type unknown” case.

Realisation status

The per-module creator endpoint is implemented (RFC-0014 S2b): declaring a module mints /net/<module>/conn, a SPEC{name, config} written there creates /net/<module>/<name> and a NAME{<name>} removes it, the conn name is reserved in both directions, and the endpoint is hidden from /net/<module>:children[] (S4) — that listing returns the module’s member connections only, while the endpoint itself stays addressable for the §6 creatability probe. The liveness engine (S5) is implemented too (tr::net::self_heal_link_t) for kinds registered with transport_kind_traits_t::self_heal_dial, and since #1548 the built-in point-to-point kinds (udp, tcp, ws) are opted in: a built-in DIAL connection is minted DORMANT with no socket and dials on first use, so creation no longer fails when the peer is down. LISTEN connections still bind eagerly and report LISTENING at creation, and bus kinds (can) stay eager by design. The module-side half of S3 is in too (#1815): a module may declare a conn:schema catalog at register_module, which its endpoint serves inside Amendment 3’s SETTINGS and validates each SPEC against; a module that declares none answers the conforming empty SETTINGS. A routed subscription holds its link’s refcount since #1816, and the CREATE/WRITE gating split (S2c) is in. The addressing half was already there — a created connection mounts and routes at /net/<module>/<name>, with the module name declared by the application (never library-derived — ADR-0073 §4) — and the :children[] creation spelling RFC-0014 supersedes is gone: S7 unregistered the client and listener child types, so the endpoint is now the only door and a :children[] creation SPEC naming either type answers SCHEMA_NOT_FOUND. RFC-0014’s byte-level clauses (the SPEC/NAME/config layout, the catalog reply bytes, the liveness encoding, the gate order, the error identities) are normative since Amendment 4; its declaring clauses stand on acceptance.

Binding — consumer-initiated subscribe-writes

Data flow is established by the consumer acting as a client (ADR-0026): a write into the producer’s :subscribers[], carrying the consumer as target. The edge is producer-held — the producer fans out; the consumer holds nothing.

write /A/sensor:subscribers[]      += SUBSCRIBER{ target = /B/ctrl/0/in }
write /B/ctrl/0/out:subscribers[]  += SUBSCRIBER{ target = /B/actuator }

These two writes are a chain, not a loop. A delivery landing on /B/ctrl/0/in stops there and never re-fans from that vertex (RFC-0007 — SUBSCRIBER delivery terminates at the target, ADR-0051 — Delivery terminates at the target; see 04 — Communication flows). The second hop happens because the controller’s own logic writes /B/ctrl/0/out, and that write fans out to /B/ctrl/0/out’s own subscribers. Propagation past a target is the target’s logic, never the dispatcher’s.

The orchestrator issues these writes on the consumer’s behalf; a device’s firmware or NVS config issues the identical write on boot. Same operation, different driver — there is no privileged “default binding”.

Delivery and the two ACLs

Delivery to a target is an ordinary write, indistinguishable from a direct one (the target is subscription-unaware at runtime). Protection is the two endpoints’ ordinary ACLs, with no extra machinery:

Direction

Guard

Question it answers

Fan-out / confidentiality

producer’s :acl

who may subscribe to me?

Fan-in / sink protection

consumer’s :acl on the target (+ firmware arity)

who may write into me?

So “multiple publishers will not feed a single sink” is enforced device-locally, even with no orchestrator present — a single-input sink rejects a second writer via its own ACL. Rejection lands at delivery time on the consumer (REST-server-auth shape), not at bind time on the producer.

Departure

The orchestrator disconnects. The created controllers and transport connections remain in the devices — RAM, or NVS where the device persists them.

Note

A subscriber edge survives the orchestrator’s departure when — and only when — its target routes through a mount. A SUBSCRIBER whose PATH names a path through a transport mount, spelled in the producer’s frame (/net/<module>/<link>/<consumer-path>), binds the edge to that mount’s link and the residual below it (graph_t::subscribe_wire, core/src/graph.cpp:delivery_link.assign(split.link);), so fwd_router_t::link_down → graph_t::evict_link_edges on the orchestrator’s session no longer matches it and the producer keeps delivering. That is RFC-0021 §4.B.1/§4.C, and it is what makes the departure above real for a third-party wire.

A SUBSCRIBER with no PATH, or one whose PATH matches no mount, still binds to the arrival session and delivers back along the accumulated src — the consumer-subscribes-for-itself shape, unchanged. A target that names a mount it cannot deliver through (the mount exactly, or a bus link’s own NAME — RFC-0020) is refused with tr::path::invalid, never silently degraded. See §Boundaries.

Two devices keep talking with no third party present; the patch cable stays. A rebooted leaf re-establishes its links and subscriptions by re-issuing the same client-writes from firmware or NVS config.

The orchestrator never carries the data. It issues create and bind writes and leaves, and the devices deliver to each other directly; proxying the data through the orchestrator is the browser relay this model exists to retire.

Connection direction and folding

The default that pairs with consumer-initiated subscription is the consumer dials, the producer pushes (SSE / server-streaming shape) — it also lets a constrained leaf dial out through NAT. Direction is fixed by the module the connection was created under, so it is chosen per connection: a constrained producer with many consumers, or NAT on both sides, is served by dialing out to any peer that has ≥2 transports (a forwarding hop, not a “router” role).

Any node with ≥2 transports forwards — forwarding is a required capability the moment a node has two wires (07 — Host embedding). There is no privileged router node, so the network folds arbitrarily — elided-CAN leaf → full-TLV QUIC backbone → another fold — with the forwarder stateless and uniform across framing modes. The bounds to design within:

  • Depth is capped by the route, and the route by its bytes. A FWD frame’s dst names every hop and is consumed monotonically, so a delivery travels exactly as far as its explicit source route — segment count ≤ 255 (03 — Addressing; RFC-0023; kMaxSegments, core/include/libtracer/path.hpp:kMaxSegments). For realistically named mounts the 1024-byte PATH budget binds first, not the segment count: a 3-segment mount run (ADR-0061) costs its NAME headers plus its bytes — 20 B/hop for /net/can/c0, 32 B/hop for /net/ws-client/board-01 — so the diameter is ≈ 30–50 hops at 3-segment mount runs, and ≈ 25 at a 5-segment one. (Arithmetic over the encoding rule at 05 — Protocol TLVs §0x06, not a routed measurement — RFC-0023 §4.3, §10.)

  • A bound route is capped by hosts, not by bytes. The second address form (PATH_REF, 05 — Protocol TLVs §0x14; RFC-0024 §4.3) spells one fixed 8-byte element per host rather than a run of names per hop, so its diameter is a host count: ≤ 255 normatively, and ≤ 69 reachable today, since every bound path is minted from a canonical one and inherits that form’s 1024-byte budget (≤ 171 under packed segments). The canonical ceiling above is therefore the binding one in practice, and the bound form’s own cap sits above it deliberately — it is a property of the element, not of whichever body grammar PATH happens to use.

  • Loops cannot form. Because dst is consumed by at least one segment per hop (a whole net/<module>/<name>[/<peer>] mount run, RFC-0014 S2a), a physical cycle is harmless per-op rather than rejected. There is no revisit check — loop-freedom is by construction, not by rejection — and no flooding, so no duplicate deliveries and no dedup state anywhere. Parallel links to one peer are distinct explicit addresses: deliberate redundancy a consumer subscribes to knowingly. A recursive topology walk is not protected by this. A walking orchestrator must carry its own client-side, identity-keyed visited set, keyed on the node identity served by the :identity facet (RFC-0011 — Node identity facet) — a node-scoped, pre-serialized record that every vertex of a node serves byte-identically, resolved above the READ gate so an unauthenticated peer can pin it.

  • No global ordering across folds. Per-producer ordering only; cross-node coherence needs a coordinated trigger.

Self-healing without a coordinator

Because nothing privileged holds the graph together, recovery is local and automatic:

  • Subscriptions re-form themselves. On reconnect a consumer re-issues its subscribe-write from firmware or NVS config (ADR-0026) — the binding repairs without anyone re-provisioning it.

  • Transport-native bindings re-learn in-band. Elided or lean bindings — a CAN id↔path map held inside the transport — re-establish from advertise frames (advertise + id-match), so a rejoining node re-announces its own mappings; see 14 — CAN transport.

  • The link layer re-dials itself. Under §Link liveness, a DIAL link with a standing binding retries toward up with backoff and no give-up bound, so a lost socket repairs beneath an unchanged subscriber edge.

  • There is no central authority to lose. Any peer can wear any hat; a departed orchestrating peer or a downed forwarding hop costs only the paths through it, and the rest of the mesh is unaffected.

Boundaries of the formation model

  • The model is imperative, not declarative. It specifies the writes that produce a wiring, not a desired-state manifest. A reconciler that diffs a manifest against live state (read of :children[] / :subscribers[]) and converges by issuing exactly these create-and-bind writes is tooling over this wire model; it adds no wire behaviour and is out of scope here.

  • The subject token is opaque and key management is out of scope. The ACL model takes a subject token from a pluggable resolver and compares it byte-for-byte; how a peer obtains, rotates or revokes one is not specified by the formation model. The :identity facet publishes a node’s public key for pinning; it is not a key-management protocol.

  • Origination of a remote-owned subscription is specified only for a mount-routed target. A cross-device wire is a subscription, not a link. A departing orchestrator leaves board A subscribed to a producer on board B by writing a SUBSCRIBER whose PATH is spelled in B’s frame — composable offline from the link name the orchestrator itself minted, no read-back and no route walk (RFC-0021 §4.B.1, #491). What is not specified is a wire SUBSCRIBER naming a purely local target on the producer (RFC-0021 §4.B.2, unruled): that spelling keeps the arrival-session binding. The bus-crossing variant is refused outright, not deferred — a bus link’s own NAME is not a routable next hop (RFC-0020).

  • Teardown is soft or hard. Soft — drop the last binding; the DIAL link goes dormant, the vertex persists, and it self-heals on next use. Hard — a NAME write retires the vertex; subscriptions routed through the link are cascade-evicted (RFC-0009 §D) rather than left dangling, and a far-side producer’s edge back through the retired link fails fast with link-down and is reaped.

Pitfalls

  • Reading listening as “a peer is attached”. A LISTEN vertex’s liveness reports only that its listen socket is bound and accepting. An implementation that surfaces it as connectivity shows a server as healthy with zero peers attached and as unhealthy with many.

  • Re-SPECing to reconfigure. A SPEC naming an existing connection is rejected PATH_IN_USE. An implementation that treats reconfiguration as “create again” never changes the address — and there is no :settings edit to reach for instead. addr and port are creation-time config (§Creation): they travel in the SPEC’s config and are parsed into the transport-private tr::net::conn_settings_t, whose only accessor hands out a const view (transport_vertex_t::settings_of, core/include/libtracer/transport_vertex.hpp:transport_vertex_t::settings_of). The vertex :settings core namespace holds nothing to write — RFC-0022 §3.B deleted settings_t, so every flat knob name under it answers SCHEMA_NOT_FOUND caller-independently (core/src/graph_fields.cpp:field_surface_t::write_settings), leaving only the read container and its reserved app subkey (core/src/graph_fields.cpp:graph_t::read_settings). Moving a peer therefore means retiring the connection (NAME) and re-creating it (SPEC), which un-routes the link and cascade-evicts the subscriptions routed through it (§Boundaries of the formation model, hard teardown).

  • Expecting a data write to revive a retired connection. Connection vertices are an exception to write-creates; the write fails and the peer stays unreachable until a SPEC recreates it.

  • Walking a folded topology without a visited set. FWD loop-freedom protects a delivery, because dst is consumed monotonically; it protects nothing about a recursive enumeration. A walker that keys its visited set on transport address rather than on the :identity record revisits the same node through a second link and does not terminate.

  • Treating a successful :subscribers[] append as proof the flow works. The fan-out gate answers on the producer; the fan-in gate answers on the consumer at delivery time. A bind that the producer accepts can still be denied on every delivery, and the orchestrator that issued it is gone by then.

  • Assuming a one-shot op leaves a link retrying. The transient hold is released before self-heal is evaluated, so a one-shot against an unreachable peer leaves the link dormant with no retry in progress. Only a standing binding makes the link self-heal. (A one-shot that succeeds is the other case and the opposite mistake: its socket may well still be up, because keep-up-until-loss at refcount 0 is a MAY the reference implementation takes — RFC-0014 §4.1. Neither is a state to rely on; take a standing binding if you need the peer reachable.)