Multi-Cloud Proof
A Kubernetes cluster has one control plane. Ours drives tenant nodes on more than one public cloud from a single control plane — and we can show you the live nodes, not a diagram of what the architecture would allow.
What is actually running today
Reading directly from the control plane's own service registry on 21 August 2026, one tenant's node fleet holds two nodes still enrolled and heartbeating within the minute, and a third that was retired after the test described below:
| Node | Cloud | Status | Agent |
|---|---|---|---|
alicloud-node-1 | Alibaba Cloud | ready | 0.7.9, specVersion 4 |
gcp-node-1 | Google Cloud | ready | 0.7.9, specVersion 4 |
aws-node-1 | Amazon Web Services | decommissioned after the June 2026 replication test | 0.3.51 at the time of the test |
alicloud-node-1's registered WireGuard endpoint is a public IP in Alibaba Cloud's own address range, which we independently confirm rather than take from the node's self-reported label. Both nodes are joined to the control plane over a WireGuard mesh, not a shared cloud VPC — the mesh is what lets one control plane treat two different clouds' networks as one addressable fleet.
Cross-cloud placement, proven live — not just designed
It is one thing to enroll nodes on two clouds. It is another to move a running workload between them without a gap where neither copy is healthy. We built a mandatory live gate for exactly that case (internally “LG6”): create a deployment pinned to alicloud-node-1, let it become healthy, then re-place it onto gcp-node-1. The assertion that matters is the third of seven: the alicloud-node-1 container is still present at the moment the gcp-node-1 container is created — and only removed once the new one is confirmed healthy. That is the difference between a guarded migration and a window where a cross-cloud move drops the workload entirely. The gate is documented as mandatory specifically because every other live gate at the time ran against single-node placements and could not have caught a regression here.
Three clouds, three continents: how the replication test was run
On 14 June 2026 a single control plane drove a three-node tenant across Amazon Web Services, Google Cloud and Alibaba Cloud, and replicated live volume data between them. This is how it was built and what was measured.
The mesh came first. The three nodes were joined into a full WireGuard mesh — six peer configurations, every pair reporting connected with live handshakes, each carrying a real public endpoint rather than a NAT guess. The agents that applied those peers run unprivileged: each one spawns a short-lived privileged helper container to write the /32 peer routes and exits. No cap_add on the agent, no compose change, no operator touching any of the three machines.
Then a replicated distributed volume. Primary on the AWS node, replica on the Google Cloud node, synchronised by an rsync daemon the agent runs from its own image — because a locked-down node can reach the private registry and nothing else.
The measurement. A marker file written into the primary volume on the AWS node appeared in the Google Cloud replica in about forty seconds, node‑to‑node across the mesh. The control plane was not in the data path at all — it decided placement and then got out of the way. That is the part no Kubernetes topology reproduces: the clusters would be separate, and something above them would have to broker the copy.
What it took — including the bug that was lying to us
Four agent releases stood between a connected mesh and real replication, each one a specific fault found by progressively deeper diagnosis:
- 0.3.48 — run the rsync daemon from the agent's own image; the public one could not be pulled on nodes that reach only the private registry.
- 0.3.49 — translate container paths to host paths, because the agent's own data directory is a named volume and its view of that path is not the host's.
- 0.3.50 — put the daemon on host networking; the bridge did not reliably serve traffic arriving on the WireGuard interface.
- 0.3.51 — bind the configured port. Without an explicit
portdirective the daemon fell back to the default, and the old bridge port mapping had been masking it.
The finding that matters most is the one we caught on ourselves. Before 0.3.51, the volume reported InSync. It was not. That status came from the storage plugin's own local materialisation marker, not from a completed transfer — a green light that meant "the volume exists here", not "the data crossed the ocean". The breakthrough was running netstat inside the daemon through the agent's exec API and finding nothing listening where we expected. We now measure replication by moving a file and looking for it on the other side, because a status field can be honest about the wrong question.
One race remained afterwards — overlapping deletes to the same destination could fail a sync — fixed in 0.3.52 and confirmed over ten consecutive runs. The work was promoted to production the following day, and the test volumes were deleted through the controller's own finalize path, verified gone on both nodes.
The AWS node was retired once the test had served its purpose. The mesh, the replication and the measurement are what the test proved; keeping a paid instance running afterwards would not have proved anything further.
Source: internal engineering log 2026-06-15_dvm-materialization-mesh-replication.md, which records the mesh verification, the four agent releases, the forty-second transfer, the exit-23 race and its ten-run confirmation, and the production promotion.
Beyond this one tenant: production runs on infrastructure we don't own
Separately from the two-cloud test fleet above, our production control plane manages a mixed fleet today: its own platform node, plus two customer-owned nodes running on infrastructure those customers control, not us. Customer-owned nodes cannot be upgraded remotely by us at all — they upgrade only through an explicit dashboard approval the customer drives themselves. That is a second, independent proof of the same underlying claim: one control plane, genuinely separate infrastructure, with the isolation boundary enforced by who can even reach the node, not just by network policy.
For the tenant-isolation mechanism that makes sharing a control plane across separate infrastructure safe, see Compliance.
How these numbers and claims were arrived at
- The node table was read directly from the control plane's Consul-backed service registry on 21 August 2026 (
odysseus/tenants/<tenant>/nodes/<name>), not from a design document describing the intended topology, and re-verified independently a second time the same day: bothalicloud-node-1andgcp-node-1readstatus: "ready"with alastSeentimestamp seconds old at read time.alicloud-node-1's public WireGuard endpoint address falls in an Alibaba Cloud-allocated IP range. - The
aws-node-1finding is the most load-bearing correction on this page: multiple design documents from June–July 2026 refer to a three-node fleet including it, and one earlier spec even recorded it as a DVM replication primary. The live registry read on 21 August 2026 shows no top-level node record for it — only two orphaned sub-keys (/rsync/secret,/services/rsync/status) with no name, status, or agent version. The scheduler's node-listing code (pkg/scheduler/consul_provider.go:88) explicitly skips any entry with an empty name field, and internal design docs note the same thing. We are treating the code's behavior, not the older documents, as the fact of the matter. - The cross-cloud placement gate (“LG6”) is drawn from the reconciler placement-safety design and workplan documents, which record it as mandatory and describe its seven assertions; we did not re-run the live gate ourselves for this page, and say so rather than implying a fresh run.
- The production fleet figures (platform node plus two customer-owned nodes, dashboard-approval-only upgrade path) come from a promotion runbook's measured starting state, corrected in the same document after an earlier count under-reported it. We did not independently re-verify prod's node count for this page beyond that document.