Skip to content

Athena, the assistant

Operating a platform means knowing which of several hundred addresses to call, in what order, with what body. Most of that knowledge is not interesting: it is the cost of asking a simple question like “why is this deployment restarting”. Athena is the surface where the question is the input.

Athena is two services, and the split is the whole design.

The assistant holds the conversation. It talks to a model provider, keeps the session, and decides which operations to attempt. It never calls the control plane itself.

The tool server is the only thing that does. It publishes a fixed set of named operations — deployment_list, container_logs and sre_incidents are three of them — each with a declared input shape and a required permission, and it turns a call into a request against the control plane’s own API. Everything the assistant can do to the platform, it does through one of those named operations, inside one tenant, with a permission check in front of it.

The consequence is that the assistant’s judgement decides what to try, and nothing else. It cannot reach an address no tool names, and it cannot exceed a permission the tenant’s principal does not hold.

A person types into the console. The console calls the control plane’s Athena routes. The control plane strips any tenant identity the caller supplied and injects the one it verified itself — a request it cannot attribute to exactly one tenant is refused rather than served broadly, and the refusal is recorded. The assistant then mints a short-lived credential carrying that same tenant, and the tool server mints its own to call the control plane, with a value in the credential matched against a header on every write.

None of those hops widens scope. Being a platform operator is a privilege signal, not a wider view: an operator still acts inside one named tenant, and the code says so in three separate places because it is the assumption most likely to be re-broken.

Every tool declares the permission it needs. Which roles hold that permission is not written down in the assistant — it is generated from the control plane’s own grant.

That is a repair, not a preference. The two used to be hand-maintained copies, and they disagreed: a developer was offered the tool that raises an incident and then refused by the server halfway through the conversation. Three more disagreements of the same shape were found at the same time, including two permission names the control plane has never had. Generating one half from the other makes that class of defect impossible to reintroduce quietly — the roles reference is where that generated result now lives: every permission and the roles that hold it, read from the control plane’s own grant.

Of everything that is registered — the count is in the next section, and it is derived from the registrations themselves rather than typed here — eight are advertised in the model’s default list: five common reads and the three helpers that find everything else. An agent searches for an operation by intent, reads its input shape, and asks for it by name. A tool reached that way is checked against the same permission as a direct call — the asking is not the doing.

This keeps the advertised surface small without hiding anything. Every registered operation still carries a permission; the full list of those permissions and the roles that hold them is the roles reference.

Athena reaches the platform through 198 named operations, and every one of them falls into one of these areas. This is what you can ask it for:

  • Finding the right operation. Asking in plain words for something whose exact name nobody knows. Only a small window of the set is advertised to the model by default, and this is how the rest is reached.
  • What a spec may contain. The authority on what a deployment document is allowed to say — the fields the control plane really accepts and their shapes, taken from the control plane’s own generated schema rather than from a page someone kept up to date. Also the manifests it currently holds, read back exactly as they were uploaded. Reading only. Nothing here uploads a manifest or changes a schema.
  • Deployments. Creating, reading, changing, scaling, restarting and rolling back a deployment, and driving a canary release.
  • Jobs and schedules. Work that runs to completion: one-off jobs, their runs, and the schedules that create them.
  • Containers. The running containers of a deployment: listing them, inspecting one, reading its output, and the operations that change it — restarting, stopping, removing, and running a command inside.
  • Secrets. Secret metadata and rotation. No tool here returns secret material, and none accepts it — a value typed into a conversation cannot be un-written.
  • Volumes and backups. Persistent volumes and the backups taken of them.
  • Networks. Container networks and what is attached to each.
  • Cluster health and diagnostics. The fleet’s health and resources, and the diagnostic reads an agent uses when something is wrong.
  • The audit trail. What happened, who did it, and what was refused.
  • Incidents. The automated incident responder, from the agent’s side: reading incidents, approving a proposed action, and raising one.
  • Vulnerability scanning. What is known to be wrong with the images a tenant runs: scanning one, reading what a scan found, testing an image against the rules, and the upgrade that fixes it. The rules themselves — including which findings to ignore, and the thresholds that decide what blocks a deployment — can be read, rewritten and deleted here too, but only by a platform administrator, and they govern every tenant rather than one. Recording a risk exception against a rule is not here at all: neither raising one nor approving one has a tool, and both are human acts in the dashboard.
  • Replicated volumes. Storage kept on more than one node so losing one does not lose the data: where a volume’s copies are, its snapshots, moving it to another node, and switching over to a copy. These are not the same objects as the container volumes above.
  • Nodes. One node in the fleet: what it is, its certificates and edge-routing state, whether it accepts new work, moving work off it before maintenance, and what an agent upgrade would change. Admitting a node and approving its upgrade are not here — both are recorded human acts.
  • Configuration history. The kept history of a deployment’s configuration: which versions exist, what the configuration was at one of them, and putting an earlier one back.
  • Autoscaling. The rules that decide when a deployment gets more or fewer copies of itself, what each deployment’s rule is now, and what the platform has actually done about it — every scaling decision taken, and the totals across them.
  • Compliance reports. Evidence for an audit: asking for a report against one of the compliance standards this platform knows how to answer for, seeing which standards are available, listing what has already been produced, and reading one back.
  • Stacks. A group of deployments managed as one thing: what it contains, and what moving it to a different management model would change.
  • Migration tasks. Long-running platform work that moves stacks between management models: what is running, and how far it has got.
  • Platform-wide reads. Questions about the platform rather than about one deployment: what the edge router is serving, whether an address’s certificate is valid and when it expires, what the platform is costing, and how far a platform-wide environment migration has got. Also the scheduling priority classes this installation defines — which work is placed first when a node cannot hold everything. Also the residency regions this control plane declares, and whether geographic placement is being enforced against them.
  • Reclaiming disk space. Finding out what a platform-wide clean-up of unused images, stopped containers and dead networks would remove, and — as a separate request that must be confirmed — running it. Asking what would go is deliberately not the same sentence as making it go. Platform administrators only. A prune never touches volumes. Reclaiming space cannot cost a tenant their data.
  • Why nothing is converging. The machinery between an instruction and a running container, for when the instruction was accepted and nothing happened: the gateway that carries work to each node and the breakers that stop it when a node is failing, the per-tenant reconcilers whose job is to make reality match the spec, what this installation is actually running, and whether the platform’s own service registry is reachable. Platform administrators only.

Secret material never enters a conversation. The control plane’s own API accepts inline secret values; the tools deliberately do not, and a call that supplies one is refused rather than quietly stripped. Anything typed into a chat is in the transcript, the session store and the tool-call audit record, none of which are the secret store and none of which can be un-written afterwards. What the tools accept instead is a declaration that the platform should mint the material, a path to material that already exists, or a template resolved inside the control plane — see secrets.

Nothing crosses a tenant boundary, and running is not building: an assistant can start, change and stop deployments, and cannot produce the container images they run.

Three providers are supported, and the key can belong to the platform or to the tenant: a per-tenant key is read from the secret store under that tenant’s own path, and the platform’s key is the fallback. A deployment with no secret store at all can supply keys through the environment instead, at platform scope only.

Not this

This page does not list the tools or their permissions — that is the roles reference — and it does not describe the assistant’s own configuration surface, which is a platform-operator matter rather than part of a tenant’s API.

It also does not describe how a conversation is stored or how long for. That is a data-retention question this page cannot answer honestly from the code alone.