Skip to content

A write is refused and it is not your request

Symptom

A deployment write is refused with the code registry_credential_store_unavailable, and the message says the control plane could not resolve registry credentials. Nothing about the request is wrong, and no edit to the spec will make it succeed — the same body will be accepted once the platform is healthy again.

This is the platform reporting a fault in itself rather than judging what you sent. It is worth knowing the difference on sight: image_not_found means the registry answered and the image is not there, which is yours to fix. This code means the registry was never asked, because the control plane could not find the credentials to ask with.

Diagnose

  1. Confirm it is the platform and not the image. Read the code, not the prose. image_not_found is a confirmed-absent image. registry_credential_store_unavailable is the control plane saying it could not check.
  2. Ask the control plane about itself. GET /healthz answers in one line: it reports consul not configured when the control plane started without a key-value store client at all, and consul unhealthy with the underlying detail when it had one and lost it. Those are different faults with different fixes, which is why the endpoint distinguishes them.
  3. For more than a verdict, GET /api/v1/cluster/health breaks the platform into components and gives the key-value store its own status — unavailable with “Consul client not configured”, or unhealthy with “Consul is disconnected”. A platform administrator can read it; it is the same view the dashboard’s cluster health renders.
  4. If it says not configured, the address the control plane was started with is the thing to check — ODYSSEUS_CONSUL_ADDRESS, or the consul.address setting it overrides. A control plane that came up without one does not acquire it later on its own.
  5. If it says disconnected, the store is configured and unreachable now: check that the Consul container is running and that the control plane can reach it on the address it was given. A restart of the store alone is not enough evidence that the control plane reconnected — read /healthz again afterwards rather than assuming.

Resolve

  • As the caller: nothing, and that is the honest answer. Retry the identical request once the platform reports healthy. Nothing was written, so there is no half-applied change to undo and no container was disturbed.
  • As an operator, not configured: set the key-value store address in the control plane’s configuration and restart it. Do not work around it by turning the image check off — the check exists because a write admitted without it tears down a healthy container and only then discovers the image cannot be pulled.
  • As an operator, disconnected: restore the store, then confirm the control plane sees it again on /healthz before telling anyone the platform is back. The refusal clears by itself once the connection is real.

Prevent

Alert on the key-value store component of cluster health, not only on the control plane’s process being up. A control plane in this state answers reads, serves the dashboard and looks entirely alive while refusing every deployment write — so “the service is running” is not evidence that anyone can deploy.