A write is refused and it is not your request
Symptom
A deployment write is refused with the code registry_credential_store_unavailable, and the message
says the control plane could not resolve registry credentials. Nothing about the request is wrong,
and no edit to the spec will make it succeed — the same body will be accepted once the platform is
healthy again.
This is the platform reporting a fault in itself rather than judging what you sent. It is worth
knowing the difference on sight: image_not_found means the registry answered and the image is not
there, which is yours to fix. This code means the registry was never asked, because the control
plane could not find the credentials to ask with.
Diagnose
- Confirm it is the platform and not the image. Read the code, not the prose.
image_not_foundis a confirmed-absent image.registry_credential_store_unavailableis the control plane saying it could not check. - Ask the control plane about itself.
GET /healthzanswers in one line: it reportsconsul not configuredwhen the control plane started without a key-value store client at all, andconsul unhealthywith the underlying detail when it had one and lost it. Those are different faults with different fixes, which is why the endpoint distinguishes them. - For more than a verdict,
GET /api/v1/cluster/healthbreaks the platform into components and gives the key-value store its own status —unavailablewith “Consul client not configured”, orunhealthywith “Consul is disconnected”. A platform administrator can read it; it is the same view the dashboard’s cluster health renders. - If it says not configured, the address the control plane was started with is the thing to
check —
ODYSSEUS_CONSUL_ADDRESS, or theconsul.addresssetting it overrides. A control plane that came up without one does not acquire it later on its own. - If it says disconnected, the store is configured and unreachable now: check that the Consul
container is running and that the control plane can reach it on the address it was given. A
restart of the store alone is not enough evidence that the control plane reconnected — read
/healthzagain afterwards rather than assuming.
Resolve
- As the caller: nothing, and that is the honest answer. Retry the identical request once the platform reports healthy. Nothing was written, so there is no half-applied change to undo and no container was disturbed.
- As an operator, not configured: set the key-value store address in the control plane’s configuration and restart it. Do not work around it by turning the image check off — the check exists because a write admitted without it tears down a healthy container and only then discovers the image cannot be pulled.
- As an operator, disconnected: restore the store, then confirm the control plane sees it again
on
/healthzbefore telling anyone the platform is back. The refusal clears by itself once the connection is real.
Prevent
Alert on the key-value store component of cluster health, not only on the control plane’s process being up. A control plane in this state answers reads, serves the dashboard and looks entirely alive while refusing every deployment write — so “the service is running” is not evidence that anyone can deploy.