Troubleshooting
A platform whose probes are green may very well never execute a single run.
That is the sentence to keep in mind: most first-install failures do not show up in a
kubectl get pods. This page lists the ones that were actually met, with the exact symptom rather
than a generic category.
1. The API refuses to start, and says why
Section titled “1. The API refuses to start, and says why”Symptom: the API pod stops before listening, with a message that names a variable.
That is deliberate. An init container checks the mandatory keys before start-up and fails by
naming them — it never names their value. It also refuses a key that is present but literally
empty or equal to <no value>, which is what an unsubstituted template produces.
Fix: set the key named in the message. Do not bypass the guard: it saves you from a platform that starts and fails later, elsewhere, with no apparent connection.
2. Creating a tenant fails with a configuration error
Section titled “2. Creating a tenant fails with a configuration error”Symptom: creating a tenant returns a tenant configuration error, although the platform is running.
Cause: the internal address of the object storage is not configured. The application needs to reach the storage from inside the cluster, under a name that may differ from the public address.
3. The token is rejected although authentication works
Section titled “3. The token is rejected although authentication works”Symptom: sign-in succeeds in the browser, but the API rejects the token.
Cause: the public issuer address and the internal one diverge. The token carries the issuer as seen by the browser; the API compares it to the one it knows. Both must designate the same issuer.
4. Runs fail silently
Section titled “4. Runs fail silently”Symptom: the platform answers, the probes are green, a run starts — and never completes, with no useful message.
Cause: the internal URL of the API is not configured. The launched workload must be able to call the platform back to report progress; without that address, it runs into the void.
This is the most expensive failure on the list, because every visible signal is green. If a run reports nothing, check this variable before anything else.
Before opening a ticket
Section titled “Before opening a ticket”Three checks that settle most cases:
- Did the init container pass? If it failed, it names the key at fault.
- Are the internal addresses (object storage, API, authentication issuer) set and reachable from inside the cluster?
- Are the object storage and the image registry reachable from the execution namespace, and not only from the platform one?