Start with the observed failure#

SymptomFirst checks
Login failsIdentity provider/local account, callback origin, session and MFA state.
No clustersConnection enablement, user role scope and configured cluster identity.
Resource list forbiddenOrkiva authorization, then the Kubernetes connection’s permissions.
Resource kind missingOptional navigation setting, resource registry and API availability in the cluster.
Logs/terminal disconnectSession state, container lifecycle and proxy WebSocket support.
Metrics emptyPrometheus discovery/URL, reachable service and available metric series.
Delivery queuedWorkers, quotas, environment concurrency, windows, dependencies and maintenance state.
Approval rejectedCurrent evidence, independent approver, MFA and recent authentication.
AI unavailableEnabled model, conformance, connection, policy limits and network route.

Collect actionable evidence#

  1. Record the time, affected version, cluster, namespace, resource or run identifier.
  2. Capture the exact error code and request identifier when available.
  3. Inspect the related resource Events, run timeline or dependency-health page.
  4. For an installation issue, download a support bundle through Operations and review it before sharing.
  5. Describe the intended action and observed result separately.
Operations provides a support bundle with a visible file allowlist. Review the exported information before sending it to support.
Operations provides a support bundle with a visible file allowlist. Review the exported information before sending it to support. View full size ↗

Avoid compounding uncertain operations#

Do not treat a timeout as proof of failure. An operation may have reached Kubernetes or the database before the connection failed. Inspect current state and the available idempotency/run evidence before issuing another mutation. For unknown-outcome delivery, use its recovery guidance rather than deleting history.