Debug your deployment
The self-hosted-controller (SHC) provides a set of diagnostic commands that let you validate configuration, inspect cluster health, audit secrets, diagnose database issues, and analyse resource usage. This article is a full reference for those commands, including expected outputs and remediation steps.
All commands are executed from the orchestration node using the exec subcommand:
exec <COMMAND>
Configuration validation
CheckLocalConfig
Validates every key in your config.yml against the SHC schema.
exec CheckLocalConfig
The check reports:
- Missing required fields.
- Values that do not match the expected format or regex.
Expected output (no errors)
Validating configuration...
Configuration is valid entries_checked=42
Example failure output
Missing required config key=global.version.platform.version
Format mismatch key=utils.oci_registry.host value=https://registry.lab expected_format=^[a-zA-Z0-9]
Configuration validation failed: 2 error(s) found
What to do after a failure: Review each error line. Add any missing required keys to config.yml. Fix values that do not match the expected format. For example, utils.oci_registry.host must be a bare hostname with no http:// or https:// scheme.
To inspect all resolved environment variable values before the schema check runs, add the -v flag:
-v exec CheckLocalConfig
The verbose output includes the fully resolved in-memory config tree, including every ${env.VAR_NAME} value substituted with its actual content. Use this to confirm that secrets injected via environment variables are correctly loaded.
Infrastructure connectivity
CheckServersAreReachable
Tests SSH connectivity to all nodes listed in utils.ansible.inventory.
exec CheckServersAreReachable
Expected output
Pinging all configured servers via Ansible...
All configured servers are reachable
What to do after a failure:
- Verify the SSH key path in
utils.ansible.ssh-keyis correct and exported in your environment. - Confirm the target node is running and reachable from the orchestration node on TCP port 22.
- Verify the username in
utils.ansible.userhas SSH access to the node. - If you use password-based sudo, confirm
SERVERS_SUDO_PASSWORDis set correctly.
CheckServerSpec
Runs the check_servers_spec Ansible playbook against every manager and worker node in utils.ansible.inventory, and reports the first requirement each node fails.
exec CheckServerSpec
The command checks the following, in this order.
| Check | Requirement |
|---|---|
| Unique hostnames | No two nodes share a hostname |
| OS family | Debian-based |
| OS version | Debian 12 or later |
| CPU | 44 cores or more per node |
| RAM | 120 GiB or more per node |
| NTP enabled | timedatectl reports NTP=yes |
| Clock synchronized | timedatectl reports NTPSynchronized=yes |
| Port availability | TCP 80, 443, 2379, 2380, 4240, 4250, 6443, 10514, and 11514 can be bound before K3s is installed |
| Dedicated storage disk | One unused block device of 200 GB or more, with no partition table and no filesystem, before K3s is installed |
The last two checks are skipped once K3s is installed on the node. After the installation, the ports are bound by the cluster and the storage disk is consumed by Ceph and Longhorn, so requiring them to be free would fail every re-run.
What to do after a failure:
| Failure | Remediation |
|---|---|
| Duplicate hostnames | The error lists the inventory host to hostname mapping. Rename the affected nodes with hostnamectl set-hostname <NEW_HOSTNAME>, then re-run the command. |
| Debian version too old | Reinstall the node on Debian 12 or later. See Technical requirements. |
| CPU or RAM below the minimum | Resize the node. The error reports the detected value. |
| NTP not enabled | Install and enable systemd-timesyncd or chrony, then run timedatectl set-ntp true. |
| Clock not synchronized | Verify that the node reaches your NTP servers, then run systemctl restart systemd-timesyncd or systemctl restart chrony. |
| A required port is already bound | Run ss -tlnp on the node to identify the process holding the port, then stop it or reconfigure it to another port. |
| No unused extra disk | Attach a block device of 200 GB or more to the node and leave it unformatted, with no partition table. A disk that already carries a filesystem or partitions is rejected. |
CheckKubernetesCluster
Connects to the Kubernetes API and verifies that every node has a Ready=True condition and that the actual node count matches your Ansible inventory.
exec CheckKubernetesCluster
Expected output
Kubernetes cluster is healthy nodes=6 expected=6
What to do after a failure:
- Run
kubectl get nodesto identify which nodes are notReady. - Run
kubectl describe node <node-name>and look for taints, conditions, or resource pressure. - Check K3s system logs on the failing node:
journalctl -u k3s -n 100.
Local artifact checks
CheckLocalReleaseFiles
Verifies that every asset declared under global.version exists in the expected directory on the orchestration node.
exec CheckLocalReleaseFiles
What to do after a failure: Confirm that the release archive was fully extracted and that global.version.platform.path points to the correct directory.
CheckLocalGit
Clones the repository configured in utils.git.repo_url and tests both pull and push access.
exec CheckLocalGit
What to do after a failure:
- Verify
GIT_HTTP_USERNAMEandGIT_HTTP_PASSWORDare set correctly. - Confirm the repository exists and is reachable from the orchestration node.
- Ensure the user has both read and write permissions on the repository.
CheckLocalOCIRegistry
Tests push, pull, and delete access to your OCI registry. If push fails, pull and delete are skipped and reported as untested.
exec CheckLocalOCIRegistry
What to do after a failure:
- Verify
REGISTRY_USERNAMEandREGISTRY_PASSWORDare set correctly. - Confirm the registry URL is reachable from the orchestration node.
- Verify the
check_repopath inutils.oci_registry.check_repopoints to an existing image in your registry.
Application health
DebugArgoCD
Renders a three-panel status dashboard for all ArgoCD repositories, the root application, and every managed application.
exec DebugArgoCD
Reading the application table:
| Sync status | Health status | Action |
|---|---|---|
| Synced | Healthy | No action required. |
| Synced | Progressing | Wait 2-3 minutes, then re-run. Normal during deployments. |
| OutOfSync | Any | Run DebugArgoCDSyncAll to force re-synchronization. |
| Any | Degraded or Missing | Inspect pod logs and run DebugDatabases. See remediation below. |
What to do for Degraded or Missing applications:
- Run
kubectl get pods -n <namespace>to identify failing pods. - Run
kubectl logs -n <namespace> <pod-name>to view pod logs. - Run
kubectl get events -n <namespace> --sort-by='.lastTimestamp'to view recent events.
DebugArgoCDSyncAll
Forces a three-phase full re-synchronization of all ArgoCD applications in parallel.
exec DebugArgoCDSyncAll
The three phases are:
- Partial sync of
secretgeneratorandconfigmapresources to refresh secrets before regeneration. - Restart of the
sekoiaio-secret-operatordeployment to pick up refreshed SecretGenerators. - Full sync of every application.
Sync behavior
The sync timeout per application defaults to 300 seconds. Up to 32 applications are synced concurrently. Override these values in config.yml under modules.debug_argocd_sync_all.
Use this command when applications are stuck in OutOfSync, after a manual change to the ArgoCD repository, or after a platform upgrade.
Secret diagnostics
DebugMissingSecrets
Compares declared SecretGenerator CRDs against actual Kubernetes Secret objects and reports:
- Secrets that are entirely missing.
- Secrets that exist but have incomplete keys.
- The Vault path expected for each missing secret.
exec DebugMissingSecrets
What to do after a failure:
- Check the
sekoiaio-secret-operatorstatus:kubectl get pods -n support | grep secret-operator. - Run
DebugArgoCDSyncAllto restart the operator and trigger secret regeneration. - If specific Vault paths are missing, contact Sekoia support with the full command output.
DebugKustomizeStacksTemplates
Clones the ArgoCD Git repository and scans every YAML file for unrendered SH_TMPL template placeholders. For each match, it reports the file path, resource type, and the YAML field that was not substituted.
exec DebugKustomizeStacksTemplates
What to do after a failure:
- Each unrendered placeholder corresponds to a missing or incorrect value in
config.yml. - Correct the relevant parameter in
config.ymland re-runPushArgoStacksto regenerate the stack manifests.
Database diagnostics
DebugDatabases
Inspects all StatefulSets and CloudNativePG clusters in the support namespace.
exec DebugDatabases
Status definitions:
| Status | Meaning |
|---|---|
| Healthy | All replicas are running and ready with no recent restarts. |
| Warning | All replicas running, but recent restarts or waiting containers detected. Monitor and re-run. |
| Unhealthy | One or more replicas are not ready. Investigate immediately. |
What to do for Unhealthy status:
- For a pod in
CrashLoopBackOff:kubectl logs -n support <pod-name> --previous - For a pod in
Pendingstate:kubectl describe pod -n support <pod-name>and check for resource pressure or missing PVCs. - For CNPG clusters:
kubectl describe cluster -n support <cluster-name>
Resource management
DebugResourceAllocation
Queries the Kubernetes Metrics API and compares live memory consumption against declared memory requests for every pod.
exec DebugResourceAllocation
Metrics Server required
This command requires the Kubernetes Metrics Server. If the Metrics API is unavailable, the command exits with an error. Run the HelmInstall module to deploy it.
The output shows:
- A waste report for all pods with a memory request, sorted by wasted RAM. Red rows indicate 80% or more waste.
- A list of pods without memory requests, alongside their live usage. These pods present a scheduling risk.
Platform installer debug
DebugPlatformInstallation
Deploys the platform-installer Helm chart with a pause command override, creating a pod that stays alive without performing any changes. The command returns as soon as the Helm install succeeds; it does not wait for the pod to start.
Use this to open an interactive shell inside the installer container and inspect its runtime environment, mounted secrets, and configuration files.
exec DebugPlatformInstallation
Once the pod is running, open a shell with:
kubectl exec -it -n support <pod-name> -- /bin/bash
Any previous debug session is cleaned up automatically before the new pod is created.
Operations
RebootNodes
Reboots all nodes listed in utils.ansible.inventory and waits for them to come back online.
exec RebootNodes
Use this when patching the OS or applying kernel updates. The playbook waits for SSH to become available again on each node before reporting success.
KubeCrashRecovery
Restarts all pods across the cluster in ordered namespace phases and waits for each phase to become healthy before moving to the next. Use this after an unexpected node crash or manual cluster restart that left pods stuck in a failed state.
exec KubeCrashRecovery
The phases run in this order: kube* namespaces → rook-ceph → vault → *system* → support → all remaining namespaces. Each phase waits up to modules.kube_crash_recovery.pod_ready_timeout seconds (default: 300) before giving up.
What to do after a timeout: Run kubectl get pods -A --field-selector=status.phase!=Running,status.phase!=Succeeded to identify the stuck pods. Check their events and logs, then re-run KubeCrashRecovery.
WipeStorageDisks
Detects and wipes disks previously used by Ceph. Only runs when modules.wipe_storage.enabled is true.
Destructive and irreversible
This command permanently destroys all data on detected Ceph disks. Only run this when explicitly instructed by Sekoia, for example during a full cluster reinstallation.
exec WipeStorageDisks
Service diagnostics
Diagnostic
Evaluates rule-based health checks against the platform's Prometheus metrics and reports the state of each platform area.
exec Diagnostic
Use this command once the infrastructure and application layers are healthy but a platform feature misbehaves, for example alerts that stop being created or assets that stop being discovered. For the available targets, the output formats, and how to read a result, see Run platform diagnostics.
Collecting logs for a support ticket
When you escalate an issue to Sekoia L3 support, include the following in your request:
-
Copy the full terminal output of the failing command.
-
Share your
config.ymlwith all secrets redacted (replace all passwords and keys with***). -
Collect K3s system logs from the affected nodes:
journalctl -u k3s -n 500 --no-pager > k3s.log -
If applications are degraded, collect ArgoCD and pod logs:
exec DebugArgoCD kubectl get pods -A --field-selector=status.phase!=Running,status.phase!=Succeeded kubectl logs -n <namespace> <pod-name> --previous
Related links
- Monitor your platform: Continuous monitoring with Grafana and on-demand diagnostic workflows.
- Deploy the platform: Post-deployment validation commands.
- Deployment configuration reference: How to fix configuration errors flagged by
CheckLocalConfig. - Run platform diagnostics: Targeted Prometheus health checks per platform area.
- Use the SHC interface: Run commands and diagnostics from the interactive interface.