Observability Guide
Use HyperDX to search logs, traces, and metrics from your install. Kindo ships ClickStack in the cluster: the OpenTelemetry Collector collects telemetry, ClickHouse stores it, and HyperDX provides search and dashboards.
Architecture
Section titled “Architecture”+-----------------------------------------------------------+| Kindo Services |+-----------+-----------+-----------+-----------+-----------+| API | Task | Credits | LiteLLM | Next.js || | Worker | | | |+-----+-----+-----+-----+-----+-----+-----+-----+-----+-----+ | | | | | +-----------+-----------+-----------+-----------+ | OTLP (traces, metrics, logs) +------------+ | +---------------+ | Agent | | | Logs DaemonSet| | Prometheus | | | pod stdout | +-----+------+ | +-------+-------+ | OTLP | | OTLP +-------------------+---------------+ | v +-------+-------+ | Gateway | | OTel Collector| +-------+-------+ | v +-------+-------+ | ClickHouse | +-------+-------+ ^ | queries +-------+-------+ | HyperDX (UI) | +---------------+To check collector health, inspect the otel-collector release in the kindo-monitoring namespace. Its workloads are:
- Gateway: receives telemetry and exports it to ClickHouse.
- Agent: collects Prometheus metrics and forwards them to the gateway.
- Logs DaemonSet: collects pod stdout from the namespaces configured under Log collection.
See Choose a size for storage and compute requirements.
Application telemetry
Section titled “Application telemetry”Kindo services send traces, metrics, and logs to the gateway automatically.
For custom applications in the cluster, configure OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector-gateway.kindo-monitoring:4317 and OTEL_EXPORTER_OTLP_PROTOCOL=grpc to send telemetry to ClickStack.
Log collection
Section titled “Log collection”Kindo application logs arrive through OTLP. The logs DaemonSet also collects pod stdout from the hatchet, nango, and ingress-nginx namespaces by default, excluding the linkerd-proxy sidecar container. Hatchet logs require this collection even when Hatchet traces and metrics are visible.
Use Helm overrides to enable collection or add namespaces:
-
Enable the logs DaemonSet if it is disabled:
Terminal window kindo config helm-override set otel-collector logs.enabled=truekindo config helm-override apply otel-collector -
Add a namespace for another component that logs to stdout. Helm replaces lists wholesale, so include the defaults along with your namespace:
Terminal window kindo config helm-override set otel-collector \'logs.namespaces[0]=hatchet' 'logs.namespaces[1]=nango' \'logs.namespaces[2]=ingress-nginx' 'logs.namespaces[3]=<your-namespace>'kindo config helm-override apply otel-collector -
Check HyperDX for logs from that namespace.
Collect stdout only for components that do not already send logs over OTLP, to avoid storing duplicate records. If a component’s logs are missing, check logs.enabled and logs.namespaces even when its traces or metrics appear.
Accessing HyperDX
Section titled “Accessing HyperDX”-
Open
https://hyperdx.<your-base-domain>. -
Read the initial admin credentials from the secrets config:
Terminal window kindo config editThe email is
clickstack.hyperdxAdminEmail, which defaults tohyperdx@<base domain>. The password isclickstack.hyperdxAdminPassword. Quit the editor without saving. -
Sign in. Change the password in the HyperDX UI when needed.
Default dashboards
Section titled “Default dashboards”Find the installed dashboards under Dashboards in HyperDX, tagged kindo-defaults:
- Agent Monitoring: Overview: run outcomes, success rate, and execution duration across agents.
- Agent Monitoring: Agent Run Details: one agent’s runs, and for a selected run, its steps and errors.
- Issue Inspector: why a chat or agent run went wrong. Paste a conversation, agent run, agent, organization, or user ID into its filters to see turn outcomes, each failed turn with what failed (the AI model or which tool) and its error, and integrations that couldn’t connect.
Start with Agent Monitoring: Overview and select an agent. On Agent Run Details, select a run to load it, then use the Run tab for its steps and the Errors tab for its errors.
Create your own dashboard with a new name for custom views. Default dashboards are overwritten on deploy and upgrade; dashboards outside the default set persist.
Signals and retention
Section titled “Signals and retention”Logs, traces, and metrics are enabled by default and stored in ClickHouse. Retention defaults to INTERVAL 7 DAY in the clickhouse release’s schema.tablesTtl setting.
To change retention:
-
Check disk capacity before increasing the retention period.
-
Set and apply the retention interval, for example:
Terminal window kindo config helm-override set clickhouse 'schema.tablesTtl=INTERVAL 14 DAY'kindo config helm-override apply clickhouse
ClickHouse stores logs, traces, and metrics on the same persistent volumes. Watch disk usage on the ClickHouse PVCs. When free space drops to the chart’s insert floor (10% of the volume), ClickHouse rejects new telemetry until retention frees space. Telemetry is dropped in the meantime. Lower schema.tablesTtl with the commands above or expand the volumes. Volumes can grow but cannot shrink. See Choose a size for sizing.
During a fresh install, collector pods wait in init until the ClickHouse schema migration finishes. If they stay there, check the migration Job and ClickHouse pod health.
Troubleshooting
Section titled “Troubleshooting”| Symptom | Cause | Solution |
|---|---|---|
| No traces/metrics in HyperDX | Application telemetry is disabled or missing | Confirm the collector is enabled and the service has its OTLP endpoint |
| No logs from an OTLP service | Gateway is not ingesting | Check that gateway pods are Running and have completed initialization |
| No logs from Hatchet, Nango, or another tailed component | Logs DaemonSet is disabled or its namespace is missing | Set logs.enabled and logs.namespaces with kindo config helm-override set otel-collector, then run kindo config helm-override apply otel-collector; see Log collection |
| Collector pods stuck in init | ClickHouse schema migration has not finished | Check the migration Job and ClickHouse pod health |
| Custom app telemetry missing | Wrong endpoint or protocol | Use the gateway OTLP endpoint above with grpc |
| ClickHouse disk filling | Retained telemetry exceeds disk headroom | Lower schema.tablesTtl with kindo config helm-override set clickhouse, then run kindo config helm-override apply clickhouse, or expand the persistent volumes |
| Long log fields appear truncated | Large string fields in application logs are truncated before export, with a marker showing the original size | Expected; no action needed. |
