Skip to content

Observability Guide

Use HyperDX to search logs, traces, and metrics from your install. Kindo ships ClickStack in the cluster: the OpenTelemetry Collector collects telemetry, ClickHouse stores it, and HyperDX provides search and dashboards.

+-----------------------------------------------------------+
| Kindo Services |
+-----------+-----------+-----------+-----------+-----------+
| API | Task | Credits | LiteLLM | Next.js |
| | Worker | | | |
+-----+-----+-----+-----+-----+-----+-----+-----+-----+-----+
| | | | |
+-----------+-----------+-----------+-----------+
| OTLP (traces, metrics, logs)
+------------+ | +---------------+
| Agent | | | Logs DaemonSet|
| Prometheus | | | pod stdout |
+-----+------+ | +-------+-------+
| OTLP | | OTLP
+-------------------+---------------+
|
v
+-------+-------+
| Gateway |
| OTel Collector|
+-------+-------+
|
v
+-------+-------+
| ClickHouse |
+-------+-------+
^
| queries
+-------+-------+
| HyperDX (UI) |
+---------------+

To check collector health, inspect the otel-collector release in the kindo-monitoring namespace. Its workloads are:

  • Gateway: receives telemetry and exports it to ClickHouse.
  • Agent: collects Prometheus metrics and forwards them to the gateway.
  • Logs DaemonSet: collects pod stdout from the namespaces configured under Log collection.

See Choose a size for storage and compute requirements.

Kindo services send traces, metrics, and logs to the gateway automatically.

For custom applications in the cluster, configure OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector-gateway.kindo-monitoring:4317 and OTEL_EXPORTER_OTLP_PROTOCOL=grpc to send telemetry to ClickStack.

Kindo application logs arrive through OTLP. The logs DaemonSet also collects pod stdout from the hatchet, nango, and ingress-nginx namespaces by default, excluding the linkerd-proxy sidecar container. Hatchet logs require this collection even when Hatchet traces and metrics are visible.

Use Helm overrides to enable collection or add namespaces:

  1. Enable the logs DaemonSet if it is disabled:

    Terminal window
    kindo config helm-override set otel-collector logs.enabled=true
    kindo config helm-override apply otel-collector
  2. Add a namespace for another component that logs to stdout. Helm replaces lists wholesale, so include the defaults along with your namespace:

    Terminal window
    kindo config helm-override set otel-collector \
    'logs.namespaces[0]=hatchet' 'logs.namespaces[1]=nango' \
    'logs.namespaces[2]=ingress-nginx' 'logs.namespaces[3]=<your-namespace>'
    kindo config helm-override apply otel-collector
  3. Check HyperDX for logs from that namespace.

Collect stdout only for components that do not already send logs over OTLP, to avoid storing duplicate records. If a component’s logs are missing, check logs.enabled and logs.namespaces even when its traces or metrics appear.

  1. Open https://hyperdx.<your-base-domain>.

  2. Read the initial admin credentials from the secrets config:

    Terminal window
    kindo config edit

    The email is clickstack.hyperdxAdminEmail, which defaults to hyperdx@<base domain>. The password is clickstack.hyperdxAdminPassword. Quit the editor without saving.

  3. Sign in. Change the password in the HyperDX UI when needed.

Find the installed dashboards under Dashboards in HyperDX, tagged kindo-defaults:

  • Agent Monitoring: Overview: run outcomes, success rate, and execution duration across agents.
  • Agent Monitoring: Agent Run Details: one agent’s runs, and for a selected run, its steps and errors.
  • Issue Inspector: why a chat or agent run went wrong. Paste a conversation, agent run, agent, organization, or user ID into its filters to see turn outcomes, each failed turn with what failed (the AI model or which tool) and its error, and integrations that couldn’t connect.

Start with Agent Monitoring: Overview and select an agent. On Agent Run Details, select a run to load it, then use the Run tab for its steps and the Errors tab for its errors.

Create your own dashboard with a new name for custom views. Default dashboards are overwritten on deploy and upgrade; dashboards outside the default set persist.

Logs, traces, and metrics are enabled by default and stored in ClickHouse. Retention defaults to INTERVAL 7 DAY in the clickhouse release’s schema.tablesTtl setting.

To change retention:

  1. Check disk capacity before increasing the retention period.

  2. Set and apply the retention interval, for example:

    Terminal window
    kindo config helm-override set clickhouse 'schema.tablesTtl=INTERVAL 14 DAY'
    kindo config helm-override apply clickhouse

ClickHouse stores logs, traces, and metrics on the same persistent volumes. Watch disk usage on the ClickHouse PVCs. When free space drops to the chart’s insert floor (10% of the volume), ClickHouse rejects new telemetry until retention frees space. Telemetry is dropped in the meantime. Lower schema.tablesTtl with the commands above or expand the volumes. Volumes can grow but cannot shrink. See Choose a size for sizing.

During a fresh install, collector pods wait in init until the ClickHouse schema migration finishes. If they stay there, check the migration Job and ClickHouse pod health.

SymptomCauseSolution
No traces/metrics in HyperDXApplication telemetry is disabled or missingConfirm the collector is enabled and the service has its OTLP endpoint
No logs from an OTLP serviceGateway is not ingestingCheck that gateway pods are Running and have completed initialization
No logs from Hatchet, Nango, or another tailed componentLogs DaemonSet is disabled or its namespace is missingSet logs.enabled and logs.namespaces with kindo config helm-override set otel-collector, then run kindo config helm-override apply otel-collector; see Log collection
Collector pods stuck in initClickHouse schema migration has not finishedCheck the migration Job and ClickHouse pod health
Custom app telemetry missingWrong endpoint or protocolUse the gateway OTLP endpoint above with grpc
ClickHouse disk fillingRetained telemetry exceeds disk headroomLower schema.tablesTtl with kindo config helm-override set clickhouse, then run kindo config helm-override apply clickhouse, or expand the persistent volumes
Long log fields appear truncatedLarge string fields in application logs are truncated before export, with a marker showing the original sizeExpected; no action needed.