Install on Kubernetes¶
Kubernetes needs two pieces:
- ClickHouse must write
system.opentelemetry_span_logand expose a read-only monitoring user. - Click-Dog must run in exactly one supported topology: a centralized Deployment that reads the cluster, or one same-pod sidecar per ClickHouse pod.
The shipped generator creates the centralized Deployment, which is the usual Kubernetes starting point. The same-pod sidecar is available for very large clusters or environments that require node-local reads.
Prerequisites¶
- a Kubernetes cluster and
kubectl - a ClickHouse Service reachable from the Click-Dog namespace
- permission to configure ClickHouse and create its monitoring user
- an OTLP/gRPC receiver address
- either a local
click-dogbinary, or Docker for the installer-wrapper generator
If ClickHouse is managed by the Altinity operator, install its
ClickHouseInstallation CRD before applying the example below:
kubectl apply -f https://raw.githubusercontent.com/Altinity/clickhouse-operator/master/deploy/operator/clickhouse-operator-install-bundle.yaml
1. Enable the ClickHouse span log¶
Click-Dog only reads telemetry; it cannot make ClickHouse emit it. Every
ClickHouse pod that should contribute traces must define
system.opentelemetry_span_log.
With the Altinity operator, add this to the ClickHouseInstallation:
apiVersion: "clickhouse.altinity.com/v1"
kind: "ClickHouseInstallation"
metadata:
name: clickhouse
namespace: clickhouse
spec:
configuration:
files:
config.d/opentelemetry.xml: |
<clickhouse>
<opentelemetry_span_log>
<engine>
ENGINE = MergeTree
PARTITION BY toYYYYMM(finish_date)
ORDER BY (finish_date, finish_time_us, trace_id)
</engine>
<database>system</database>
<table>opentelemetry_span_log</table>
<flush_interval_milliseconds>7500</flush_interval_milliseconds>
</opentelemetry_span_log>
</clickhouse>
The operator puts the file on every pod and performs the required rolling restart. ClickHouse only starts writing the table after restart.
Span generation must also be sampled. ClickHouse emits spans for queries with a
parent trace context, or you can temporarily set
opentelemetry_start_trace_probability to 1 while validating the pipeline
and lower it afterward.
The complete checked-in example is
deploy/kubernetes/clickhouse-operator-spanlog.yaml.
2. Create the read-only monitoring user¶
Click-Dog needs SELECT on two system tables and enforces readonly=2 on every
connection. With the Altinity operator, extend the same installation:
spec:
configuration:
profiles:
click_dog_readonly/readonly: 2
users:
click_dog_monitor/networks/ip:
- "::/0" # tighten to the Click-Dog pod CIDR in production
click_dog_monitor/password_sha256_hex: REPLACE_WITH_SHA256_HEX_OF_PASSWORD
click_dog_monitor/profile: click_dog_readonly
click_dog_monitor/quota: default
click_dog_monitor/grants/query:
- "GRANT SELECT ON system.opentelemetry_span_log TO click_dog_monitor"
- "GRANT SELECT ON system.query_log TO click_dog_monitor"
# The centralized Deployment reads through cluster(), a table function
# ClickHouse gates behind REMOTE; system.columns feeds the replica
# capability probes. A same-pod sidecar needs only the two SELECTs.
- "GRANT REMOTE ON *.* TO click_dog_monitor"
- "GRANT SELECT ON system.columns TO click_dog_monitor"
Generate the password hash while keeping the plaintext for the Click-Dog Secret:
If you do not use the operator, execute equivalent SQL through your existing secured ClickHouse administration path:
CREATE USER IF NOT EXISTS click_dog_monitor IDENTIFIED WITH sha256_hash BY '<hash>';
GRANT SELECT ON system.opentelemetry_span_log TO click_dog_monitor;
GRANT SELECT ON system.query_log TO click_dog_monitor;
-- Centralized Deployment (cluster() reads) only; a same-pod sidecar needs just the two SELECTs.
GRANT REMOTE ON *.* TO click_dog_monitor;
GRANT SELECT ON system.columns TO click_dog_monitor;
Run the grants on every ClickHouse node (ON CLUSTER where the user is
managed through SQL). Without REMOTE the centralized Deployment fails every
read with Not enough privileges ... READ ON REMOTE, and click-dog check
names the missing grant. See
Configuration · Cluster Mode for the full set.
Administrator credentials are needed only to create the user. Never put them in the Click-Dog runtime Secret.
3. Choose the topology¶
Centralized Deployment¶
One Click-Dog Deployment reaches ClickHouse through its Service and uses
ClickHouse cluster() queries to read every shard:
clickhouse:
host: clickhouse.clickhouse.svc.cluster.local
port: 9000
use_cluster_queries: true
cluster: main
use_cluster_queries: true requires a non-empty cluster name. Validation fails
instead of silently falling back to one Service-routed pod. A single-node
ClickHouse installation can omit --cluster; the generator warns that the
Deployment will then read only the pod selected by the Service.
Same-pod sidecar¶
One Click-Dog container runs inside each ClickHouse pod, connects to
localhost, and keeps use_cluster_queries: false. Because it shares the
ClickHouse pod's network namespace, it reads only that pod's local span log.
Do not combine sidecars with cluster queries
A sidecar on every pod with use_cluster_queries: true makes every instance
read every shard, multiplying exports and cluster load. Use either one
centralized cluster reader or node-local sidecars.
A standalone Deployment or DaemonSet cannot use localhost: that address
points back to the Click-Dog container's own pod. The Kubernetes generator
therefore rejects loopback --ch-host values.
4. Generate a centralized Deployment¶
Export the monitoring password and point the generator at the ClickHouse Service and OTLP receiver:
export CLICKHOUSE_PASSWORD="your-monitoring-password"
click-dog deploy kubernetes \
-c otel-collector.monitoring:4317 \
--ch-host clickhouse.clickhouse.svc.cluster.local \
--cluster main
On a fresh workstation, the release installer can run the same generator from the published image:
curl -fsSL https://github.com/coltconsulting/click-dog/releases/latest/download/install.sh -o install.sh
export CLICKHOUSE_PASSWORD="your-monitoring-password"
bash install.sh kubernetes \
-c otel-collector.monitoring:4317 \
--ch-host clickhouse.clickhouse.svc.cluster.local \
--cluster main
The wrapper requires Docker but generates files only; it does not install a host binary or systemd service.
Both paths create click-dog-k8s/ containing a Namespace, ServiceAccount,
ConfigMap, Secret, single-replica Deployment, .gitignore, and
kustomization.yaml. The Secret contains credentials and is written locally
with mode 0600; do not commit it.
Review the generated YAML, then apply it:
The ConfigMap enables health on :8686 and Prometheus metrics on :9090. The
Deployment probes /healthz for liveness and /readyz for readiness.
5. Verify¶
Check the workload, then run Click-Dog's own data-path check:
kubectl get pods -n click-dog -o wide
kubectl logs -n click-dog deploy/click-dog --tail=50
kubectl exec -n click-dog deploy/click-dog -- \
click-dog check --config /etc/click-dog/click-dog.yaml
Confirm ClickHouse is writing spans and the monitoring user can read them:
EXISTS TABLE system.opentelemetry_span_log;
SELECT count()
FROM system.opentelemetry_span_log
WHERE finish_date >= today();
A non-zero count after traced queries confirms the source is filling. If it stays at zero, verify sampling and confirm the ClickHouse pods restarted after the span-log configuration landed.
Update¶
Resolve a stable release, update only the generated image reference, then reapply:
LATEST_URL="$(curl -fsSL -o /dev/null -w '%{url_effective}' \
https://github.com/coltconsulting/click-dog/releases/latest)" || exit 1
VERSION="${LATEST_URL##*/v}"
click-dog deploy kubernetes --update -v "$VERSION" -o click-dog-k8s
kubectl apply -k click-dog-k8s/
The command creates deployment.yaml.bak and changes only the image tag. It
does not rewrite the ConfigMap or Secret. The equivalent installer-wrapper
command is:
See Updating for the shared config-default and prerelease contract.
Migrating an old DaemonSet output directory
Output generated before the centralized Deployment change contains
daemonset.yaml. The update command refuses it because changing an image
tag cannot migrate topology. Regenerate the directory with the current
generator, apply it, then delete the old DaemonSet—kubectl apply -k does
not prune the obsolete workload:
Same-pod sidecar example¶
When ClickHouse runs as a StatefulSet or operator-managed pod, add a Click-Dog container to that pod template:
- name: click-dog
image: ghcr.io/coltconsulting/click-dog:${VERSION}
args: ["--config", "/etc/click-dog/click-dog.yaml"]
env:
- name: CLICKHOUSE_USERNAME
valueFrom:
secretKeyRef:
name: click-dog-secret
key: CLICKHOUSE_USERNAME
- name: CLICKHOUSE_PASSWORD
valueFrom:
secretKeyRef:
name: click-dog-secret
key: CLICKHOUSE_PASSWORD
resources:
requests: { cpu: 50m, memory: 64Mi }
limits: { cpu: 200m, memory: 128Mi }
volumeMounts:
- name: click-dog-config
mountPath: /etc/click-dog
readOnly: true
The mounted config must use clickhouse.host: localhost and
use_cluster_queries: false. Pin ${VERSION} to the bare release version in
the rendered manifest rather than using latest in production.
TLS with cert-manager¶
The repository includes a certificate example for environments using cert-manager:
kubectl apply -f https://github.com/cert-manager/cert-manager/releases/latest/download/cert-manager.yaml
vi deploy/kubernetes/certificate.yaml
kubectl apply -f deploy/kubernetes/certificate.yaml
Set the certificate identity and issuer for your environment, mount the resulting Secret, then enable the matching TLS fields in the Click-Dog config. See Configuration: TLS.
Managed ClickHouse¶
ClickHouse Cloud and other managed offerings often do not permit config.d/
files, so the span-log settings above do not apply. Sampling has to be set on
the application's user or role instead, where the provider allows it.
On providers that expose no span log at all, the live path has nothing to read; the historical query-log backfill path may still be available, as a separate source and mode.
Check the provider's current settings and system-table support before deploying. Where SQL user management is available, keep the monitoring grants restricted to whichever required system tables the service exposes.
ClickHouse Cloud is a separate case. It does expose
system.opentelemetry_span_log, and is not supported yet for a different
reason. See ClickHouse Cloud.
Remove¶
Remove the generated Kubernetes resources with their orchestrator:
Retain or remove the local generated directory and Secret according to your
configuration-management policy. click-dog deploy uninstall is systemd-only
and deliberately does not alter Kubernetes resources.