Changelog¶
Notable changes in the current release. History begins at the first public release below.
This project uses calendar versioning (vYY.MM.idx)
stamped at release: the month is when a version actually shipped, so the number
is not known until it ships. Work in progress lives under Unreleased, and
that heading is renamed to the version as part of cutting the release. Beta
builds publish from Unreleased — they are dated test builds and do not name a
future release.
[v26.09.1] — 2026-09-14¶
The first public release, and a public beta: pre-1.0, Apache 2.0. What it guarantees and what it deliberately does not are set out in the operating contract.
ClickHouse generates OpenTelemetry spans into system.opentelemetry_span_log
and gives you no way to ship them anywhere. Click-Dog is a single static binary
that reads those spans and exports them over OTLP gRPC to any OTEL-compatible
backend — Datadog, Honeycomb, Grafana, a collector of your own.
Added¶
Export¶
- Scheduled mode. A continuous polling loop over
system.opentelemetry_span_logon a configurable interval, with batching, chunking, and a per-cycle span cap (monitor.max_spans_per_cycle, default 1000, ceiling 100000). - Backfill mode. A one-time historical export of
system.query_logbetween two timestamps, emitting synthetic spans for queries that qualify, then exit. Backfill covers the query log only; it does not replaysystem.opentelemetry_span_log. - Live span selection. Duration filters run in ClickHouse to find qualifying
trace_ids, and spans for those traces are then fetched — so selection is driven by whole traces rather than by individual slow spans.monitor.max_spans_per_cyclebounds each cycle rather than discarding the remainder: a large trace continues across cycles, and spans the walk falls more than three lookback windows behind are dropped. - OTLP gRPC and Splunk HEC exporters, configurable independently, with TLS and mutual TLS on the OTLP path.
- Fan-out to multiple sinks. Backends start concurrently and each receives the full configured
monitor.export_timeout_s, so a slow sink cannot consume the deadline belonging to a healthy one. This build delivers under all-required semantics; any-success exists in the code but is not operator-selectable, so do not design around it during the beta. - Forward-progress guarantees. Each cycle continues through the lookback window oldest first from where the last successfully exported cycle stopped, so already-exported rows cannot repeatedly consume the cycle budget, and every span is read while inflow stays below
max_spans_per_cycle / check_interval_s. A cycle whose export fails does not advance, and the next successful cycle resumes from the last exported position up to three lookback windows back. - In-memory deduplication. An LRU cache (default 10,000 entries) suppresses repeats within a run. It is not persisted: restarts, retries, and leader-partition windows can resend spans carrying stable
(trace_id, span_id)identities, and how a backend treats those repeats is backend-specific.
Filtering and query-text privacy¶
- Two-tier duration filtering, at trace level and span level, applied in the ClickHouse query rather than in Go, to minimize data transfer.
- Operation whitelists with wildcards, IP whitelists, and regex query blacklists, plus query-length limits applied at fetch time and at export time.
- Four query-text modes.
filters.query_text_modeacceptsraw(the compatibility default),redacted,normalized_only, andnone, applied to scheduled native spans and query-log backfill alike. Privacy-restricting modes never fall back to raw text, including URL-encoded HTTPqueryparameters. A configuration carrying the olderfilters.redact_querieswith no mode set migrates to fail-closedredactedand warns at startup: unlike legacy redaction, a statement matching no rule is now omitted rather than exported. - Per-exporter
max_query_length, applied to livedb.statementvalues and backfill output alike.
Resilience and high availability¶
- Circuit breaker and adaptive backoff, layered: the breaker blocks cycles outright while backoff widens the interval.
- Leader election via ClickHouse Keeper, so a cluster-mode deployment (
use_cluster_queries: true) exports its whole-cluster read once rather than once per instance. Sidecars keep exporting their own node's spans and use leadership only for coordination. Instances create ephemeral znodes in Keeper to claim and hold leadership. - Topology self-audit. Detects the sidecar-plus-cluster anti-pattern that produces duplicate whole-cluster readers, and warns without ever gating readiness.
- Connection pooling, query timeouts, and startup health checks against every configured node.
Query analysis¶
click-dog analyze queriesproduces a deterministicanalysis.report.v1JSON contract oversystem.query_log— four analyzers, stable condition identities, no LLM and no query rewriting.- Baselines and regression detection.
-save-baselinecaptures an exact-hash known-good artifact;-baselinecompares against it non-mutatingly and emitsanalysis.report.v2. Conservativelatency_regressionandfailure_spikeanalyzers suppress low-volume, partially matched, stale, and incompatible comparisons. - Policy gates and notifications.
-fail-on criticalor-fail-on warningmakes the report return exit3at that threshold for CI use, and-notifysends a boundedanalysis.notification.v1summary to a Slack-compatible webhook and/or the Datadog Events API. click-dog analyze tracedrills from a query family into its traces, with a guided-wizardflow.
Operating¶
- Health probes —
/healthz,/readyz,/status, and leader-served/clusterzon a listener separate from metrics, so a scrape failure cannot look like a liveness failure. The listener is opt-in: sethealth.enabled(theproductionstarter profile does) and it defaults to:8686. A separate admin listener carriesPOST /flush, kept off the scrape endpoint so/metricscan be bound publicly without exposing it. - Self-metrics — a Prometheus scrape endpoint (also serving a legacy
/health) plus optional OTLP push, with per-sink export counters. Their state is reported at startup and byclick-dog validate, including an explicit disabled message, so a blank dashboard is diagnosable. - Three Datadog dashboards —
click-dog create-dashboardsprovisions application query analysis, exported user activity, and Click-Dog's own health. Activity counts are labeled as covering the qualified and exported stream, not as an audit record. - Outbound webhooks for lifecycle events, including a bounded synchronous delivery attempt on shutdown so process exit cannot race the final event.
Commands¶
click-dog initandinit --wizardgenerate a starter configuration;checkvalidates configuration and connectivity;validatevalidates configuration offline without connecting to anything;backfill --start/--endexports a historical time range one-shot;test exportandtest tracingsmoke-test an exporter and native trace propagation;analyze queriesandanalyze tracerun query analysis;deploy kubernetesanddeploy dockerrender manifests,deploy statusinspects a local installation, anddeploy uninstallremoves a standard systemd installation after confirmation;create-dashboardsprovisions Datadog;self-updatereplaces the binary in place after verifying the release signature chain described under Security.test-spanremains as a deprecated alias fortest export, and--dry-runreads real data, discards the exports, and reports what would have shipped.click-dog flushasks a running service to export now, and exports nothing itself. Without Keeper configured it signals a systemd-managed unit, so a click-dog running in Docker, Kubernetes, or the foreground is not reachable this way and the command reports that and exits non-zero. With Keeper configured it writes a request for the elected leader and returns success once the request is recorded, whether or not a leader picks it up.- Configuration is YAML with
${VAR}and$VARexpansion, validated byclick-dog validate, with logging at debug/info/warn/error to stderr or a rotating file.
Installation¶
install.shdownloads the release archive, verifies it against the signature chain described under Security before executing anything, and installs a systemd unit when asked (--systemd; the quickstart path always does). Plaininstall.shresolves the latest GA, so reaching a-beta.Nbuild needs--prerelease. Kubernetes manifests, Docker deployment, and an Ansible playbook for multi-node rollouts ship alongside.
Fixed¶
- A capped cycle skipped the newest traces. When a cycle continued a page the previous cycle had capped, it returned the older remainder without reading anything newer, so traces that finished between the two reads were never selected. At inflow the docs call sustainable (26–32 spans/s at the defaults) roughly a third of traces were lost. The reader now walks the window oldest first from a committed position and fills the rest of the cap with newer traces in the same cycle.
- A failed export was skipped instead of retried. The reader moved its position when it fetched a page, before the page was exported, so a collector error or export timeout on a capped page sent the next cycle past it. The position now moves only after every export in the cycle succeeds, and the next cycle resumes from it even when it has fallen behind the lookback window, up to three windows back, so one failed export followed by a backoff step loses nothing at the defaults. A strict IP or user whitelist that dropped spans because a lookup failed also leaves the page to be read again.
- Every other scheduled cycle was blind. The live span reader's keyset walk deferred its wrap to the next cycle and returned an empty page in between, so traces that arrived after a walk began were only read every second cycle. With the lookback sized as
check_interval_s + lookback_buffer_s, that was a rolling gap at the defaults; a ground-truth join against a 3-node cluster's span log measured roughly a fifth to a third of qualifying traces never exported. The walk now wraps within the same cycle, and a trace page shorter than the cap ends the walk outright. - A standby's startup cycle exported the leader's window. In cluster mode the election joined in the background while the startup cycle ran immediately, so the gate failed open and every standby start delivered one full lookback window of duplicates. A cluster reader now waits up to the Keeper session timeout for its election to join before its startup cycle; an unreachable Keeper still fails open after the wait.
- Promotion waited for the next tick. A promoted cluster reader now runs a cycle immediately, so the
now - lookbackre-read lands within the session timeout instead of up to one check interval later. click-dog checknamed the wrong grant in cluster mode. A monitoring user holding exactly the two documented per-table grants was told to grant them again; the remedy now namesGRANT REMOTE ON *.*, whichcluster()requires.- Canary queries against ClickHouse. Scan the native unsigned result of
count()before converting it to the signed public result type, preventing canary checks from failing on clickhouse-go's strict integer decoding. - Query text leaked through exception messages. ClickHouse quotes a failing statement, literals included, in the
clickhouse.exceptionattribute it records on the query span, andredacted,normalized_only, andnoneexported it unchanged. Every privacy-restricting mode now drops exception messages and keeps exception codes. - Exact IPv4 entries in
whitelist_ipsnever matched. ClickHouse reports an IPv4 client as::ffff:a.b.c.d, and exact entries such as10.0.1.50were compared as strings, so every span from that client was dropped. Entries now compare as addresses; CIDR entries were unaffected. click-dog --config x validatestarted the daemon. Flag parsing stopped at the verb and nothing checked what was left. Arguments after the top-level flags now exit2and say the command goes first.- The Ansible playbook failed as documented. Self-referencing defaults aborted the run with a recursive template error and overrode inventory settings, untagged updates ran the rollback tasks and restored the previous binary, and the unit required the TLS directory to exist. Defaults now yield to inventory and extra vars, rollback runs only with
--tags rollback, and the TLS path is optional.
Security¶
- Read-only against ClickHouse. Every ClickHouse SQL connection sets
readonly=2, so Click-Dog cannot write to the database it observes. The only writes to the observed data plane go to ClickHouse Keeper — not ClickHouse itself — for HA leader-election znodes and flush requests. The binary does write elsewhere:analyze queries --save-baselinestores a baseline file,create-dashboardscalls the Datadog API,self-updatereplaces the binary on disk, and the webhook notifier makes outbound POSTs. - Signed artifacts. Container images are cosign-signed directly. For archives, cosign signs
checksums.txtand each archive is bound to it by SHA-256 — that is the chaininstall.shandself-updateverify: cosign-verify the checksum list against the public release workflow identity, then match the downloaded archive's digest against it. Verifying an archive's signature directly will fail, because there isn't one. The documented exception is a binary you supply yourself:install.sh -band the Ansibleclick_dog_local_binaryvariable install a pre-verified binary and skip the download and cosign steps, so that path is yours to audit. - TLS throughout. Available for ClickHouse and OTLP connections, with mutual TLS for OTLP. Incomplete mTLS pairs and unreadable certificate files are rejected at configuration load rather than at first connection — a config setting only one of
client_cert/client_keywill not start — and a plaintext OTLP exporter warns at startup about SQL-literal exposure. - Built with Go 1.26.6, which the release gate enforces with
govulncheck. That toolchain clearsGO-2026-6218(net/url),GO-2026-6091(html/template),GO-2026-6090(crypto/tls),GO-2026-6089andGO-2026-5026(net/http), andGO-2026-5972(encoding/asn1) — advisories reachable through the health server, the Keeper TLS dialer, the self-updater's download and checksum verification, and the ClickHouse reader's certificate pool. - gRPC v1.82.1 on the OTLP export transport, clearing
GO-2026-6061, a reachable authorization/HTTP-2 advisory reported by the samegovulncheckgate. - Log files are created
0640. A log-shipping agent running under its own UID must share the click-dog service user's group. - Apache 2.0, with a DCO
Signed-off-bytrailer required on every contribution.
Deprecated¶
click-dog test-spanis a compatibility alias forclick-dog test export. Existing flags, stdout behavior, and exit status remain available; stderr names the replacement. No removal release is assigned.- The
-validateand-backfill-start/-backfill-endmode flags are compatibility aliases for thevalidateandbackfill --start/--endsubcommands. The flags keep working and now warn on stderr; the subcommands are the documented surface. No removal release is assigned.