Configuration Reference¶
Click-Dog uses YAML configuration with environment variable support. When
--config is omitted, it first looks for
/etc/click-dog/click-dog.yaml, then falls back to ./click-dog.yaml in the
current directory. Override with --config:
Unknown keys are rejected at load time, so a typo like check_intervla_s fails fast instead of being silently ignored. Use click-dog validate to check that your config parses, passes validation, and references readable files for active TLS settings without starting the service or making network connections.
Configuration fragments
Unless a YAML block is explicitly labeled as complete, it is a fragment to merge into a standalone configuration. In particular, an exporter-only or observability-only block does not include the required scheduled-monitor threshold. Start from the complete minimal profile, then merge the focused fragment into it.
Validate and run¶
For a manual foreground run, save your configuration as click-dog.yaml and
replace the example addresses, credentials, and any TLS file paths. Complete
templates are available in the minimal profile
and production profile.
If your config uses ${CLICKHOUSE_PASSWORD}, set it in the same shell first:
Validate the file offline, without opening network connections:
Then preview one polling cycle against your real ClickHouse. The dry-run discards span exports, prints a summary, and exits. Enabled self-metrics and startup webhooks can still send traffic; see dry-run behavior.
When the result looks right, start continuous polling and export:
Press Ctrl+C to stop the foreground process. For an installed service, follow the service setup instructions instead of starting a second instance.
Validation and migrations¶
Run the offline validator before a rollout:
It checks YAML shape, configuration relationships, and readability of
certificate files used by active ClickHouse/OTLP TLS connections. It does not
connect to ClickHouse or an exporter. click-dog check adds network and
data-plane checks; see what the HEC check proves
for that exporter's intentionally limited result.
The load-bearing validation rules are:
| Area | Contract |
|---|---|
| Standalone minimum | Scheduled mode defaults on. When monitor.enabled: true, both monitor.min_trace_duration_ms and monitor.check_interval_s must be greater than zero. At least one entry under exporters.otel or exporters.splunk_hec is also required. Consequently, omitting monitor: is not a valid scheduled-mode config. Set monitor.enabled: false only for a deliberately backfill-only config. |
| ClickHouse addressing | clickhouse.port must be 1–65535. When present, clickhouse.cluster must start with an ASCII letter and then contain only letters, digits, _, or -; it is required when use_cluster_queries: true. |
| Exporter entries | Every exporters.otel[] entry requires a nonblank collector_address. Every exporters.splunk_hec[] entry requires an endpoint and a nonempty token (inline or from token_file); endpoint scheme rules are described under Multiple Export Backends. |
| OTLP self-metrics | When metrics.otlp.enabled: true, interval_seconds must be positive. Inheritance requires at least one exporters.otel entry. With inherit_otel_connection: false, metrics.otlp.collector_address is required. |
| User filters | filters.whitelist_users and filters.blacklist_users require monitor.enrich_from_query_log: true, because the span log does not provide the originating user by itself. |
| Query-text privacy | filters.query_text_mode accepts raw, redacted, normalized_only, or none. redacted requires a valid rule; the ambiguous normalized spelling is rejected. Rules under raw are rejected, while inert rules under normalized_only/none produce a warning. |
| Topology audit | Negative audit timing/count values are rejected. When enabled, the effective interval_s, debounce_count, and query_log_lookback_minutes are positive; a YAML value of 0 is normalized to its default (300, 2, and 15, respectively). |
| Cluster health | health.cluster.enabled: true requires health.enabled: true and a routable health.cluster.self. self and every peer must be bare host:port (bracket IPv6), without a scheme, path, query, or fragment. Empty and duplicate peer entries are rejected after trimming whitespace. |
| TLS coherence | clickhouse.insecure_skip_verify and exporters.otel[].insecure_skip_verify require their corresponding secure: true; HEC insecure_skip_verify requires an https:// endpoint. An OTLP client certificate and key must be supplied together. TLS files used by active ClickHouse/OTLP connections must be readable at load time. Inactive CA or complete OTLP client-certificate material is accepted but emits a warning, as documented in the relevant TLS tables below. |
filters.redact_queries[].pattern expressions are compiled by configuration
validation. filters.blacklist_queries has a different lifecycle: its regexes
are compiled when the daemon or backfill query filter is initialized, not by
click-dog validate or click-dog check. An invalid blacklist expression can therefore
pass both checks and then stop daemon/backfill startup with Failed to initialize
query filter. See Filtering · Regex syntax.
Removed keys are errors, not compatibility aliases:
- Replace the removed top-level
otel:mapping with a list entry underexporters.otel:. The strict YAML decoder reports the old key as unknown; click-dog does not migrate it automatically. - Delete the removed
ha.enabledkey. Configureha.keeper.hoststo enable election, or omit the hosts to run standalone. This key receives a targeted removal error rather than the generic unknown-field error.
There are currently no legacy keys that load with only a deprecation warning.
Unknown and removed keys fail closed so a migration cannot be silently ignored.
For query-text migration only, a config with legacy redact_queries rules and
no query_text_mode resolves to redacted and emits a validation warning;
this preserves the privacy intent without silently returning to raw mode. The
new mode is stricter than legacy redaction: unmatched statements are omitted,
and the migration warning states that data-visibility change explicitly.
Environment Variable Expansion¶
Selected string fields support ${VAR} and $VAR syntax. Unset variables
resolve to empty strings and produce a startup warning when the braced form was
used.
clickhouse:
host: ${CLICKHOUSE_HOST}
username: ${CLICKHOUSE_USERNAME}
password: ${CLICKHOUSE_PASSWORD}
exporters:
otel:
- collector_address: ${OTEL_COLLECTOR_ADDRESS}
service_name: ${OTEL_SERVICE_NAME}
ca_cert: ${OTEL_CA_CERT}
client_cert: ${OTEL_CLIENT_CERT}
client_key: ${OTEL_CLIENT_KEY}
The following fields support environment variable expansion:
- clickhouse.host, clickhouse.username, clickhouse.password, clickhouse.password_file, clickhouse.ca_cert
- exporters.otel[].collector_address, exporters.otel[].service_name, exporters.otel[].ca_cert, exporters.otel[].client_cert, exporters.otel[].client_key
- exporters.splunk_hec[].endpoint, exporters.splunk_hec[].token, exporters.splunk_hec[].token_file
- metrics.otlp.collector_address, metrics.otlp.host, metrics.otlp.service_name
- ha.keeper.auth_user, ha.keeper.auth_password, ha.keeper.auth_password_file
- webhook.url
- datadog_events.site, datadog_events.api_key,
datadog_events.application_key, datadog_events.environment, and
datadog_events.service
The *_file secret-path fields (clickhouse.password_file,
exporters.splunk_hec[].token_file, ha.keeper.auth_password_file,
datadog_events.api_key_file, and datadog_events.application_key_file) are
also expanded, so the path can reference an environment variable. For example:
password_file: ${SECRET_PATH}.
ClickHouse Connection¶
clickhouse:
host: localhost # ClickHouse hostname
port: 9000 # Native protocol port (default: 9000)
database: default # Database name
username: default # Username (supports env vars)
password: ${CLICKHOUSE_PASSWORD} # Password (supports env vars)
# TLS
secure: false # Enable TLS (default: false)
insecure_skip_verify: false # Skip cert verification (TESTING ONLY)
ca_cert: "" # Path to CA cert for self-signed certs
# Cluster
cluster: "" # Cluster name for distributed queries
use_cluster_queries: false # Use cluster() function (default: false)
# Connection pool
max_open_conns: 2 # Max open connections (default: 2)
max_idle_conns: 1 # Max idle connections (default: 1)
query_timeout_s: 30 # Query timeout in seconds (default: 30)
# Memory
max_memory_usage: 104857600 # Per-query memory limit in bytes (default: 100MB)
Connection Details¶
| Configuration path | Default | Description |
|---|---|---|
clickhouse.host |
None | ClickHouse server hostname |
clickhouse.port |
9000 |
Native protocol port; must be between 1 and 65535 |
clickhouse.database |
None | Database name |
clickhouse.username |
None | Username |
clickhouse.password |
None | Password (supports ${ENV} expansion). Mutually exclusive with password_file |
clickhouse.password_file |
None | Path to a file holding the password (read at load; one trailing CRLF or LF is trimmed). Path supports ${VAR} expansion. Keeps the secret out of the process environment. Mutually exclusive with password. An empty file is rejected; use inline password: "" for an intentionally empty password |
TLS¶
| Configuration path | Default | Description |
|---|---|---|
clickhouse.secure |
false |
Enable TLS encryption. When false, a configured ca_cert is ignored and produces a validation warning |
clickhouse.insecure_skip_verify |
false |
Skip certificate verification. Insecure. Use only for testing. Requires secure: true; otherwise validation fails |
clickhouse.ca_cert |
None | Path to a CA certificate file for TLS verification with self-signed or internal CA certs. Supports env expansion and must be readable at config load when secure: true |
Cluster Mode¶
Click-Dog supports two reader topologies. A single deployment is a sidecar
deployment with one instance. The config is identical; only the instance count
differs.
Sidecar (default, recommended): Deploy one Click-Dog instance per ClickHouse
node. Each instance reads only its local system.opentelemetry_span_log;
the partitions are disjoint by host, so the union gives full cluster coverage
with zero coordination. This is the topology that distributes the expensive
cluster-wide read. Each instance has minimal permissions and generates minimal
load. No Keeper, no leader election. Deploy the same config once per node.
Cluster (use_cluster_queries: true): A Click-Dog instance reads the
whole cluster by wrapping the span query as
cluster('cluster_name', system.opentelemetry_span_log) (one replica per shard).
Requires network access to all nodes and a non-empty cluster: name. Cluster
mode is always leader-gated (see below): run 1 instance (the leader of
an election of one) or 2–3 for automatic failover. Adding instances doesn't
divvy up the read: they stand by behind one active exporter.
Grants. cluster() is a table function, and ClickHouse gates it behind the
REMOTE privilege on top of the per-table SELECT grants, so the monitoring
user needs, on every node:
GRANT SELECT ON system.opentelemetry_span_log TO click_dog_monitor;
GRANT SELECT ON system.query_log TO click_dog_monitor;
GRANT REMOTE ON *.* TO click_dog_monitor; -- cluster() / clusterAllReplicas()
GRANT SELECT ON system.columns TO click_dog_monitor; -- replica capability probes
GRANT SELECT ON system.processes TO click_dog_monitor; -- only for `analyze trace --source current`
Without REMOTE, every read fails with Not enough privileges ... READ ON
REMOTE and click-dog check fails the span_log readable probe naming the
grant. Without system.columns, the startup probes that verify
normalized_query_hash and query_kind on every replica cannot run, so they
warn and the normalized query attributes and operation enrichment stay off
(click_dog_normalized_query_supported and
click_dog_query_operation_supported read 0). Grant ON CLUSTER so every
node the reader can connect to agrees. Which identity the remote shard uses
for the read depends on the cluster definition: without an inter-server
<secret>, ClickHouse connects to the other shards as the <user> named in
remote_servers (commonly default), not as the monitoring user.
A shard that is down fails the whole read. cluster() runs without
skip_unavailable_shards, so while any shard is unreachable every cycle errors
(Error fetching spans: ... All connection tries failed), adaptive backoff
grows, the circuit breaker opens after failure_threshold consecutive failures,
and /readyz reports 503 until the shard returns. Nothing is exported for the
reachable shards in the meantime, and the lookback window bounds what is
recovered afterwards. This is deliberate: a partial read would look complete
while silently missing that shard's spans. Sidecars on the surviving nodes are
unaffected, which is one reason sidecar is the recommended topology for larger
clusters.
One distributed query is not necessarily one trace. Whether the shard work
shares the initiator's trace_id is ClickHouse's context propagation, not
Click-Dog's. A trace started by opentelemetry_start_trace_probability is not
propagated to remote shards (verified on ClickHouse 25.8): each shard records
its part as a separate trace, and each qualifies for export on its own
duration, so one slow distributed query can arrive as 1 + shards traces. A
trace started from a client-supplied traceparent header is propagated, and
every shard's spans land in one trace, each tagged with its own hostname.
| Configuration path | Default | Description |
|---|---|---|
clickhouse.cluster |
None | Cluster name. Required when use_cluster_queries is true; must match [A-Za-z][A-Za-z0-9_-]* because it is interpolated into cluster() queries |
clickhouse.use_cluster_queries |
false |
When true, uses cluster('name', table) to read one replica per shard across the cluster (cluster topology). When false, reads only this node's local tables (sidecar topology). Optional normalized query attributes are enabled in cluster query mode only after clusterAllReplicas('name', system.columns) verifies every replica exposes normalized_query_hash; mixed or unknown clusters keep the safe fallback |
Cluster mode is always leader-gated¶
In cluster mode the fetch/export cycle runs only on the election leader; other instances stand by (recorded as skipped cycles, so a standby's last-success gauge legitimately stays at zero). The rule is export iff no election is active, or this instance is the leader, so:
- A single cluster reader with no Keeper configured is valid and common:
no election starts, so it always exports (no regression for existing
use_cluster_queries: trueusers). - With Keeper (2–3 instances), one is elected leader and exports; the rest
defer. On leader loss, promotion normally takes about the configured Keeper
session timeout (10s default; valid range 5-30s), plus election overhead. The
new leader runs a cycle as soon as it is promoted rather than waiting for its
next tick, re-reading from
now - lookback. When that lookback covers the full failover interval (the defaults do: 40s against roughly 10s), the boundary is duplicate-prone overlap; a shorter explicit lookback can let spans age out before promotion. - A cluster reader waits up to the session timeout for its election to join before running its startup cycle, so an instance that starts as a standby under a healthy Keeper exports nothing and its last-success gauge stays at zero. If Keeper is unreachable the wait times out and the startup cycle runs fail-open.
- Leader election is driven solely by the presence of
ha.keeper.hosts. There is no separate enable flag. Configure Keeper hosts and instances coordinate; omit them and each instance runs standalone. Only the cluster-query data path is leader-gated. Sidecars always export their disjoint node-local scopes; Keeper leadership on a sidecar is used only for coordination duties such as/clusterzand shared flush requests.
The delivery contract is best-effort with at-least-once retries, not
exactly-once. Click-Dog preserves (trace_id, span_id), which lets operators
identify repeat deliveries, but OTLP does not require collectors or backends to
collapse them. Under a healthy Keeper, cluster mode guarantees no steady-state
duplication, not "never duplicates." Every Keeper disruption after an
instance has joined fails open: a session loss, reconnect, partition, or watch error
drops candidate state so the instance keeps exporting (rather than stalling as a
gated standby), and the election goroutine retries until it rejoins. The
resulting duplicate window is therefore bounded and recovers to a single
exporter when Keeper returns. An election-constructor failure is the
exception: host parsing or DNS resolution, or an initial authentication
failure, leaves that process standalone until restart. Click-Dog does not
fail closed: it prefers availability.
Pick one topology per fleet
Don't run per-node sidecars with use_cluster_queries: true. Leader-gating
now prevents steady-state duplication when those instances share an
election, but sidecars that are not sharing an election (different
base_path, or no Keeper) would each go cluster-wide and duplicate. If you
want whole-cluster reads, run cluster mode with a shared Keeper; if you
want one reader per node, leave use_cluster_queries off. See
Operating · Cluster topology.
This is enforced at runtime. When use_cluster_queries: true,
click-dog runs an in-process topology self-audit that detects more than
one instance doing whole-cluster reads and surfaces it via the
click_dog_topology_warning metric, the /status topology_warning field,
a throttled WARN log, and the Datadog health dashboard. It is
observability-only: it never gates readiness or de-rotates a node and is a
backstop for the case leader-gating can't reach (instances not sharing an
election). See monitor.topology_audit below.
Topology self-audit¶
When clickhouse.use_cluster_queries: true, click-dog periodically checks
whether more than one instance is running whole-cluster span reads: the sidecar
+ use_cluster_queries anti-pattern that causes N× duplicate exports. It is a
hard no-op unless this instance itself uses cluster queries (zero cost for
the default per-node sidecar topology), and is observability-only: it emits
the click_dog_topology_warning{reason} gauge, the /status topology_warning
field, and a throttled WARN, but never affects /readyz or /healthz.
| Configuration path | Default | Description |
|---|---|---|
monitor.topology_audit.enabled |
true |
Enable the self-audit (still a no-op unless use_cluster_queries is on) |
monitor.topology_audit.interval_s |
300 |
Audit tick interval in seconds. Negative values fail validation; 0 loads as the default when read from YAML |
monitor.topology_audit.debounce_count |
2 |
Consecutive positive ticks before a warning latches (absorbs rollout/restart transients). Negative values fail; 0 loads as the default |
monitor.topology_audit.query_log_lookback_minutes |
15 |
Window of system.query_log scanned for distinct cluster readers. Negative values fail; 0 loads as the default |
First detection lands roughly interval_s × debounce_count after a misconfig
appears: ~10 minutes at the defaults. Lower interval_s for faster
detection. The reason label is sidecar_cluster_queries when this instance's
clickhouse.host is loopback (co-located sidecars) and
multi_instance_cluster_queries otherwise (separate readers not sharing an
election).
Keep
interval_swell above the poll cadence. The auditor counts a host as a live cluster reader only if its most recent whole-cluster read falls inside a recency window it derives frommonitor.check_interval_s, capped atinterval_s / 2. If you raisecheck_interval_spastinterval_s / 2, that cap drops the window below the poll cadence, so a concurrent reader that only reads once percheck_interval_scan fall outside it between ticks. In that case, the misconfig goes undetected, or the warning flaps. Keepinterval_s ≥ 6 × check_interval_sto preserve the full window. click-dog emits a load-timeWARNwhen this combination is detected.
Connection Pool¶
| Configuration path | Default | Description |
|---|---|---|
clickhouse.max_open_conns |
2 |
Maximum number of open connections |
clickhouse.max_idle_conns |
1 |
Maximum number of idle connections |
clickhouse.query_timeout_s |
30 |
Query timeout in seconds. Prevents runaway queries |
clickhouse.max_memory_usage |
104857600 |
Per-query memory limit in bytes (100 MB). Prevents a single monitoring query from consuming excessive memory |
OTEL Exporter¶
exporters:
otel:
- collector_address: localhost:4317 # OTEL collector gRPC address
service_name: click-dog-monitor # Service name in traces
max_query_length: 100000 # Truncate queries longer than this (default: 100000)
# TLS
secure: false # Enable TLS for gRPC (default: false)
insecure_skip_verify: false # Skip cert verification (TESTING ONLY)
ca_cert: "" # Path to CA cert
client_cert: "" # Client cert for mTLS
client_key: "" # Client key for mTLS
| Configuration path | Default | Description |
|---|---|---|
exporters.otel[].collector_address |
None | Required for every entry. OTEL collector gRPC endpoint (e.g., localhost:4317). Supports env expansion |
exporters.otel[].service_name |
click-dog-monitor |
Service name that appears in traces. Supports env expansion |
exporters.otel[].max_query_length |
100000 |
Truncate SQL query text longer than this in exported spans. 0 or omitted is coerced to the 100000 default; the limit is not disabled at config load |
exporters.otel[].secure |
false |
Enable TLS for the gRPC connection. When false, configured CA or complete mTLS material is ignored and produces a validation warning |
exporters.otel[].insecure_skip_verify |
false |
Skip TLS cert verification. Insecure. Requires secure: true; otherwise validation fails |
exporters.otel[].ca_cert |
None | CA certificate path for TLS verification. Supports env expansion and must be readable at config load when secure: true |
exporters.otel[].client_cert |
None | Client certificate for mutual TLS (mTLS). Supports env expansion and must be used with client_key, regardless of secure; it must be readable at config load when secure: true |
exporters.otel[].client_key |
None | Client key for mTLS. Supports env expansion and must be used with client_cert, regardless of secure; it must be readable at config load when secure: true |
Plaintext transport and raw-query compatibility defaults
With secure: false (the default), span export runs over unencrypted gRPC, and click-dog logs a startup warning. filters.query_text_mode also defaults to raw, so query literals can leave the host in db.statement. Set secure: true for any non-local collector and use normalized_only or none for privacy-sensitive production environments. Use redacted only with rules broad enough for the statements you intend to retain; unmatched statements are omitted.
Authentication
Click-Dog supports TLS and mTLS for the gRPC connection. It does not currently support bearer tokens, API keys, or custom gRPC metadata headers. If your OTEL collector requires header-based auth, place an authenticating proxy in front of it or use mTLS for client authentication.
Multiple Export Backends¶
The exporters section accepts any number of OTLP/gRPC sinks under exporters.otel[], and can be combined with Splunk HEC under exporters.splunk_hec[]:
exporters:
otel:
- collector_address: primary-collector:4317
service_name: click-dog-monitor
- collector_address: backup-collector:4317
service_name: clickdog-backup
splunk_hec:
- endpoint: https://splunk.internal:8088
token: ${SPLUNK_HEC_TOKEN}
index: clickhouse
source: click-dog
source_type: _json
allow_insecure_http: false
max_query_length: 100000
Each exporters.splunk_hec[] entry takes a Splunk HTTP Event Collector endpoint, a token (env-expandable), and optional index, source, source_type, max_query_length, and TLS fields. Use https:// endpoints by default because the HEC token is sent in the Authorization header; http:// endpoints are accepted only when allow_insecure_http: true is set explicitly for local, development, or test HEC receivers. The token can also be supplied through token_file, whose path supports ${VAR} expansion. One trailing CRLF or LF is trimmed at load. token and token_file are mutually exclusive.
| Configuration path | Default | Description |
|---|---|---|
exporters.splunk_hec[].endpoint |
None | Required. HEC base URL. Must use https:// unless allow_insecure_http is explicitly enabled; click-dog appends the event path as described in the Splunk HEC guide |
exporters.splunk_hec[].token |
None | Required unless token_file resolves a nonempty token. Inline or env-expanded HEC token. Mutually exclusive with token_file |
exporters.splunk_hec[].token_file |
None | Env-expandable path to the required HEC token file. Mutually exclusive with token; an empty file is rejected |
exporters.splunk_hec[].index |
None | Optional target index |
exporters.splunk_hec[].source |
click-dog |
Event source |
exporters.splunk_hec[].source_type |
_json |
Event source type |
exporters.splunk_hec[].allow_insecure_http |
false |
Permit an http:// endpoint for local development or tests |
exporters.splunk_hec[].insecure_skip_verify |
false |
Skip TLS certificate verification for an https:// endpoint. Insecure. Requires an https:// endpoint; otherwise validation fails |
exporters.splunk_hec[].max_query_length |
100000 |
Truncate SQL in HEC events |
When more than one exporter is configured, spans are sent to all backends concurrently under the default PolicyAllRequired contract. Any backend failure, whether partial or total, is an export error. Starting each call together gives every backend the full monitor.export_timeout_s interval; result aggregation and per-sink statuses remain in configuration order. The cycle is marked an error (click_dog_cycle_results_total{result="error"}), feeds the circuit breaker and adaptive backoff, and shows up as the most recent error on /status and /readyz. The per-sink counters (click_dog_export_{attempts,accepted,errors}_total{sink}) and warning log summaries show which backend failed, so a dead sink cannot silently drop out of the fan-out.
A span is marked as "seen" in the dedup cache only if every backend accepts it. If any backend fails, no span from that batch is marked seen, and the next cycle re-delivers the full batch to every sink. Backends that already accepted those spans will see the same (trace_id, span_id) pair again. Stable IDs make duplicates discoverable, but whether they are collapsed, stored, or rejected is backend-specific. The same contract applies to backfill mode: a partial multi-sink failure counts the query as Failed in the run summary, the process exits non-zero, and re-running the window re-delivers to every sink.
An alternate PolicyAnySuccess policy exists in code: spans are marked seen as soon as one sink accepts them, while failed sinks remain visible through ExportResult and the per-sink metrics. It is not exposed via YAML today; production configs use the strict default policy. The per-call deadline applied to every backend in the fan-out is monitor.export_timeout_s (see Monitor below).
High Availability¶
Leader election via ClickHouse Keeper for multi-node deployments. Configuring
keeper.hosts is what turns election on: there is no separate enable flag.
Omit the ha block entirely and each instance runs standalone.
ha:
keeper:
hosts:
- keeper-01:9181
- keeper-02:9181
- keeper-03:9181
secure: false # Enable TLS to Keeper (default: false)
session_timeout_s: 10 # Session timeout (default: 10, range: 5-30)
base_path: /click-dog/election # Znode path prefix
# auth_user: "" # Optional digest auth
# auth_password: "" # Optional digest auth
| Configuration path | Default | Description |
|---|---|---|
ha.keeper.hosts |
None | ClickHouse Keeper (ZooKeeper-compatible) endpoints. Presence enables leader election; empty/omitted means standalone |
ha.keeper.secure |
false |
Enable TLS for the Keeper connection. TLS uses the host's system trust roots and supports neither a configured custom CA nor client-certificate mTLS; private-CA Keeper requires installing that CA into the service host/container trust store. Use TLS with Keeper digest auth on untrusted networks so credentials are not sent in plaintext |
ha.keeper.session_timeout_s |
10 |
With nonempty keeper.hosts, 0/omitted becomes 10 and effective values must be 5–30. With no hosts, election is inactive and this field is ignored and not range-validated |
ha.keeper.base_path |
/click-dog/election |
Znode path prefix. Must be exactly /click-dog or under /click-dog/; paths overlapping ClickHouse's own Keeper paths (/clickhouse...) are rejected |
ha.keeper.auth_user |
None | Optional digest authentication username (supports ${ENV} expansion) |
ha.keeper.auth_password |
None | Optional digest authentication password (supports ${ENV} expansion). Mutually exclusive with keeper.auth_password_file |
ha.keeper.auth_password_file |
None | Path to a file holding the digest auth password (read at load; one trailing CRLF or LF is trimmed). Path supports ${VAR} expansion. Mutually exclusive with keeper.auth_password |
Each click-dog instance creates an ephemeral sequential znode. The lowest sequence number becomes leader. When the leader dies, its session expires, the znode is deleted, and the next candidate auto-promotes (~10 seconds failover).
When keeper.auth_user is omitted, election znodes use Keeper world ACLs; click-dog emits a startup warning so that choice is visible. Configure digest auth for shared Keeper ensembles, and set keeper.secure: true when Keeper supports TLS.
Leader gating depends on topology
In cluster mode (use_cluster_queries: true), leadership gates the data path: only the leader fetches and exports; standbys skip each cycle (recorded as a skipped cycle, no export) so the cluster's spans aren't duplicated. In sidecar mode each instance reads its own node's disjoint local scope and always exports regardless of leader status: there leadership only drives coordination duties (flush, /clusterz). Either way, configuring ha.keeper.hosts is what activates election.
Keeper startup behavior depends on how far initialization gets:
- With resolvable Keeper addresses and no digest authentication, the pinned ZooKeeper client starts network dialing asynchronously. The instance fails open and keeps exporting while the election loop retries, then joins automatically when Keeper becomes reachable.
- If election construction returns an error, click-dog logs
Failed to join leader electionand leaves that process standalone until restart. Host parsing or DNS resolution can take this path. InitialAddAuthis synchronous, so an authentication error or an unavailable Keeper whileauth_useris configured can take it too.
Monitor¶
monitor:
enabled: true # Enable scheduled monitoring (default: true)
# Duration thresholds (milliseconds)
min_trace_duration_ms: 1000 # Find traces with spans >= this (required)
min_span_duration_ms: 0 # Only export spans >= this (0 = all)
max_trace_duration_ms: 0 # Skip traces with spans > this (0 = no limit)
max_span_duration_ms: 0 # Skip spans > this (0 = no limit)
max_query_length: 100000 # Skip queries with SQL > this chars (see also exporters.otel[].max_query_length which truncates instead)
# Polling
check_interval_s: 30 # Polling interval in seconds
lookback_s: 40 # How far back to look (default: check_interval_s + lookback_buffer_s)
lookback_buffer_s: 10 # Extra seconds added when defaulting lookback_s (default: 10)
# Rate limiting
max_spans_per_cycle: 1000 # Max spans fetched and exported per polling cycle (default: 1000)
# Batch processing
batch_size: 0 # Batch size (0 = no batching)
batch_delay_ms: 0 # Delay between batches (0 = no delay)
export_timeout_s: 30 # Per-call exporter deadline in seconds (0 disables)
# Deduplication
dedup_cache_size: 10000 # LRU cache entries (default: 10000)
# Degraded-mode canary
canary:
enabled: false
threshold_duration_ms: 60000
Duration Filtering¶
Click-Dog uses a two-tier duration filtering approach:
- Trace-level (
min_trace_duration_ms/max_trace_duration_ms): Find traces that contain at least one span within this duration range - Span-level (
min_span_duration_ms/max_span_duration_ms): From those traces, only export individual spans within this duration range
All duration filtering happens in SQL for maximum efficiency.
| Configuration path | Default | Description |
|---|---|---|
monitor.enabled |
true |
Enable scheduled monitoring. When true, min_trace_duration_ms and check_interval_s must both be greater than zero. Set false for deliberately backfill-only usage |
monitor.min_trace_duration_ms |
None | Required. Find traces containing spans >= this duration (ms) |
monitor.min_span_duration_ms |
0 |
Only export spans >= this (ms). 0 = export all spans from matching traces |
monitor.max_trace_duration_ms |
0 |
Skip traces with spans > this (ms). 0 = no upper limit |
monitor.max_span_duration_ms |
0 |
Skip individual spans > this (ms). 0 = no upper limit |
monitor.max_query_length |
100000 |
Skip queries with SQL text longer than this (chars) and truncate live query_log.normalized_query previews to this length. 0 or omitted is coerced to the 100000 default; the limit is not disabled at config load. Different from exporters.otel[].max_query_length which truncates at export time |
Polling¶
| Configuration path | Default | Description |
|---|---|---|
monitor.check_interval_s |
30 |
How often to poll ClickHouse (seconds). Must be greater than 0 when enabled is true |
monitor.lookback_s |
check_interval_s + lookback_buffer_s |
How far back to look for spans. Defaults to interval + the buffer below, to avoid gaps between polls. Once spans have been exported, a later cycle resumes from the last exported position even when it has fallen behind this window, up to three lookback windows back (120s at the defaults, which covers a failed export followed by the first backoff step); spans further behind are dropped. A process with nothing exported yet starts at the plain window. Once a cycle has read everything new, spare max_spans_per_cycle goes to re-reading already-exported spans in the window for rows that flushed late; the dedup cache skips the repeats. Set this explicitly to override the computed default |
monitor.lookback_buffer_s |
10 |
Seconds added to check_interval_s when computing the default lookback_s. Absorbs clock skew and slow poll cycles; increase under high load if you see missed spans. Ignored when lookback_s is set explicitly |
Rate Limiting & Batching¶
| Configuration path | Default | Description |
|---|---|---|
monitor.max_spans_per_cycle |
1000 |
Maximum number of spans fetched and exported per polling cycle. Applied as a SQL LIMIT on the spans query (and as an upper bound on each trace-ID page, since each trace contributes at least one span). Each cycle continues through the lookback window oldest first from where the last successfully exported cycle stopped, so the cap is spent on spans not yet exported, and a trace with more spans than the cap is finished over the following cycles. The walk only moves once a cycle's exports succeed. Inflow below max_spans_per_cycle / check_interval_s is read in full; a trace the walk falls more than three lookback windows behind is dropped, which is what eventually happens to the oldest spans when inflow stays above that rate. A debug-level log line reports the total span count per cycle. The absolute ceiling is 100000 spans per cycle to bound memory; config validation rejects larger values at load |
monitor.batch_size |
0 |
Process spans in batches of this size. 0 = process all at once |
monitor.batch_delay_ms |
0 |
Milliseconds to wait between batches. 0 = no delay |
monitor.export_timeout_s |
30 |
Per-call deadline for exporter calls. Applies to live span batches, backfill query exports, and degraded-mode canary span exports so a stuck collector fails into the normal retry, backoff, and circuit-breaker path. Set to 0 to disable the client-side export deadline |
monitor.dedup_cache_size |
10000 |
Size of the LRU cache for span deduplication, keyed by the composite OTLP span identity (trace_id, span_id). OTLP span IDs are only unique within a trace, so dedup must include the trace ID; a span-id-only cache would silently drop sibling spans in different traces that collide on the 64-bit ID. Each entry's SpanKey is 24 B (16 B trace_id + 8 B span_id), but the LRU container adds doubly-linked-list pointers and a map bucket, so the practical cost is ~80–100 B per entry on 64-bit. 10,000 entries is approximately 1 MiB. Memory scales linearly. The cache is in-memory only; a restart can resend spans still inside lookback_s, and downstream duplicate handling is backend-specific |
monitor.extract_log_comment |
true |
Extract log_comment from URI attributes and promote JSON keys as span attributes prefixed with log_comment.. See Span Attributes |
monitor.enrich_from_query_log |
true |
Enrich spans with metadata from system.query_log (user, client, tables, query stats) by joining on query_id (from the clickhouse.query_id span attribute). Note: includes query_log.user, query_log.client_address, and query_log.client_hostname, which may contain PII or internal network information. See the privacy note and Span Attributes |
Degraded-mode canary¶
| Configuration path | Default | Description |
|---|---|---|
monitor.canary.enabled |
false |
Run and export a lightweight canary query while the circuit breaker is open or adaptive backoff is elevated |
monitor.canary.threshold_duration_ms |
60000 |
Count ClickHouse spans exceeding this duration when constructing the canary span. Must not be negative |
The canary queries ClickHouse and exports the result, so it exercises both
sides even when the breaker was tripped by export failures. When a circuit
breaker is configured, that success is reported to it — but it counts toward
success_threshold only once the breaker has reached half-open. While the
breaker is still open, the report clears the failure count without advancing
recovery. With monitor.circuit_breaker.enabled: false the canary can still
run on elevated backoff alone, and there is no breaker to report to.
See Troubleshooting · Canary Queries in Degraded Mode for emitted attributes and failure behavior.
Circuit Breaker¶
Protects ClickHouse by stopping queries after repeated failures.
monitor:
circuit_breaker:
enabled: false # Enable circuit breaker (default: false)
failure_threshold: 3 # Open after N consecutive failures
success_threshold: 1 # Close after N successes in half-open state
reset_timeout_s: 60 # Try again after N seconds when open
| Configuration path | Default | Description |
|---|---|---|
monitor.circuit_breaker.enabled |
false |
Enable the circuit breaker. Note: the install script enables this by default in generated configs |
monitor.circuit_breaker.failure_threshold |
3 |
Number of consecutive failures before opening the circuit |
monitor.circuit_breaker.success_threshold |
1 |
Number of successful requests needed in half-open state to close the circuit |
monitor.circuit_breaker.reset_timeout_s |
60 |
Seconds to wait before transitioning from open to half-open |
See Resilience for details on circuit breaker behavior.
Adaptive Backoff¶
Increases polling interval when errors occur, reducing load on a struggling ClickHouse.
monitor:
backoff:
enabled: false # Enable adaptive backoff (default: false)
max_interval_s: 300 # Maximum polling interval (default: 300 = 5 min)
backoff_factor: 2.0 # Multiply interval on failure (default: 2.0)
| Configuration path | Default | Description |
|---|---|---|
monitor.backoff.enabled |
false |
Enable adaptive backoff. Note: the install script enables this by default in generated configs |
monitor.backoff.max_interval_s |
300 |
Maximum polling interval in seconds (5 minutes) |
monitor.backoff.backoff_factor |
2.0 |
Multiply the current interval by this factor on each failure. Must be greater than 1; values in (0, 1] are rejected at config load |
See Resilience for details on backoff behavior.
Filters¶
filters:
query_text_mode: redacted # raw | redacted | normalized_only | none
whitelist_operations: [] # Operation names to INCLUDE (supports * wildcard)
blacklist_operations: # Operation name substrings to EXCLUDE at SQL level
- "MergeTreeIndex"
- "VFSWrite"
whitelist_ips: [] # IP addresses to INCLUDE
whitelist_users: [] # ClickHouse users to INCLUDE (strict: drops spans with unknown user)
blacklist_users: [] # ClickHouse users to EXCLUDE (defense-in-depth alongside CH user grants)
blacklist_queries: # Regex patterns to EXCLUDE
- "^SELECT \\* FROM system\\."
- "SHOW TABLES"
- "(?i)healthcheck"
redact_queries: # Regex replacements applied before export
- pattern: "(?i)identified\\s+by\\s+'[^']*'"
replacement: "IDENTIFIED BY '[REDACTED]'"
| Configuration path | Default | Description |
|---|---|---|
filters.query_text_mode |
raw |
Final query-text representation applied after filtering/enrichment and before exporter fan-out. normalized_only and none never fall back to raw text and are recommended for privacy-sensitive production environments |
filters.whitelist_operations |
[] |
If set, only export spans with matching operation names. Supports * wildcard (e.g., DB::Interpreter*::execute()) |
filters.blacklist_operations |
[] |
Substring patterns matched against operation_name. Pushed down to ClickHouse as NOT LIKE '%pattern%' so matching spans are never fetched. Use this to drop high-volume internal pipeline spans (MergeTreeIndex, VFSWrite, etc.) without affecting query-level spans |
filters.whitelist_ips |
[] |
If set, only export spans from these client IP addresses. Entries may be exact IPs (10.0.1.50) or CIDR ranges (10.0.0.0/8) |
filters.whitelist_users |
[] |
If set, only export spans whose query_log.user is on the list. Pushed down to enrichment SQL as AND user IN (...) and re-checked per span. This is strict: spans with an unresolvable user are dropped |
filters.blacklist_users |
[] |
Never export spans whose query_log.user is on the list. Enforced at the Go layer for the span path (pushing it to enrichment SQL would strip the user and let blacklisted spans bypass the check); pushed down to SQL on the slow-query / backfill path |
filters.blacklist_queries |
[] |
Regex patterns matched against SQL query text. Matching spans are excluded |
filters.redact_queries[].pattern |
None | Regex pattern matched against SQL query text after blacklist filtering and before export. Used only with query_text_mode: redacted; at least one valid rule is required. An unmatched statement is omitted rather than exported raw. Rules under raw are invalid; stale rules under normalized_only/none are ignored with a warning |
filters.redact_queries[].replacement |
[REDACTED] |
Replacement text used for matches. Empty or omitted uses the default replacement |
Note: blacklist_operations runs at the SQL level before data is fetched. whitelist_operations runs after fetch in Go. The blacklist always takes precedence because matching spans never reach the whitelist check.
See Filtering for detailed examples and evaluation order.
Logging¶
log_level: info # debug, info, warn, error (default: info)
log_file: "" # Log file path (default: stderr)
log_format: text # text or json (default: text)
log_rotation:
max_size_mb: 100 # Rotate when the active file passes this size (default: 100; must be >= 1)
max_files: 3 # Keep this many rotated files (default: 3; must be >= 1)
| Configuration path | Default | Description |
|---|---|---|
log_level |
info |
Log verbosity: debug, info, warn, error |
log_file |
None | Optional file path. If empty, logs go to stderr (container-friendly) |
log_format |
text |
Output format: text (human-readable) or json (structured). JSON emits {"ts":"...","level":"...","msg":"..."} |
log_rotation.max_size_mb |
100 |
Rotation threshold per file. Omit the key to take the default; an explicit 0 is rejected (would disable rotation and produce unbounded growth). |
log_rotation.max_files |
3 |
Number of rotated files to retain. Must be >= 1. |
Log file permissions¶
When log_file is set, click-dog creates the file with mode 0640
(owner: read/write, group: read, world: none). Operational metadata such as
client IPs, operation names, and error traces is sensitive enough to keep
off world-readable storage but the service group still needs read access
for log shipping.
If a separate log-shipping agent runs alongside click-dog (Filebeat, Fluent Bit node agent, journald forwarder, Vector, Promtail, etc.) under a different UID, that agent must be in the click-dog service user's group for log shipping to work. Otherwise the agent will silently fail to read rotated files. Typical setup:
# Add the log shipper's user to click-dog's group
usermod -aG click-dog filebeat
systemctl restart filebeat
Earlier releases created the log file with mode 0644 (world-readable).
Upgrading deployments that relied on an unprivileged log shipper reading
the file via the world bit will need to adjust group membership as
above.
Metrics¶
metrics:
enabled: false # Enable scrape + admin endpoints
listen_address: ":9090" # Scrape listener: /metrics + legacy /health (default: :9090)
admin_listen_address: "127.0.0.1:9091" # Admin listener: POST /flush (default: 127.0.0.1:9091)
otlp:
enabled: false # Optional push path for self-metrics
inherit_otel_connection: true # Reuse exporters.otel[0] when enabled
interval_seconds: 10
| Configuration path | Default | Description |
|---|---|---|
metrics.enabled |
false |
Start both the scrape and admin HTTP listeners. This setting gates both; false also disables HTTP POST /flush (SIGUSR1 still works) |
metrics.listen_address |
:9090 |
Scrape listener (/metrics, legacy /health); safe to expose to your scraper |
metrics.admin_listen_address |
127.0.0.1:9091 |
Admin listener (POST /flush); loopback by default. Non-loopback binds emit a startup WARN |
metrics.otlp.enabled |
false |
Push click-dog self-metrics over OTLP. Additive to the /metrics scrape endpoint; it may run with metrics.enabled: false |
metrics.otlp.inherit_otel_connection |
true when OTLP metrics are enabled and the key is omitted |
Reuse exporters.otel[0] connection settings for the self-metrics exporter. Requires at least one OTLP trace exporter when enabled |
metrics.otlp.collector_address |
None | Standalone OTLP metrics collector address, required when inherit_otel_connection: false. Supports env expansion |
metrics.otlp.secure |
false |
Enable TLS for a standalone OTLP metrics connection. Use inheritance when CA or mTLS settings are needed |
metrics.otlp.host |
OS hostname | Host identity used on pushed self-metrics. Supports env expansion |
metrics.otlp.service_name |
exporters.otel[0].service_name, then click-dog-monitor |
Service name used on pushed self-metrics. Supports env expansion |
metrics.otlp.interval_seconds |
10 |
Push interval. Must be greater than 0 when otlp.enabled is true |
metrics.otlp.rename |
{} |
Optional map from canonical self-metric keys to emitted OTLP metric names. Keys are validated at load time |
The
metrics:andhealth:blocks are independent.metrics.enableddoes not control the Kubernetes-shaped/healthz//readyz//statusprobes. Those live in the separatehealth:block below. Disabling metrics will not turn off the health listener, and vice versa.
See Observability for the full endpoint table, the rationale for splitting scrape from admin, and deployment recipes (systemd / Docker / Kubernetes).
Health Endpoints¶
health:
enabled: false # Off by default; install.sh enables it on :8686
listen_address: ":8686" # /healthz, /readyz, /status (default when enabled)
cluster:
enabled: false # Leader-only /clusterz aggregate endpoint
self: "" # Routable host:port for this instance
peer_timeout_ms: 3000
peers: []
| Configuration path | Default | Description |
|---|---|---|
health.enabled |
false |
Start the health listener. Independent of metrics.enabled |
health.listen_address |
:8686 when health.enabled: true |
/healthz (liveness), /readyz (readiness; pings ClickHouse), /status (JSON report). If this normalizes to the same socket as metrics.listen_address, the handlers are mounted on the metrics listener so only one port is opened |
health.cluster.enabled |
false |
Mount leader-only /clusterz aggregate health when health.enabled is also true |
health.cluster.self |
None | Routable bare host:port for this instance's /readyz; required when cluster health is enabled. Do not include http://, a path, query, or fragment |
health.cluster.peer_timeout_ms |
3000 |
Per-peer /readyz fanout deadline |
health.cluster.peers |
[] |
Static list of bare peer host:port addresses to include in /clusterz. Empty entries, schemes/paths, missing ports, and duplicate peers are rejected |
/healthz always returns 200; /readyz returns 503 while ClickHouse is
unreachable or the circuit breaker is open. /status is a reporting endpoint
that always returns 200: assess health from its JSON body, not its status
code.
See Observability: Health Endpoints
for the full endpoint table, the /status JSON shape, Kubernetes probe
recipes, and cluster-aware health roll-up.
Webhook Notifications¶
webhook:
enabled: false
url: ${SLACK_WEBHOOK_URL} # Webhook URL (supports env vars)
timeout_s: 10 # HTTP timeout in seconds (default: 10)
events: # Event filter (default: all events)
- analysis_findings
- circuit_breaker_opened
- circuit_breaker_closed
- startup
- shutdown
| Configuration path | Default | Description |
|---|---|---|
webhook.enabled |
false |
Enable webhook notifications |
webhook.url |
None | HTTP POST endpoint. Supports ${VAR} expansion. Must be http:// or https:// with a host; redirects are not followed |
webhook.timeout_s |
10 |
HTTP request timeout in seconds |
webhook.events |
[] (all) |
List of event names to fire on. Empty = all events |
Valid events: analysis_findings, circuit_breaker_opened,
circuit_breaker_closed, backfill_complete, backfill_failed,
error_spike, startup, and shutdown. analysis_findings is evaluated only
by an explicit click-dog analyze queries --notify run; enabling the destination
does not make scheduled monitoring run query analysis.
See Observability for payload format, behavior details, and compatible services.
Datadog Event Management¶
datadog_events:
enabled: false
site: datadoghq.com
api_key: ${DD_API_KEY} # or api_key_file
# api_key_file: /run/secrets/datadog-api-key
application_key: ${DD_APP_KEY} # or application_key_file
# application_key_file: /run/secrets/datadog-app-key
timeout_s: 10
environment: production
service: click-dog
| Configuration path | Default | Description |
|---|---|---|
datadog_events.enabled |
false |
Enable this destination for explicit analyze queries --notify runs |
datadog_events.site |
None | Trusted Datadog site hostname used to derive https://event-management-intake.<site>/api/v2/events; a scheme, path, port, query, fragment, IP address, or malformed DNS hostname is rejected |
datadog_events.api_key |
None | Datadog API key. Supports ${VAR} expansion; mutually exclusive with datadog_events.api_key_file |
datadog_events.api_key_file |
None | File containing the API key; exactly one trailing LF or CRLF is trimmed and other whitespace is preserved; mutually exclusive with datadog_events.api_key |
datadog_events.application_key |
None | Datadog application key. Supports ${VAR} expansion; mutually exclusive with datadog_events.application_key_file |
datadog_events.application_key_file |
None | File containing the application key; exactly one trailing LF or CRLF is trimmed and other whitespace is preserved; mutually exclusive with datadog_events.application_key |
datadog_events.timeout_s |
10 |
Single-attempt request timeout, from 1 through 60 seconds |
datadog_events.environment |
None | Required low-cardinality environment:<value> tag; letters, digits, ., _, :, /, and -, at most 200 characters |
datadog_events.service |
None | Required low-cardinality service:<value> tag with the same character and length rules |
When enabled, site, environment, service, and one source for each
credential are required. Both credentials are required because the Events API
v2 request is authenticated with DD-API-KEY and DD-APPLICATION-KEY. Secret values are
never included in validation or delivery errors. The client performs one
bounded request, follows no redirects, and treats only HTTP 2xx as intake
acceptance. It does not retry, and acceptance is not confirmation that an Event
Monitor evaluated or notified.
This destination handles the bounded query-analysis notification summary. It does not replace or share the OTLP trace or self-metric exporters. See Datadog integration for setup and Event Monitor guidance.
Complete Example¶
Start from the complete minimal profile, or
see config.yaml.example in the repository root for a fully commented
configuration file with all options.