Skip to content

Configuration Reference

Click-Dog uses YAML configuration with environment variable support. When --config is omitted, it first looks for /etc/click-dog/click-dog.yaml, then falls back to ./click-dog.yaml in the current directory. Override with --config:

./click-dog --config /path/to/config.yaml

Unknown keys are rejected at load time, so a typo like check_intervla_s fails fast instead of being silently ignored. Use click-dog validate to check that your config parses, passes validation, and references readable files for active TLS settings without starting the service or making network connections.

Configuration fragments

Unless a YAML block is explicitly labeled as complete, it is a fragment to merge into a standalone configuration. In particular, an exporter-only or observability-only block does not include the required scheduled-monitor threshold. Start from the complete minimal profile, then merge the focused fragment into it.

Validate and run

For a manual foreground run, save your configuration as click-dog.yaml and replace the example addresses, credentials, and any TLS file paths. Complete templates are available in the minimal profile and production profile.

If your config uses ${CLICKHOUSE_PASSWORD}, set it in the same shell first:

export CLICKHOUSE_PASSWORD='your-clickhouse-password'

Validate the file offline, without opening network connections:

./click-dog validate --config click-dog.yaml

Then preview one polling cycle against your real ClickHouse. The dry-run discards span exports, prints a summary, and exits. Enabled self-metrics and startup webhooks can still send traffic; see dry-run behavior.

./click-dog --dry-run --config click-dog.yaml

When the result looks right, start continuous polling and export:

./click-dog --config click-dog.yaml

Press Ctrl+C to stop the foreground process. For an installed service, follow the service setup instructions instead of starting a second instance.

Validation and migrations

Run the offline validator before a rollout:

sudo -u click-dog click-dog validate --config /etc/click-dog/click-dog.yaml

It checks YAML shape, configuration relationships, and readability of certificate files used by active ClickHouse/OTLP TLS connections. It does not connect to ClickHouse or an exporter. click-dog check adds network and data-plane checks; see what the HEC check proves for that exporter's intentionally limited result.

The load-bearing validation rules are:

Area Contract
Standalone minimum Scheduled mode defaults on. When monitor.enabled: true, both monitor.min_trace_duration_ms and monitor.check_interval_s must be greater than zero. At least one entry under exporters.otel or exporters.splunk_hec is also required. Consequently, omitting monitor: is not a valid scheduled-mode config. Set monitor.enabled: false only for a deliberately backfill-only config.
ClickHouse addressing clickhouse.port must be 1–65535. When present, clickhouse.cluster must start with an ASCII letter and then contain only letters, digits, _, or -; it is required when use_cluster_queries: true.
Exporter entries Every exporters.otel[] entry requires a nonblank collector_address. Every exporters.splunk_hec[] entry requires an endpoint and a nonempty token (inline or from token_file); endpoint scheme rules are described under Multiple Export Backends.
OTLP self-metrics When metrics.otlp.enabled: true, interval_seconds must be positive. Inheritance requires at least one exporters.otel entry. With inherit_otel_connection: false, metrics.otlp.collector_address is required.
User filters filters.whitelist_users and filters.blacklist_users require monitor.enrich_from_query_log: true, because the span log does not provide the originating user by itself.
Query-text privacy filters.query_text_mode accepts raw, redacted, normalized_only, or none. redacted requires a valid rule; the ambiguous normalized spelling is rejected. Rules under raw are rejected, while inert rules under normalized_only/none produce a warning.
Topology audit Negative audit timing/count values are rejected. When enabled, the effective interval_s, debounce_count, and query_log_lookback_minutes are positive; a YAML value of 0 is normalized to its default (300, 2, and 15, respectively).
Cluster health health.cluster.enabled: true requires health.enabled: true and a routable health.cluster.self. self and every peer must be bare host:port (bracket IPv6), without a scheme, path, query, or fragment. Empty and duplicate peer entries are rejected after trimming whitespace.
TLS coherence clickhouse.insecure_skip_verify and exporters.otel[].insecure_skip_verify require their corresponding secure: true; HEC insecure_skip_verify requires an https:// endpoint. An OTLP client certificate and key must be supplied together. TLS files used by active ClickHouse/OTLP connections must be readable at load time. Inactive CA or complete OTLP client-certificate material is accepted but emits a warning, as documented in the relevant TLS tables below.

filters.redact_queries[].pattern expressions are compiled by configuration validation. filters.blacklist_queries has a different lifecycle: its regexes are compiled when the daemon or backfill query filter is initialized, not by click-dog validate or click-dog check. An invalid blacklist expression can therefore pass both checks and then stop daemon/backfill startup with Failed to initialize query filter. See Filtering · Regex syntax.

Removed keys are errors, not compatibility aliases:

  • Replace the removed top-level otel: mapping with a list entry under exporters.otel:. The strict YAML decoder reports the old key as unknown; click-dog does not migrate it automatically.
  • Delete the removed ha.enabled key. Configure ha.keeper.hosts to enable election, or omit the hosts to run standalone. This key receives a targeted removal error rather than the generic unknown-field error.

There are currently no legacy keys that load with only a deprecation warning. Unknown and removed keys fail closed so a migration cannot be silently ignored. For query-text migration only, a config with legacy redact_queries rules and no query_text_mode resolves to redacted and emits a validation warning; this preserves the privacy intent without silently returning to raw mode. The new mode is stricter than legacy redaction: unmatched statements are omitted, and the migration warning states that data-visibility change explicitly.

Environment Variable Expansion

Selected string fields support ${VAR} and $VAR syntax. Unset variables resolve to empty strings and produce a startup warning when the braced form was used.

clickhouse:
  host: ${CLICKHOUSE_HOST}
  username: ${CLICKHOUSE_USERNAME}
  password: ${CLICKHOUSE_PASSWORD}

exporters:
  otel:
    - collector_address: ${OTEL_COLLECTOR_ADDRESS}
      service_name: ${OTEL_SERVICE_NAME}
      ca_cert: ${OTEL_CA_CERT}
      client_cert: ${OTEL_CLIENT_CERT}
      client_key: ${OTEL_CLIENT_KEY}

The following fields support environment variable expansion: - clickhouse.host, clickhouse.username, clickhouse.password, clickhouse.password_file, clickhouse.ca_cert - exporters.otel[].collector_address, exporters.otel[].service_name, exporters.otel[].ca_cert, exporters.otel[].client_cert, exporters.otel[].client_key - exporters.splunk_hec[].endpoint, exporters.splunk_hec[].token, exporters.splunk_hec[].token_file - metrics.otlp.collector_address, metrics.otlp.host, metrics.otlp.service_name - ha.keeper.auth_user, ha.keeper.auth_password, ha.keeper.auth_password_file - webhook.url - datadog_events.site, datadog_events.api_key, datadog_events.application_key, datadog_events.environment, and datadog_events.service

The *_file secret-path fields (clickhouse.password_file, exporters.splunk_hec[].token_file, ha.keeper.auth_password_file, datadog_events.api_key_file, and datadog_events.application_key_file) are also expanded, so the path can reference an environment variable. For example: password_file: ${SECRET_PATH}.


ClickHouse Connection

clickhouse:
  host: localhost              # ClickHouse hostname
  port: 9000                   # Native protocol port (default: 9000)
  database: default            # Database name
  username: default            # Username (supports env vars)
  password: ${CLICKHOUSE_PASSWORD}  # Password (supports env vars)

  # TLS
  secure: false                # Enable TLS (default: false)
  insecure_skip_verify: false  # Skip cert verification (TESTING ONLY)
  ca_cert: ""                  # Path to CA cert for self-signed certs

  # Cluster
  cluster: ""                  # Cluster name for distributed queries
  use_cluster_queries: false   # Use cluster() function (default: false)

  # Connection pool
  max_open_conns: 2            # Max open connections (default: 2)
  max_idle_conns: 1            # Max idle connections (default: 1)
  query_timeout_s: 30          # Query timeout in seconds (default: 30)

  # Memory
  max_memory_usage: 104857600  # Per-query memory limit in bytes (default: 100MB)

Connection Details

Configuration path Default Description
clickhouse.host None ClickHouse server hostname
clickhouse.port 9000 Native protocol port; must be between 1 and 65535
clickhouse.database None Database name
clickhouse.username None Username
clickhouse.password None Password (supports ${ENV} expansion). Mutually exclusive with password_file
clickhouse.password_file None Path to a file holding the password (read at load; one trailing CRLF or LF is trimmed). Path supports ${VAR} expansion. Keeps the secret out of the process environment. Mutually exclusive with password. An empty file is rejected; use inline password: "" for an intentionally empty password

TLS

Configuration path Default Description
clickhouse.secure false Enable TLS encryption. When false, a configured ca_cert is ignored and produces a validation warning
clickhouse.insecure_skip_verify false Skip certificate verification. Insecure. Use only for testing. Requires secure: true; otherwise validation fails
clickhouse.ca_cert None Path to a CA certificate file for TLS verification with self-signed or internal CA certs. Supports env expansion and must be readable at config load when secure: true

Cluster Mode

Click-Dog supports two reader topologies. A single deployment is a sidecar deployment with one instance. The config is identical; only the instance count differs.

Sidecar (default, recommended): Deploy one Click-Dog instance per ClickHouse node. Each instance reads only its local system.opentelemetry_span_log; the partitions are disjoint by host, so the union gives full cluster coverage with zero coordination. This is the topology that distributes the expensive cluster-wide read. Each instance has minimal permissions and generates minimal load. No Keeper, no leader election. Deploy the same config once per node.

Cluster (use_cluster_queries: true): A Click-Dog instance reads the whole cluster by wrapping the span query as cluster('cluster_name', system.opentelemetry_span_log) (one replica per shard). Requires network access to all nodes and a non-empty cluster: name. Cluster mode is always leader-gated (see below): run 1 instance (the leader of an election of one) or 2–3 for automatic failover. Adding instances doesn't divvy up the read: they stand by behind one active exporter.

Grants. cluster() is a table function, and ClickHouse gates it behind the REMOTE privilege on top of the per-table SELECT grants, so the monitoring user needs, on every node:

GRANT SELECT ON system.opentelemetry_span_log TO click_dog_monitor;
GRANT SELECT ON system.query_log TO click_dog_monitor;
GRANT REMOTE ON *.* TO click_dog_monitor;          -- cluster() / clusterAllReplicas()
GRANT SELECT ON system.columns TO click_dog_monitor; -- replica capability probes
GRANT SELECT ON system.processes TO click_dog_monitor; -- only for `analyze trace --source current`

Without REMOTE, every read fails with Not enough privileges ... READ ON REMOTE and click-dog check fails the span_log readable probe naming the grant. Without system.columns, the startup probes that verify normalized_query_hash and query_kind on every replica cannot run, so they warn and the normalized query attributes and operation enrichment stay off (click_dog_normalized_query_supported and click_dog_query_operation_supported read 0). Grant ON CLUSTER so every node the reader can connect to agrees. Which identity the remote shard uses for the read depends on the cluster definition: without an inter-server <secret>, ClickHouse connects to the other shards as the <user> named in remote_servers (commonly default), not as the monitoring user.

A shard that is down fails the whole read. cluster() runs without skip_unavailable_shards, so while any shard is unreachable every cycle errors (Error fetching spans: ... All connection tries failed), adaptive backoff grows, the circuit breaker opens after failure_threshold consecutive failures, and /readyz reports 503 until the shard returns. Nothing is exported for the reachable shards in the meantime, and the lookback window bounds what is recovered afterwards. This is deliberate: a partial read would look complete while silently missing that shard's spans. Sidecars on the surviving nodes are unaffected, which is one reason sidecar is the recommended topology for larger clusters.

One distributed query is not necessarily one trace. Whether the shard work shares the initiator's trace_id is ClickHouse's context propagation, not Click-Dog's. A trace started by opentelemetry_start_trace_probability is not propagated to remote shards (verified on ClickHouse 25.8): each shard records its part as a separate trace, and each qualifies for export on its own duration, so one slow distributed query can arrive as 1 + shards traces. A trace started from a client-supplied traceparent header is propagated, and every shard's spans land in one trace, each tagged with its own hostname.

Configuration path Default Description
clickhouse.cluster None Cluster name. Required when use_cluster_queries is true; must match [A-Za-z][A-Za-z0-9_-]* because it is interpolated into cluster() queries
clickhouse.use_cluster_queries false When true, uses cluster('name', table) to read one replica per shard across the cluster (cluster topology). When false, reads only this node's local tables (sidecar topology). Optional normalized query attributes are enabled in cluster query mode only after clusterAllReplicas('name', system.columns) verifies every replica exposes normalized_query_hash; mixed or unknown clusters keep the safe fallback

Cluster mode is always leader-gated

In cluster mode the fetch/export cycle runs only on the election leader; other instances stand by (recorded as skipped cycles, so a standby's last-success gauge legitimately stays at zero). The rule is export iff no election is active, or this instance is the leader, so:

  • A single cluster reader with no Keeper configured is valid and common: no election starts, so it always exports (no regression for existing use_cluster_queries: true users).
  • With Keeper (2–3 instances), one is elected leader and exports; the rest defer. On leader loss, promotion normally takes about the configured Keeper session timeout (10s default; valid range 5-30s), plus election overhead. The new leader runs a cycle as soon as it is promoted rather than waiting for its next tick, re-reading from now - lookback. When that lookback covers the full failover interval (the defaults do: 40s against roughly 10s), the boundary is duplicate-prone overlap; a shorter explicit lookback can let spans age out before promotion.
  • A cluster reader waits up to the session timeout for its election to join before running its startup cycle, so an instance that starts as a standby under a healthy Keeper exports nothing and its last-success gauge stays at zero. If Keeper is unreachable the wait times out and the startup cycle runs fail-open.
  • Leader election is driven solely by the presence of ha.keeper.hosts. There is no separate enable flag. Configure Keeper hosts and instances coordinate; omit them and each instance runs standalone. Only the cluster-query data path is leader-gated. Sidecars always export their disjoint node-local scopes; Keeper leadership on a sidecar is used only for coordination duties such as /clusterz and shared flush requests.

The delivery contract is best-effort with at-least-once retries, not exactly-once. Click-Dog preserves (trace_id, span_id), which lets operators identify repeat deliveries, but OTLP does not require collectors or backends to collapse them. Under a healthy Keeper, cluster mode guarantees no steady-state duplication, not "never duplicates." Every Keeper disruption after an instance has joined fails open: a session loss, reconnect, partition, or watch error drops candidate state so the instance keeps exporting (rather than stalling as a gated standby), and the election goroutine retries until it rejoins. The resulting duplicate window is therefore bounded and recovers to a single exporter when Keeper returns. An election-constructor failure is the exception: host parsing or DNS resolution, or an initial authentication failure, leaves that process standalone until restart. Click-Dog does not fail closed: it prefers availability.

Pick one topology per fleet

Don't run per-node sidecars with use_cluster_queries: true. Leader-gating now prevents steady-state duplication when those instances share an election, but sidecars that are not sharing an election (different base_path, or no Keeper) would each go cluster-wide and duplicate. If you want whole-cluster reads, run cluster mode with a shared Keeper; if you want one reader per node, leave use_cluster_queries off. See Operating · Cluster topology.

This is enforced at runtime. When use_cluster_queries: true, click-dog runs an in-process topology self-audit that detects more than one instance doing whole-cluster reads and surfaces it via the click_dog_topology_warning metric, the /status topology_warning field, a throttled WARN log, and the Datadog health dashboard. It is observability-only: it never gates readiness or de-rotates a node and is a backstop for the case leader-gating can't reach (instances not sharing an election). See monitor.topology_audit below.

Topology self-audit

When clickhouse.use_cluster_queries: true, click-dog periodically checks whether more than one instance is running whole-cluster span reads: the sidecar + use_cluster_queries anti-pattern that causes N× duplicate exports. It is a hard no-op unless this instance itself uses cluster queries (zero cost for the default per-node sidecar topology), and is observability-only: it emits the click_dog_topology_warning{reason} gauge, the /status topology_warning field, and a throttled WARN, but never affects /readyz or /healthz.

Configuration path Default Description
monitor.topology_audit.enabled true Enable the self-audit (still a no-op unless use_cluster_queries is on)
monitor.topology_audit.interval_s 300 Audit tick interval in seconds. Negative values fail validation; 0 loads as the default when read from YAML
monitor.topology_audit.debounce_count 2 Consecutive positive ticks before a warning latches (absorbs rollout/restart transients). Negative values fail; 0 loads as the default
monitor.topology_audit.query_log_lookback_minutes 15 Window of system.query_log scanned for distinct cluster readers. Negative values fail; 0 loads as the default

First detection lands roughly interval_s × debounce_count after a misconfig appears: ~10 minutes at the defaults. Lower interval_s for faster detection. The reason label is sidecar_cluster_queries when this instance's clickhouse.host is loopback (co-located sidecars) and multi_instance_cluster_queries otherwise (separate readers not sharing an election).

Keep interval_s well above the poll cadence. The auditor counts a host as a live cluster reader only if its most recent whole-cluster read falls inside a recency window it derives from monitor.check_interval_s, capped at interval_s / 2. If you raise check_interval_s past interval_s / 2, that cap drops the window below the poll cadence, so a concurrent reader that only reads once per check_interval_s can fall outside it between ticks. In that case, the misconfig goes undetected, or the warning flaps. Keep interval_s ≥ 6 × check_interval_s to preserve the full window. click-dog emits a load-time WARN when this combination is detected.

Connection Pool

Configuration path Default Description
clickhouse.max_open_conns 2 Maximum number of open connections
clickhouse.max_idle_conns 1 Maximum number of idle connections
clickhouse.query_timeout_s 30 Query timeout in seconds. Prevents runaway queries
clickhouse.max_memory_usage 104857600 Per-query memory limit in bytes (100 MB). Prevents a single monitoring query from consuming excessive memory

OTEL Exporter

exporters:
  otel:
    - collector_address: localhost:4317  # OTEL collector gRPC address
      service_name: click-dog-monitor     # Service name in traces

      max_query_length: 100000     # Truncate queries longer than this (default: 100000)

      # TLS
      secure: false                # Enable TLS for gRPC (default: false)
      insecure_skip_verify: false  # Skip cert verification (TESTING ONLY)
      ca_cert: ""                  # Path to CA cert
      client_cert: ""              # Client cert for mTLS
      client_key: ""               # Client key for mTLS
Configuration path Default Description
exporters.otel[].collector_address None Required for every entry. OTEL collector gRPC endpoint (e.g., localhost:4317). Supports env expansion
exporters.otel[].service_name click-dog-monitor Service name that appears in traces. Supports env expansion
exporters.otel[].max_query_length 100000 Truncate SQL query text longer than this in exported spans. 0 or omitted is coerced to the 100000 default; the limit is not disabled at config load
exporters.otel[].secure false Enable TLS for the gRPC connection. When false, configured CA or complete mTLS material is ignored and produces a validation warning
exporters.otel[].insecure_skip_verify false Skip TLS cert verification. Insecure. Requires secure: true; otherwise validation fails
exporters.otel[].ca_cert None CA certificate path for TLS verification. Supports env expansion and must be readable at config load when secure: true
exporters.otel[].client_cert None Client certificate for mutual TLS (mTLS). Supports env expansion and must be used with client_key, regardless of secure; it must be readable at config load when secure: true
exporters.otel[].client_key None Client key for mTLS. Supports env expansion and must be used with client_cert, regardless of secure; it must be readable at config load when secure: true

Plaintext transport and raw-query compatibility defaults

With secure: false (the default), span export runs over unencrypted gRPC, and click-dog logs a startup warning. filters.query_text_mode also defaults to raw, so query literals can leave the host in db.statement. Set secure: true for any non-local collector and use normalized_only or none for privacy-sensitive production environments. Use redacted only with rules broad enough for the statements you intend to retain; unmatched statements are omitted.

Authentication

Click-Dog supports TLS and mTLS for the gRPC connection. It does not currently support bearer tokens, API keys, or custom gRPC metadata headers. If your OTEL collector requires header-based auth, place an authenticating proxy in front of it or use mTLS for client authentication.

Multiple Export Backends

The exporters section accepts any number of OTLP/gRPC sinks under exporters.otel[], and can be combined with Splunk HEC under exporters.splunk_hec[]:

exporters:
  otel:
    - collector_address: primary-collector:4317
      service_name: click-dog-monitor
    - collector_address: backup-collector:4317
      service_name: clickdog-backup
  splunk_hec:
    - endpoint: https://splunk.internal:8088
      token: ${SPLUNK_HEC_TOKEN}
      index: clickhouse
      source: click-dog
      source_type: _json
      allow_insecure_http: false
      max_query_length: 100000

Each exporters.splunk_hec[] entry takes a Splunk HTTP Event Collector endpoint, a token (env-expandable), and optional index, source, source_type, max_query_length, and TLS fields. Use https:// endpoints by default because the HEC token is sent in the Authorization header; http:// endpoints are accepted only when allow_insecure_http: true is set explicitly for local, development, or test HEC receivers. The token can also be supplied through token_file, whose path supports ${VAR} expansion. One trailing CRLF or LF is trimmed at load. token and token_file are mutually exclusive.

Configuration path Default Description
exporters.splunk_hec[].endpoint None Required. HEC base URL. Must use https:// unless allow_insecure_http is explicitly enabled; click-dog appends the event path as described in the Splunk HEC guide
exporters.splunk_hec[].token None Required unless token_file resolves a nonempty token. Inline or env-expanded HEC token. Mutually exclusive with token_file
exporters.splunk_hec[].token_file None Env-expandable path to the required HEC token file. Mutually exclusive with token; an empty file is rejected
exporters.splunk_hec[].index None Optional target index
exporters.splunk_hec[].source click-dog Event source
exporters.splunk_hec[].source_type _json Event source type
exporters.splunk_hec[].allow_insecure_http false Permit an http:// endpoint for local development or tests
exporters.splunk_hec[].insecure_skip_verify false Skip TLS certificate verification for an https:// endpoint. Insecure. Requires an https:// endpoint; otherwise validation fails
exporters.splunk_hec[].max_query_length 100000 Truncate SQL in HEC events

When more than one exporter is configured, spans are sent to all backends concurrently under the default PolicyAllRequired contract. Any backend failure, whether partial or total, is an export error. Starting each call together gives every backend the full monitor.export_timeout_s interval; result aggregation and per-sink statuses remain in configuration order. The cycle is marked an error (click_dog_cycle_results_total{result="error"}), feeds the circuit breaker and adaptive backoff, and shows up as the most recent error on /status and /readyz. The per-sink counters (click_dog_export_{attempts,accepted,errors}_total{sink}) and warning log summaries show which backend failed, so a dead sink cannot silently drop out of the fan-out.

A span is marked as "seen" in the dedup cache only if every backend accepts it. If any backend fails, no span from that batch is marked seen, and the next cycle re-delivers the full batch to every sink. Backends that already accepted those spans will see the same (trace_id, span_id) pair again. Stable IDs make duplicates discoverable, but whether they are collapsed, stored, or rejected is backend-specific. The same contract applies to backfill mode: a partial multi-sink failure counts the query as Failed in the run summary, the process exits non-zero, and re-running the window re-delivers to every sink.

An alternate PolicyAnySuccess policy exists in code: spans are marked seen as soon as one sink accepts them, while failed sinks remain visible through ExportResult and the per-sink metrics. It is not exposed via YAML today; production configs use the strict default policy. The per-call deadline applied to every backend in the fan-out is monitor.export_timeout_s (see Monitor below).


High Availability

Leader election via ClickHouse Keeper for multi-node deployments. Configuring keeper.hosts is what turns election on: there is no separate enable flag. Omit the ha block entirely and each instance runs standalone.

ha:
  keeper:
    hosts:
      - keeper-01:9181
      - keeper-02:9181
      - keeper-03:9181
    secure: false                 # Enable TLS to Keeper (default: false)
    session_timeout_s: 10        # Session timeout (default: 10, range: 5-30)
    base_path: /click-dog/election  # Znode path prefix
    # auth_user: ""              # Optional digest auth
    # auth_password: ""          # Optional digest auth
Configuration path Default Description
ha.keeper.hosts None ClickHouse Keeper (ZooKeeper-compatible) endpoints. Presence enables leader election; empty/omitted means standalone
ha.keeper.secure false Enable TLS for the Keeper connection. TLS uses the host's system trust roots and supports neither a configured custom CA nor client-certificate mTLS; private-CA Keeper requires installing that CA into the service host/container trust store. Use TLS with Keeper digest auth on untrusted networks so credentials are not sent in plaintext
ha.keeper.session_timeout_s 10 With nonempty keeper.hosts, 0/omitted becomes 10 and effective values must be 5–30. With no hosts, election is inactive and this field is ignored and not range-validated
ha.keeper.base_path /click-dog/election Znode path prefix. Must be exactly /click-dog or under /click-dog/; paths overlapping ClickHouse's own Keeper paths (/clickhouse...) are rejected
ha.keeper.auth_user None Optional digest authentication username (supports ${ENV} expansion)
ha.keeper.auth_password None Optional digest authentication password (supports ${ENV} expansion). Mutually exclusive with keeper.auth_password_file
ha.keeper.auth_password_file None Path to a file holding the digest auth password (read at load; one trailing CRLF or LF is trimmed). Path supports ${VAR} expansion. Mutually exclusive with keeper.auth_password

Each click-dog instance creates an ephemeral sequential znode. The lowest sequence number becomes leader. When the leader dies, its session expires, the znode is deleted, and the next candidate auto-promotes (~10 seconds failover).

When keeper.auth_user is omitted, election znodes use Keeper world ACLs; click-dog emits a startup warning so that choice is visible. Configure digest auth for shared Keeper ensembles, and set keeper.secure: true when Keeper supports TLS.

Leader gating depends on topology

In cluster mode (use_cluster_queries: true), leadership gates the data path: only the leader fetches and exports; standbys skip each cycle (recorded as a skipped cycle, no export) so the cluster's spans aren't duplicated. In sidecar mode each instance reads its own node's disjoint local scope and always exports regardless of leader status: there leadership only drives coordination duties (flush, /clusterz). Either way, configuring ha.keeper.hosts is what activates election.

Keeper startup behavior depends on how far initialization gets:

  • With resolvable Keeper addresses and no digest authentication, the pinned ZooKeeper client starts network dialing asynchronously. The instance fails open and keeps exporting while the election loop retries, then joins automatically when Keeper becomes reachable.
  • If election construction returns an error, click-dog logs Failed to join leader election and leaves that process standalone until restart. Host parsing or DNS resolution can take this path. Initial AddAuth is synchronous, so an authentication error or an unavailable Keeper while auth_user is configured can take it too.

Monitor

monitor:
  enabled: true                # Enable scheduled monitoring (default: true)

  # Duration thresholds (milliseconds)
  min_trace_duration_ms: 1000  # Find traces with spans >= this (required)
  min_span_duration_ms: 0      # Only export spans >= this (0 = all)
  max_trace_duration_ms: 0     # Skip traces with spans > this (0 = no limit)
  max_span_duration_ms: 0      # Skip spans > this (0 = no limit)

  max_query_length: 100000     # Skip queries with SQL > this chars (see also exporters.otel[].max_query_length which truncates instead)

  # Polling
  check_interval_s: 30         # Polling interval in seconds
  lookback_s: 40               # How far back to look (default: check_interval_s + lookback_buffer_s)
  lookback_buffer_s: 10        # Extra seconds added when defaulting lookback_s (default: 10)

  # Rate limiting
  max_spans_per_cycle: 1000    # Max spans fetched and exported per polling cycle (default: 1000)

  # Batch processing
  batch_size: 0                # Batch size (0 = no batching)
  batch_delay_ms: 0            # Delay between batches (0 = no delay)
  export_timeout_s: 30         # Per-call exporter deadline in seconds (0 disables)

  # Deduplication
  dedup_cache_size: 10000      # LRU cache entries (default: 10000)

  # Degraded-mode canary
  canary:
    enabled: false
    threshold_duration_ms: 60000

Duration Filtering

Click-Dog uses a two-tier duration filtering approach:

  1. Trace-level (min_trace_duration_ms / max_trace_duration_ms): Find traces that contain at least one span within this duration range
  2. Span-level (min_span_duration_ms / max_span_duration_ms): From those traces, only export individual spans within this duration range

All duration filtering happens in SQL for maximum efficiency.

Configuration path Default Description
monitor.enabled true Enable scheduled monitoring. When true, min_trace_duration_ms and check_interval_s must both be greater than zero. Set false for deliberately backfill-only usage
monitor.min_trace_duration_ms None Required. Find traces containing spans >= this duration (ms)
monitor.min_span_duration_ms 0 Only export spans >= this (ms). 0 = export all spans from matching traces
monitor.max_trace_duration_ms 0 Skip traces with spans > this (ms). 0 = no upper limit
monitor.max_span_duration_ms 0 Skip individual spans > this (ms). 0 = no upper limit
monitor.max_query_length 100000 Skip queries with SQL text longer than this (chars) and truncate live query_log.normalized_query previews to this length. 0 or omitted is coerced to the 100000 default; the limit is not disabled at config load. Different from exporters.otel[].max_query_length which truncates at export time

Polling

Configuration path Default Description
monitor.check_interval_s 30 How often to poll ClickHouse (seconds). Must be greater than 0 when enabled is true
monitor.lookback_s check_interval_s + lookback_buffer_s How far back to look for spans. Defaults to interval + the buffer below, to avoid gaps between polls. Once spans have been exported, a later cycle resumes from the last exported position even when it has fallen behind this window, up to three lookback windows back (120s at the defaults, which covers a failed export followed by the first backoff step); spans further behind are dropped. A process with nothing exported yet starts at the plain window. Once a cycle has read everything new, spare max_spans_per_cycle goes to re-reading already-exported spans in the window for rows that flushed late; the dedup cache skips the repeats. Set this explicitly to override the computed default
monitor.lookback_buffer_s 10 Seconds added to check_interval_s when computing the default lookback_s. Absorbs clock skew and slow poll cycles; increase under high load if you see missed spans. Ignored when lookback_s is set explicitly

Rate Limiting & Batching

Configuration path Default Description
monitor.max_spans_per_cycle 1000 Maximum number of spans fetched and exported per polling cycle. Applied as a SQL LIMIT on the spans query (and as an upper bound on each trace-ID page, since each trace contributes at least one span). Each cycle continues through the lookback window oldest first from where the last successfully exported cycle stopped, so the cap is spent on spans not yet exported, and a trace with more spans than the cap is finished over the following cycles. The walk only moves once a cycle's exports succeed. Inflow below max_spans_per_cycle / check_interval_s is read in full; a trace the walk falls more than three lookback windows behind is dropped, which is what eventually happens to the oldest spans when inflow stays above that rate. A debug-level log line reports the total span count per cycle. The absolute ceiling is 100000 spans per cycle to bound memory; config validation rejects larger values at load
monitor.batch_size 0 Process spans in batches of this size. 0 = process all at once
monitor.batch_delay_ms 0 Milliseconds to wait between batches. 0 = no delay
monitor.export_timeout_s 30 Per-call deadline for exporter calls. Applies to live span batches, backfill query exports, and degraded-mode canary span exports so a stuck collector fails into the normal retry, backoff, and circuit-breaker path. Set to 0 to disable the client-side export deadline
monitor.dedup_cache_size 10000 Size of the LRU cache for span deduplication, keyed by the composite OTLP span identity (trace_id, span_id). OTLP span IDs are only unique within a trace, so dedup must include the trace ID; a span-id-only cache would silently drop sibling spans in different traces that collide on the 64-bit ID. Each entry's SpanKey is 24 B (16 B trace_id + 8 B span_id), but the LRU container adds doubly-linked-list pointers and a map bucket, so the practical cost is ~80–100 B per entry on 64-bit. 10,000 entries is approximately 1 MiB. Memory scales linearly. The cache is in-memory only; a restart can resend spans still inside lookback_s, and downstream duplicate handling is backend-specific
monitor.extract_log_comment true Extract log_comment from URI attributes and promote JSON keys as span attributes prefixed with log_comment.. See Span Attributes
monitor.enrich_from_query_log true Enrich spans with metadata from system.query_log (user, client, tables, query stats) by joining on query_id (from the clickhouse.query_id span attribute). Note: includes query_log.user, query_log.client_address, and query_log.client_hostname, which may contain PII or internal network information. See the privacy note and Span Attributes

Degraded-mode canary

Configuration path Default Description
monitor.canary.enabled false Run and export a lightweight canary query while the circuit breaker is open or adaptive backoff is elevated
monitor.canary.threshold_duration_ms 60000 Count ClickHouse spans exceeding this duration when constructing the canary span. Must not be negative

The canary queries ClickHouse and exports the result, so it exercises both sides even when the breaker was tripped by export failures. When a circuit breaker is configured, that success is reported to it — but it counts toward success_threshold only once the breaker has reached half-open. While the breaker is still open, the report clears the failure count without advancing recovery. With monitor.circuit_breaker.enabled: false the canary can still run on elevated backoff alone, and there is no breaker to report to.

See Troubleshooting · Canary Queries in Degraded Mode for emitted attributes and failure behavior.


Circuit Breaker

Protects ClickHouse by stopping queries after repeated failures.

monitor:
  circuit_breaker:
    enabled: false             # Enable circuit breaker (default: false)
    failure_threshold: 3       # Open after N consecutive failures
    success_threshold: 1       # Close after N successes in half-open state
    reset_timeout_s: 60        # Try again after N seconds when open
Configuration path Default Description
monitor.circuit_breaker.enabled false Enable the circuit breaker. Note: the install script enables this by default in generated configs
monitor.circuit_breaker.failure_threshold 3 Number of consecutive failures before opening the circuit
monitor.circuit_breaker.success_threshold 1 Number of successful requests needed in half-open state to close the circuit
monitor.circuit_breaker.reset_timeout_s 60 Seconds to wait before transitioning from open to half-open

See Resilience for details on circuit breaker behavior.


Adaptive Backoff

Increases polling interval when errors occur, reducing load on a struggling ClickHouse.

monitor:
  backoff:
    enabled: false             # Enable adaptive backoff (default: false)
    max_interval_s: 300        # Maximum polling interval (default: 300 = 5 min)
    backoff_factor: 2.0        # Multiply interval on failure (default: 2.0)
Configuration path Default Description
monitor.backoff.enabled false Enable adaptive backoff. Note: the install script enables this by default in generated configs
monitor.backoff.max_interval_s 300 Maximum polling interval in seconds (5 minutes)
monitor.backoff.backoff_factor 2.0 Multiply the current interval by this factor on each failure. Must be greater than 1; values in (0, 1] are rejected at config load

See Resilience for details on backoff behavior.


Filters

filters:
  query_text_mode: redacted     # raw | redacted | normalized_only | none
  whitelist_operations: []     # Operation names to INCLUDE (supports * wildcard)
  blacklist_operations:        # Operation name substrings to EXCLUDE at SQL level
    - "MergeTreeIndex"
    - "VFSWrite"
  whitelist_ips: []            # IP addresses to INCLUDE
  whitelist_users: []          # ClickHouse users to INCLUDE (strict: drops spans with unknown user)
  blacklist_users: []          # ClickHouse users to EXCLUDE (defense-in-depth alongside CH user grants)
  blacklist_queries:           # Regex patterns to EXCLUDE
    - "^SELECT \\* FROM system\\."
    - "SHOW TABLES"
    - "(?i)healthcheck"
  redact_queries:              # Regex replacements applied before export
    - pattern: "(?i)identified\\s+by\\s+'[^']*'"
      replacement: "IDENTIFIED BY '[REDACTED]'"
Configuration path Default Description
filters.query_text_mode raw Final query-text representation applied after filtering/enrichment and before exporter fan-out. normalized_only and none never fall back to raw text and are recommended for privacy-sensitive production environments
filters.whitelist_operations [] If set, only export spans with matching operation names. Supports * wildcard (e.g., DB::Interpreter*::execute())
filters.blacklist_operations [] Substring patterns matched against operation_name. Pushed down to ClickHouse as NOT LIKE '%pattern%' so matching spans are never fetched. Use this to drop high-volume internal pipeline spans (MergeTreeIndex, VFSWrite, etc.) without affecting query-level spans
filters.whitelist_ips [] If set, only export spans from these client IP addresses. Entries may be exact IPs (10.0.1.50) or CIDR ranges (10.0.0.0/8)
filters.whitelist_users [] If set, only export spans whose query_log.user is on the list. Pushed down to enrichment SQL as AND user IN (...) and re-checked per span. This is strict: spans with an unresolvable user are dropped
filters.blacklist_users [] Never export spans whose query_log.user is on the list. Enforced at the Go layer for the span path (pushing it to enrichment SQL would strip the user and let blacklisted spans bypass the check); pushed down to SQL on the slow-query / backfill path
filters.blacklist_queries [] Regex patterns matched against SQL query text. Matching spans are excluded
filters.redact_queries[].pattern None Regex pattern matched against SQL query text after blacklist filtering and before export. Used only with query_text_mode: redacted; at least one valid rule is required. An unmatched statement is omitted rather than exported raw. Rules under raw are invalid; stale rules under normalized_only/none are ignored with a warning
filters.redact_queries[].replacement [REDACTED] Replacement text used for matches. Empty or omitted uses the default replacement

Note: blacklist_operations runs at the SQL level before data is fetched. whitelist_operations runs after fetch in Go. The blacklist always takes precedence because matching spans never reach the whitelist check.

See Filtering for detailed examples and evaluation order.


Logging

log_level: info                # debug, info, warn, error (default: info)
log_file: ""                   # Log file path (default: stderr)
log_format: text               # text or json (default: text)
log_rotation:
  max_size_mb: 100             # Rotate when the active file passes this size (default: 100; must be >= 1)
  max_files: 3                 # Keep this many rotated files (default: 3; must be >= 1)
Configuration path Default Description
log_level info Log verbosity: debug, info, warn, error
log_file None Optional file path. If empty, logs go to stderr (container-friendly)
log_format text Output format: text (human-readable) or json (structured). JSON emits {"ts":"...","level":"...","msg":"..."}
log_rotation.max_size_mb 100 Rotation threshold per file. Omit the key to take the default; an explicit 0 is rejected (would disable rotation and produce unbounded growth).
log_rotation.max_files 3 Number of rotated files to retain. Must be >= 1.

Log file permissions

When log_file is set, click-dog creates the file with mode 0640 (owner: read/write, group: read, world: none). Operational metadata such as client IPs, operation names, and error traces is sensitive enough to keep off world-readable storage but the service group still needs read access for log shipping.

If a separate log-shipping agent runs alongside click-dog (Filebeat, Fluent Bit node agent, journald forwarder, Vector, Promtail, etc.) under a different UID, that agent must be in the click-dog service user's group for log shipping to work. Otherwise the agent will silently fail to read rotated files. Typical setup:

# Add the log shipper's user to click-dog's group
usermod -aG click-dog filebeat
systemctl restart filebeat

Earlier releases created the log file with mode 0644 (world-readable). Upgrading deployments that relied on an unprivileged log shipper reading the file via the world bit will need to adjust group membership as above.


Metrics

metrics:
  enabled: false                          # Enable scrape + admin endpoints
  listen_address: ":9090"                 # Scrape listener: /metrics + legacy /health (default: :9090)
  admin_listen_address: "127.0.0.1:9091"  # Admin listener: POST /flush (default: 127.0.0.1:9091)
  otlp:
    enabled: false                        # Optional push path for self-metrics
    inherit_otel_connection: true          # Reuse exporters.otel[0] when enabled
    interval_seconds: 10
Configuration path Default Description
metrics.enabled false Start both the scrape and admin HTTP listeners. This setting gates both; false also disables HTTP POST /flush (SIGUSR1 still works)
metrics.listen_address :9090 Scrape listener (/metrics, legacy /health); safe to expose to your scraper
metrics.admin_listen_address 127.0.0.1:9091 Admin listener (POST /flush); loopback by default. Non-loopback binds emit a startup WARN
metrics.otlp.enabled false Push click-dog self-metrics over OTLP. Additive to the /metrics scrape endpoint; it may run with metrics.enabled: false
metrics.otlp.inherit_otel_connection true when OTLP metrics are enabled and the key is omitted Reuse exporters.otel[0] connection settings for the self-metrics exporter. Requires at least one OTLP trace exporter when enabled
metrics.otlp.collector_address None Standalone OTLP metrics collector address, required when inherit_otel_connection: false. Supports env expansion
metrics.otlp.secure false Enable TLS for a standalone OTLP metrics connection. Use inheritance when CA or mTLS settings are needed
metrics.otlp.host OS hostname Host identity used on pushed self-metrics. Supports env expansion
metrics.otlp.service_name exporters.otel[0].service_name, then click-dog-monitor Service name used on pushed self-metrics. Supports env expansion
metrics.otlp.interval_seconds 10 Push interval. Must be greater than 0 when otlp.enabled is true
metrics.otlp.rename {} Optional map from canonical self-metric keys to emitted OTLP metric names. Keys are validated at load time

The metrics: and health: blocks are independent. metrics.enabled does not control the Kubernetes-shaped /healthz / /readyz / /status probes. Those live in the separate health: block below. Disabling metrics will not turn off the health listener, and vice versa.

See Observability for the full endpoint table, the rationale for splitting scrape from admin, and deployment recipes (systemd / Docker / Kubernetes).


Health Endpoints

health:
  enabled: false              # Off by default; install.sh enables it on :8686
  listen_address: ":8686"     # /healthz, /readyz, /status (default when enabled)
  cluster:
    enabled: false            # Leader-only /clusterz aggregate endpoint
    self: ""                  # Routable host:port for this instance
    peer_timeout_ms: 3000
    peers: []
Configuration path Default Description
health.enabled false Start the health listener. Independent of metrics.enabled
health.listen_address :8686 when health.enabled: true /healthz (liveness), /readyz (readiness; pings ClickHouse), /status (JSON report). If this normalizes to the same socket as metrics.listen_address, the handlers are mounted on the metrics listener so only one port is opened
health.cluster.enabled false Mount leader-only /clusterz aggregate health when health.enabled is also true
health.cluster.self None Routable bare host:port for this instance's /readyz; required when cluster health is enabled. Do not include http://, a path, query, or fragment
health.cluster.peer_timeout_ms 3000 Per-peer /readyz fanout deadline
health.cluster.peers [] Static list of bare peer host:port addresses to include in /clusterz. Empty entries, schemes/paths, missing ports, and duplicate peers are rejected

/healthz always returns 200; /readyz returns 503 while ClickHouse is unreachable or the circuit breaker is open. /status is a reporting endpoint that always returns 200: assess health from its JSON body, not its status code.

See Observability: Health Endpoints for the full endpoint table, the /status JSON shape, Kubernetes probe recipes, and cluster-aware health roll-up.


Webhook Notifications

webhook:
  enabled: false
  url: ${SLACK_WEBHOOK_URL}    # Webhook URL (supports env vars)
  timeout_s: 10                # HTTP timeout in seconds (default: 10)
  events:                      # Event filter (default: all events)
    - analysis_findings
    - circuit_breaker_opened
    - circuit_breaker_closed
    - startup
    - shutdown
Configuration path Default Description
webhook.enabled false Enable webhook notifications
webhook.url None HTTP POST endpoint. Supports ${VAR} expansion. Must be http:// or https:// with a host; redirects are not followed
webhook.timeout_s 10 HTTP request timeout in seconds
webhook.events [] (all) List of event names to fire on. Empty = all events

Valid events: analysis_findings, circuit_breaker_opened, circuit_breaker_closed, backfill_complete, backfill_failed, error_spike, startup, and shutdown. analysis_findings is evaluated only by an explicit click-dog analyze queries --notify run; enabling the destination does not make scheduled monitoring run query analysis.

See Observability for payload format, behavior details, and compatible services.


Datadog Event Management

datadog_events:
  enabled: false
  site: datadoghq.com
  api_key: ${DD_API_KEY}             # or api_key_file
  # api_key_file: /run/secrets/datadog-api-key
  application_key: ${DD_APP_KEY}     # or application_key_file
  # application_key_file: /run/secrets/datadog-app-key
  timeout_s: 10
  environment: production
  service: click-dog
Configuration path Default Description
datadog_events.enabled false Enable this destination for explicit analyze queries --notify runs
datadog_events.site None Trusted Datadog site hostname used to derive https://event-management-intake.<site>/api/v2/events; a scheme, path, port, query, fragment, IP address, or malformed DNS hostname is rejected
datadog_events.api_key None Datadog API key. Supports ${VAR} expansion; mutually exclusive with datadog_events.api_key_file
datadog_events.api_key_file None File containing the API key; exactly one trailing LF or CRLF is trimmed and other whitespace is preserved; mutually exclusive with datadog_events.api_key
datadog_events.application_key None Datadog application key. Supports ${VAR} expansion; mutually exclusive with datadog_events.application_key_file
datadog_events.application_key_file None File containing the application key; exactly one trailing LF or CRLF is trimmed and other whitespace is preserved; mutually exclusive with datadog_events.application_key
datadog_events.timeout_s 10 Single-attempt request timeout, from 1 through 60 seconds
datadog_events.environment None Required low-cardinality environment:<value> tag; letters, digits, ., _, :, /, and -, at most 200 characters
datadog_events.service None Required low-cardinality service:<value> tag with the same character and length rules

When enabled, site, environment, service, and one source for each credential are required. Both credentials are required because the Events API v2 request is authenticated with DD-API-KEY and DD-APPLICATION-KEY. Secret values are never included in validation or delivery errors. The client performs one bounded request, follows no redirects, and treats only HTTP 2xx as intake acceptance. It does not retry, and acceptance is not confirmation that an Event Monitor evaluated or notified.

This destination handles the bounded query-analysis notification summary. It does not replace or share the OTLP trace or self-metric exporters. See Datadog integration for setup and Event Monitor guidance.


Complete Example

Start from the complete minimal profile, or see config.yaml.example in the repository root for a fully commented configuration file with all options.