Skip to content

Operation Modes

Click-Dog supports two long-running modes: scheduled (continuous monitoring) and backfill (one-shot historical export): plus short-lived inspection modes (validate, dry-run, and click-dog analyze), operational commands (click-dog test, test-span, and flush), and a separate click-dog deploy subcommand family for generating deployment artifacts and managing the state of a standard systemd installation. Scheduled and backfill can run simultaneously as two separate processes.

Scheduled Mode

Continuously polls ClickHouse's system.opentelemetry_span_log table at a configurable interval and exports matching spans to your configured OTEL gRPC and/or Splunk HEC sinks.

How It Works

  1. Queries system.opentelemetry_span_log for trace IDs containing spans within the configured duration range
  2. Fetches all spans for those traces (optionally filtered by span-level duration)
  3. Deduplicates against an LRU cache to avoid re-exporting
  4. Applies operation, IP, and query filters
  5. Exports to the configured OTEL gRPC and/or Splunk HEC sinks
  6. Waits check_interval_s seconds, then repeats

Usage

# With default config resolution (/etc/click-dog/click-dog.yaml first,
# then ./click-dog.yaml)
./click-dog

# With custom config
./click-dog --config /path/to/config.yaml

Configuration

monitor:
  enabled: true
  min_trace_duration_ms: 1000  # Find traces with spans >= 1 second
  min_span_duration_ms: 0      # Export all spans from those traces
  check_interval_s: 30         # Poll every 30 seconds
  lookback_s: 40               # Look back 40 seconds each poll

Key Behaviors

  • Runs on startup: the first poll happens immediately, then repeats at check_interval_s
  • Deduplication: an LRU cache (default 10,000 entries) keyed by the composite OTLP span identity (trace_id, span_id) prevents re-exporting the same spans across polling cycles. Span IDs are only unique within a trace, so the trace ID is part of the key
  • Lookback overlap: set lookback_s slightly larger than check_interval_s (the default adds 10 seconds) to avoid gaps between polls
  • Graceful shutdown: responds to SIGINT and SIGTERM by cancelling any in-progress polling cycle, including its ClickHouse query, batch delay, or exporter call. Exporter connections are then closed with a 5-second timeout. The in-memory dedup cache is not persisted. After restart, spans still within the lookback window can be re-exported once. Stable (trace_id, span_id) values make repeats identifiable, but downstream duplicate handling is backend-specific.
  • Resilience: supports circuit breaker and adaptive backoff to protect ClickHouse during failures

Data Source

Scheduled mode reads from system.opentelemetry_span_log, which contains OpenTelemetry spans generated internally by ClickHouse. This includes: - Query execution spans - Internal ClickHouse operations (merges, mutations, etc.) - Distributed query coordination spans

Before scheduled spans reach any exporter, filters.query_text_mode applies the final raw/redacted/normalized-only/none representation. Filtering may inspect raw SQL first; multi-sink fan-out receives only the shaped span. The three restrictive modes also remove URI attribute values carrying HTTP query parameters. See Query Text Export Modes.


Backfill Mode

One-time export of historical data from ClickHouse's system.query_log table for a specific time range.

How It Works

  1. Queries system.query_log for queries in the specified time range matching the duration threshold
  2. Filters by QueryFinish and ExceptionWhileProcessing event types
  3. Applies configured filters to the original in-process query
  4. Applies filters.query_text_mode without raw fallback (including removing URI attribute copies carried in HTTP query parameters)
  5. Exports each shaped query as a trace/event to the configured OTEL gRPC and/or Splunk HEC sinks
  6. Exits when complete

Usage

# Backfill a specific time range
./click-dog backfill \
  --config config.yaml \
  --start "2024-01-15T00:00:00Z" \
  --end "2024-01-15T23:59:59Z"

Time format must be RFC3339 (e.g., 2024-01-15T14:30:00Z). The pre-subcommand spelling — -backfill-start / -backfill-end mode flags on the root command — still works but is deprecated.

Configuration

monitor:
  enabled: false               # Disable scheduled mode (optional)
  min_trace_duration_ms: 1000  # Only export queries >= 1 second
  max_spans_per_cycle: 5000    # Limit total queries per run
  batch_size: 100              # Process in batches of 100
  batch_delay_ms: 50           # 50ms between batches

Key Behaviors

  • Exits when done: unlike scheduled mode, backfill runs once and exits
  • No deduplication cache: since it is a one-time run, there is no need for the LRU cache
  • Rate controllable: use max_spans_per_cycle, batch_size, and batch_delay_ms to control the export rate
  • Different data source: reads from system.query_log (not system.opentelemetry_span_log), which provides richer query metadata
  • Different trace shape: creates one synthetic span per query-log row. It does not recreate native ClickHouse trace topology, child spans, or span-log-only attributes
  • Filters applied: IP whitelist, operation whitelist, query blacklist, user filters, and query redaction use query_log fields. Because query_log entries have no operation name, a configured whitelist_operations rejects every backfill query (the sole exception is a pattern that matches the empty string, such as "*"). Leave whitelist_operations empty when using backfill
  • Exit code reflects export outcome: backfill exits 0 only when every non-filtered query exports successfully. Any export failure exits non-zero. The final log line and the backfill_failed webhook (if enabled) include queries=N exported=N filtered=N failed=N so an operator can tell whether the failure was partial (some data landed) or total (none did) and decide how to recover. The terminal backfill_complete or backfill_failed webhook attempt completes or reaches its configured timeout before the one-shot process exits

Data Source

Backfill mode reads from system.query_log, which contains finished query events. This provides attributes not available in the span log, such as: - Read/write row and byte counts - Memory usage - Databases and tables accessed - Exception codes


Running Both Modes Together

You can run scheduled monitoring and backfill simultaneously as two separate processes. This is useful for: - Recovering from an outage while continuing real-time monitoring - Backfilling a historical period for investigation

# Terminal 1: Continuous monitoring
./click-dog --config config.yaml

# Terminal 2: Backfill a specific incident window
./click-dog backfill --config config.yaml \
  --start "2024-01-15T10:00:00Z" \
  --end "2024-01-15T11:00:00Z"

Chunked Backfill

For large time ranges, split the backfill into smaller chunks to control load:

# Process hour by hour
./click-dog backfill --config config.yaml \
  --start "2024-01-15T00:00:00Z" \
  --end "2024-01-15T01:00:00Z"

./click-dog backfill --config config.yaml \
  --start "2024-01-15T01:00:00Z" \
  --end "2024-01-15T02:00:00Z"

# ... and so on

Validate Mode

Validates the configuration file and prints parsed settings without connecting to anything. Useful for checking config before deployment.

Usage

./click-dog validate --config config.yaml

The pre-subcommand spelling, a -validate mode flag on the root command, still works but is deprecated.

Output

Prints parsed settings including: - Any deprecation, environment, and validation warnings from config load - ClickHouse connection details - Configured exporters (count, plus one line per OTEL or Splunk HEC sink) - Monitor settings (minimum trace duration and check interval) and the effective query-text export mode - HA leader-election status when configured - Self-metrics (OTLP push) state, whether enabled or disabled, plus the metrics, health, and webhook listeners when enabled

Exits with code 0 if valid, non-zero if there are errors.


Dry-run Mode

Runs one cycle of the pipeline against your real ClickHouse: same queries, same filters, same dedup, then prints a cumulative summary (traces, spans, queries, top operations) and exits. Routes exports through DryRunExporter, so nothing is sent to OTEL or Splunk HEC.

Use it to: - preview what a new filter or duration threshold will export before pointing the binary at a real collector; - diagnose "why isn't this span being exported?" without sending traffic to your observability bill; - smoke-test a config on a production cluster without writing data.

Usage

# Dry-run the current lookback window (one cycle, then exit)
./click-dog --dry-run --config click-dog.yaml

# Dry-run a historical window (backfill, then exit)
./click-dog backfill --dry-run --config click-dog.yaml \
  --start "2024-01-15T00:00:00Z" \
  --end   "2024-01-15T23:59:59Z"

Key Behaviors

  • Single cycle, then exits: scheduled --dry-run runs exactly the startup-check cycle (one pipeline.Process call), prints the summary, and returns. It does not enter the ticker loop. Backfill --dry-run runs the configured window, prints the summary, and exits
  • Real reads, no writes: ClickHouse is queried normally. Span exports are discarded before any OTEL / Splunk HEC traffic (exporter clients are constructed but their lazy connections never carry spans). One exception: with metrics.otlp.enabled: true, the self-metrics push still opens its collector connection as in a normal run
  • Summary shape: Traces, Spans, Queries, and a "Top operations" table (top 10 by count) are printed to stdout via DryRunExporter.PrintSummary
  • Sink label: Prometheus metrics emitted in the cycle carry sink="dry_run" so dashboards can distinguish dry-run traffic from live exports
  • No exports: the startup webhook notification still fires when webhooks are enabled, but the shutdown webhook does not fire because dry-run exits before the scheduled loop's signal-driven shutdown path

Want a continuous dry-run loop? That isn't supported today: scheduled --dry-run exits after one cycle. If you need rolling summaries, re-invoke periodically (cron, watch, etc.) or open a follow-up issue.


click-dog analyze (read-only reports)

click-dog analyze is a local, read-only inspection command family for query behavior. It produces reports and exits; it does not start the scheduled polling loop, open exporters, start metrics or health servers, start leader election, or run the circuit breaker.

Subcommand Purpose
click-dog analyze queries Build a deterministic query-analysis report; optionally capture/compare an explicit known-good baseline, apply a CI finding threshold, or explicitly notify configured analysis destinations after report output
click-dog analyze trace Drill from a current/recent query or explicit identity to native trace spans and nearby query-family context

See Query Analysis for the operator workflow and JSON artifact boundary, including --fail-on, --notify, and exit-code precedence.


Operational commands

Test exporter delivery

test export sends one synthetic span to every configured exporter without querying ClickHouse. Use it to isolate exporter credentials, TLS, and routing from the ClickHouse data plane.

sudo -u click-dog click-dog test export --config /etc/click-dog/click-dog.yaml

The span is named click-dog.test-span and carries click_dog.source=test-span, click_dog.test=true, and click_dog.test_kind=export. Each exporter is reported independently. A PASS means the exporter accepted the span; it does not prove backend indexing.

click-dog test-span remains a deprecated compatibility alias with the same flags, output behavior, and exit status. It prints the replacement command to stderr and has no scheduled removal date.

Test native tracing

test tracing proves the native ClickHouse path with one bounded, read-only query:

sudo -u click-dog click-dog test tracing \
  --config /etc/click-dog/click-dog.yaml \
  --timeout 30s

The command supplies a sampled parent context and a unique click-dog-test-... query ID to exactly SELECT 1 /* click-dog test tracing */, polls the span log by the generated trace ID, and requires a native span to reference the local parent. It then exports one batch containing the local parent followed by every recovered native child. The single timeout covers connection, query, polling, fetch, and export.

This exact-ID path intentionally bypasses scheduled duration, query, user, IP, and operation filters. It also exports the recovered span-log rows directly: the scheduled pipeline's query-log enrichment does not run, so this command does not validate log_comment extraction or table/database string-array attributes. It tests only the configured connection's selected node, not every cluster node or shard. A scheduled Click-Dog process can later read and re-export the native rows because deduplication is process-local; that possible duplicate does not change the test result.

Trigger an immediate cycle

sudo -u click-dog click-dog flush --config /etc/click-dog/click-dog.yaml

Without Keeper, the command sends SIGUSR1 to the local systemd process. With ha.keeper.hosts, it writes a shared request that the elected coordination leader consumes. A local SIGUSR1 or HTTP POST /flush always targets that specific process: a cluster-mode non-leader records a skipped cycle, while a sidecar runs its local cycle because sidecar collection is not leader-gated.

The poll timer resets once the flush cycle finishes, so the next scheduled cycle is a full interval away rather than arriving immediately behind it.

Flush is best-effort. A lost request does not change the next scheduled cycle.


click-dog deploy (deployment management)

click-dog deploy is a separate subcommand family that manages deployment artifacts and installed state rather than running the monitoring pipeline. It does not start the monitoring loop.

Subcommand Purpose
click-dog deploy kubernetes -c COLLECTOR --ch-host CH_SVC Render a Deployment + ConfigMap + Secret to ./click-dog-k8s/
click-dog deploy docker -c HOST Render a Docker Compose stack to ./click-dog-docker/
click-dog deploy status Report installed version, systemd active/enabled state, and the local /healthz response. Add --json for machine output
sudo click-dog deploy uninstall Confirm, then remove a standard systemd installation, including config and credentials

The Kubernetes and Docker generators take the same -c (collector address), -u / -p (ClickHouse credentials), -o (output directory), and -v (version) flags. Pass --update to refresh the image tag in an existing output directory without overwriting other fields.

click-dog deploy status exits 0 only when systemd reports active and /healthz returns 200; otherwise 1. Plain-text output is printed in both cases.

See the Kubernetes and Docker guides for the equivalent installer-wrapper recipes.


Comparison

Scheduled Backfill Validate Dry-run Analyze Test deploy
Data source system.opentelemetry_span_log system.query_log None Same as scheduled / backfill system.query_log, system.processes, and/or system.opentelemetry_span_log None (export) or exact trace ID in system.opentelemetry_span_log (tracing) None
Duration Runs continuously Runs once, then exits Exits immediately Runs once, then exits Exits after producing a report Exits after the bounded smoke test Exits immediately
Time range Rolling window (lookback_s) Fixed range (backfill --start/--end) N/A Same as the mode it shadows Bounded lookback window One generated trace, one-day partition floor N/A
Deduplication LRU cache prevents re-exports Not needed (one-time run) N/A Active during the single cycle N/A None; a scheduled process may later re-export native rows N/A
Resilience Circuit breaker + adaptive backoff None (single run) N/A Single cycle; no loop to back off N/A One command deadline; bounded polling N/A
Exports anything? Yes Yes No No; discards via DryRunExporter No Yes No
Use case Real-time monitoring Historical query visibility and query-log-based outage recovery Config verification Preview / smoke-test without sending Local query triage and report artifacts Exporter and native tracing go-live checks Generate manifests, inspect installed state, or remove a standard systemd installation