Operation Modes¶
Click-Dog supports two long-running modes: scheduled (continuous
monitoring) and backfill (one-shot historical export): plus short-lived
inspection modes (validate, dry-run, and click-dog analyze),
operational commands (click-dog test, test-span, and flush), and a
separate click-dog deploy subcommand family for generating deployment
artifacts and managing the state of a standard systemd installation.
Scheduled and backfill can run simultaneously as two separate processes.
Scheduled Mode¶
Continuously polls ClickHouse's system.opentelemetry_span_log table at a configurable interval and exports matching spans to your configured OTEL gRPC and/or Splunk HEC sinks.
How It Works¶
- Queries
system.opentelemetry_span_logfor trace IDs containing spans within the configured duration range - Fetches all spans for those traces (optionally filtered by span-level duration)
- Deduplicates against an LRU cache to avoid re-exporting
- Applies operation, IP, and query filters
- Exports to the configured OTEL gRPC and/or Splunk HEC sinks
- Waits
check_interval_sseconds, then repeats
Usage¶
# With default config resolution (/etc/click-dog/click-dog.yaml first,
# then ./click-dog.yaml)
./click-dog
# With custom config
./click-dog --config /path/to/config.yaml
Configuration¶
monitor:
enabled: true
min_trace_duration_ms: 1000 # Find traces with spans >= 1 second
min_span_duration_ms: 0 # Export all spans from those traces
check_interval_s: 30 # Poll every 30 seconds
lookback_s: 40 # Look back 40 seconds each poll
Key Behaviors¶
- Runs on startup: the first poll happens immediately, then repeats at
check_interval_s - Deduplication: an LRU cache (default 10,000 entries) keyed by the composite OTLP span identity
(trace_id, span_id)prevents re-exporting the same spans across polling cycles. Span IDs are only unique within a trace, so the trace ID is part of the key - Lookback overlap: set
lookback_sslightly larger thancheck_interval_s(the default adds 10 seconds) to avoid gaps between polls - Graceful shutdown: responds to SIGINT and SIGTERM by cancelling any in-progress polling cycle, including its ClickHouse query, batch delay, or exporter call. Exporter connections are then closed with a 5-second timeout. The in-memory dedup cache is not persisted. After restart, spans still within the lookback window can be re-exported once. Stable
(trace_id, span_id)values make repeats identifiable, but downstream duplicate handling is backend-specific. - Resilience: supports circuit breaker and adaptive backoff to protect ClickHouse during failures
Data Source¶
Scheduled mode reads from system.opentelemetry_span_log, which contains OpenTelemetry spans generated internally by ClickHouse. This includes:
- Query execution spans
- Internal ClickHouse operations (merges, mutations, etc.)
- Distributed query coordination spans
Before scheduled spans reach any exporter, filters.query_text_mode applies
the final raw/redacted/normalized-only/none representation. Filtering may
inspect raw SQL first; multi-sink fan-out receives only the shaped span. The
three restrictive modes also remove URI attribute values carrying HTTP
query parameters. See
Query Text Export Modes.
Backfill Mode¶
One-time export of historical data from ClickHouse's system.query_log table for a specific time range.
How It Works¶
- Queries
system.query_logfor queries in the specified time range matching the duration threshold - Filters by
QueryFinishandExceptionWhileProcessingevent types - Applies configured filters to the original in-process query
- Applies
filters.query_text_modewithout raw fallback (including removing URI attribute copies carried in HTTPqueryparameters) - Exports each shaped query as a trace/event to the configured OTEL gRPC and/or Splunk HEC sinks
- Exits when complete
Usage¶
# Backfill a specific time range
./click-dog backfill \
--config config.yaml \
--start "2024-01-15T00:00:00Z" \
--end "2024-01-15T23:59:59Z"
Time format must be RFC3339 (e.g., 2024-01-15T14:30:00Z). The pre-subcommand
spelling — -backfill-start / -backfill-end mode flags on the root command —
still works but is deprecated.
Configuration¶
monitor:
enabled: false # Disable scheduled mode (optional)
min_trace_duration_ms: 1000 # Only export queries >= 1 second
max_spans_per_cycle: 5000 # Limit total queries per run
batch_size: 100 # Process in batches of 100
batch_delay_ms: 50 # 50ms between batches
Key Behaviors¶
- Exits when done: unlike scheduled mode, backfill runs once and exits
- No deduplication cache: since it is a one-time run, there is no need for the LRU cache
- Rate controllable: use
max_spans_per_cycle,batch_size, andbatch_delay_msto control the export rate - Different data source: reads from
system.query_log(notsystem.opentelemetry_span_log), which provides richer query metadata - Different trace shape: creates one synthetic span per query-log row. It does not recreate native ClickHouse trace topology, child spans, or span-log-only attributes
- Filters applied: IP whitelist, operation whitelist, query blacklist, user filters, and query redaction use
query_logfields. Becausequery_logentries have no operation name, a configuredwhitelist_operationsrejects every backfill query (the sole exception is a pattern that matches the empty string, such as"*"). Leavewhitelist_operationsempty when using backfill - Exit code reflects export outcome: backfill exits 0 only when every non-filtered query exports successfully. Any export failure exits non-zero. The final log line and the
backfill_failedwebhook (if enabled) includequeries=N exported=N filtered=N failed=Nso an operator can tell whether the failure was partial (some data landed) or total (none did) and decide how to recover. The terminalbackfill_completeorbackfill_failedwebhook attempt completes or reaches its configured timeout before the one-shot process exits
Data Source¶
Backfill mode reads from system.query_log, which contains finished query events. This provides attributes not available in the span log, such as:
- Read/write row and byte counts
- Memory usage
- Databases and tables accessed
- Exception codes
Running Both Modes Together¶
You can run scheduled monitoring and backfill simultaneously as two separate processes. This is useful for: - Recovering from an outage while continuing real-time monitoring - Backfilling a historical period for investigation
# Terminal 1: Continuous monitoring
./click-dog --config config.yaml
# Terminal 2: Backfill a specific incident window
./click-dog backfill --config config.yaml \
--start "2024-01-15T10:00:00Z" \
--end "2024-01-15T11:00:00Z"
Chunked Backfill¶
For large time ranges, split the backfill into smaller chunks to control load:
# Process hour by hour
./click-dog backfill --config config.yaml \
--start "2024-01-15T00:00:00Z" \
--end "2024-01-15T01:00:00Z"
./click-dog backfill --config config.yaml \
--start "2024-01-15T01:00:00Z" \
--end "2024-01-15T02:00:00Z"
# ... and so on
Validate Mode¶
Validates the configuration file and prints parsed settings without connecting to anything. Useful for checking config before deployment.
Usage¶
The pre-subcommand spelling, a -validate mode flag on the root command,
still works but is deprecated.
Output¶
Prints parsed settings including: - Any deprecation, environment, and validation warnings from config load - ClickHouse connection details - Configured exporters (count, plus one line per OTEL or Splunk HEC sink) - Monitor settings (minimum trace duration and check interval) and the effective query-text export mode - HA leader-election status when configured - Self-metrics (OTLP push) state, whether enabled or disabled, plus the metrics, health, and webhook listeners when enabled
Exits with code 0 if valid, non-zero if there are errors.
Dry-run Mode¶
Runs one cycle of the pipeline against your real ClickHouse: same
queries, same filters, same dedup, then prints a cumulative summary
(traces, spans, queries, top operations) and exits. Routes exports
through DryRunExporter, so nothing is sent to OTEL or Splunk HEC.
Use it to: - preview what a new filter or duration threshold will export before pointing the binary at a real collector; - diagnose "why isn't this span being exported?" without sending traffic to your observability bill; - smoke-test a config on a production cluster without writing data.
Usage¶
# Dry-run the current lookback window (one cycle, then exit)
./click-dog --dry-run --config click-dog.yaml
# Dry-run a historical window (backfill, then exit)
./click-dog backfill --dry-run --config click-dog.yaml \
--start "2024-01-15T00:00:00Z" \
--end "2024-01-15T23:59:59Z"
Key Behaviors¶
- Single cycle, then exits: scheduled
--dry-runruns exactly the startup-check cycle (onepipeline.Processcall), prints the summary, and returns. It does not enter the ticker loop. Backfill--dry-runruns the configured window, prints the summary, and exits - Real reads, no writes: ClickHouse is queried normally. Span exports
are discarded before any OTEL / Splunk HEC traffic (exporter clients are
constructed but their lazy connections never carry spans). One exception:
with
metrics.otlp.enabled: true, the self-metrics push still opens its collector connection as in a normal run - Summary shape:
Traces,Spans,Queries, and a "Top operations" table (top 10 by count) are printed to stdout viaDryRunExporter.PrintSummary - Sink label: Prometheus metrics emitted in the cycle carry
sink="dry_run"so dashboards can distinguish dry-run traffic from live exports - No exports: the startup webhook notification still fires when webhooks are enabled, but the shutdown webhook does not fire because dry-run exits before the scheduled loop's signal-driven shutdown path
Want a continuous dry-run loop? That isn't supported today: scheduled
--dry-runexits after one cycle. If you need rolling summaries, re-invoke periodically (cron,watch, etc.) or open a follow-up issue.
click-dog analyze (read-only reports)¶
click-dog analyze is a local, read-only inspection command family for query
behavior. It produces reports and exits; it does not start the scheduled polling
loop, open exporters, start metrics or health servers, start leader election, or
run the circuit breaker.
| Subcommand | Purpose |
|---|---|
click-dog analyze queries |
Build a deterministic query-analysis report; optionally capture/compare an explicit known-good baseline, apply a CI finding threshold, or explicitly notify configured analysis destinations after report output |
click-dog analyze trace |
Drill from a current/recent query or explicit identity to native trace spans and nearby query-family context |
See Query Analysis for the operator workflow and JSON
artifact boundary, including --fail-on, --notify, and exit-code precedence.
Operational commands¶
Test exporter delivery¶
test export sends one synthetic span to every configured exporter without
querying ClickHouse. Use it to isolate exporter credentials, TLS, and routing
from the ClickHouse data plane.
The span is named click-dog.test-span and carries
click_dog.source=test-span, click_dog.test=true, and
click_dog.test_kind=export. Each exporter is reported independently. A PASS
means the exporter accepted the span; it does not prove backend indexing.
click-dog test-span remains a deprecated compatibility alias with the same
flags, output behavior, and exit status. It prints the replacement command to
stderr and has no scheduled removal date.
Test native tracing¶
test tracing proves the native ClickHouse path with one bounded, read-only
query:
The command supplies a sampled parent context and a unique
click-dog-test-... query ID to exactly
SELECT 1 /* click-dog test tracing */, polls the span log by the generated
trace ID, and requires a native span to reference the local parent. It then
exports one batch containing the local parent followed by every recovered
native child. The single timeout covers connection, query, polling, fetch, and
export.
This exact-ID path intentionally bypasses scheduled duration, query, user, IP,
and operation filters. It also exports the recovered span-log rows directly:
the scheduled pipeline's query-log enrichment does not run, so this command
does not validate log_comment extraction or table/database string-array
attributes. It tests only the configured connection's selected node, not every
cluster node or shard. A scheduled Click-Dog process can later read and
re-export the native rows because deduplication is process-local; that possible
duplicate does not change the test result.
Trigger an immediate cycle¶
Without Keeper, the command sends SIGUSR1 to the local systemd process. With
ha.keeper.hosts, it writes a shared request that the elected coordination
leader consumes. A local SIGUSR1 or HTTP POST /flush always targets that
specific process: a cluster-mode non-leader records a skipped cycle, while a
sidecar runs its local cycle because sidecar collection is not leader-gated.
The poll timer resets once the flush cycle finishes, so the next scheduled cycle is a full interval away rather than arriving immediately behind it.
Flush is best-effort. A lost request does not change the next scheduled cycle.
click-dog deploy (deployment management)¶
click-dog deploy is a separate subcommand family that manages deployment
artifacts and installed state rather than running the monitoring pipeline. It
does not start the monitoring loop.
| Subcommand | Purpose |
|---|---|
click-dog deploy kubernetes -c COLLECTOR --ch-host CH_SVC |
Render a Deployment + ConfigMap + Secret to ./click-dog-k8s/ |
click-dog deploy docker -c HOST |
Render a Docker Compose stack to ./click-dog-docker/ |
click-dog deploy status |
Report installed version, systemd active/enabled state, and the local /healthz response. Add --json for machine output |
sudo click-dog deploy uninstall |
Confirm, then remove a standard systemd installation, including config and credentials |
The Kubernetes and Docker generators take the same -c (collector
address), -u / -p (ClickHouse credentials), -o (output directory),
and -v (version) flags. Pass --update to refresh the image tag in an
existing output directory without overwriting other fields.
click-dog deploy status exits 0 only when systemd reports active
and /healthz returns 200; otherwise 1. Plain-text output is
printed in both cases.
See the Kubernetes and Docker guides for the equivalent installer-wrapper recipes.
Comparison¶
| Scheduled | Backfill | Validate | Dry-run | Analyze | Test | deploy |
|
|---|---|---|---|---|---|---|---|
| Data source | system.opentelemetry_span_log |
system.query_log |
None | Same as scheduled / backfill | system.query_log, system.processes, and/or system.opentelemetry_span_log |
None (export) or exact trace ID in system.opentelemetry_span_log (tracing) |
None |
| Duration | Runs continuously | Runs once, then exits | Exits immediately | Runs once, then exits | Exits after producing a report | Exits after the bounded smoke test | Exits immediately |
| Time range | Rolling window (lookback_s) |
Fixed range (backfill --start/--end) |
N/A | Same as the mode it shadows | Bounded lookback window | One generated trace, one-day partition floor | N/A |
| Deduplication | LRU cache prevents re-exports | Not needed (one-time run) | N/A | Active during the single cycle | N/A | None; a scheduled process may later re-export native rows | N/A |
| Resilience | Circuit breaker + adaptive backoff | None (single run) | N/A | Single cycle; no loop to back off | N/A | One command deadline; bounded polling | N/A |
| Exports anything? | Yes | Yes | No | No; discards via DryRunExporter |
No | Yes | No |
| Use case | Real-time monitoring | Historical query visibility and query-log-based outage recovery | Config verification | Preview / smoke-test without sending | Local query triage and report artifacts | Exporter and native tracing go-live checks | Generate manifests, inspect installed state, or remove a standard systemd installation |