Datadog Monitors¶
This guide is part of the Datadog integration. Configure these alerts after the Click-Dog health metrics and packaged dashboards are reporting data.
The canonical setup is the Click-Dog: Health dashboard wired to OTLP self-metrics. Build alerts on top of it in this order:
Tier 1: Exporter health (canonical)¶
These monitors fire on click-dog itself and use the default OTLP self-metric
names (Datadog adds .count to monotonic sums). They are the primary
signal that click-dog is doing its job.
-
Stale exports: last successful cycle is too old. The single most important alert: if click-dog stops, you lose visibility into ClickHouse. Configure this as a formula metric monitor (Datadog → Monitors → New Monitor → Metric → "Use a formula"), not a single-metric monitor:
-
Query
a:min:click_dog.last_success_timestamp_seconds{role:active} - Formula:
now() - a - Alert threshold:
> 600
last_success_timestamp_seconds is absent until the first successful
cycle, so a fresh or failed-start instance reads as no-data rather than
epoch zero. In HA mode, the role:active filter excludes leader-gated
standbys from stale-export paging while keeping the active exporter covered.
The min: aggregation is intentional: it pages when any active exporter is
stale, matching the Health dashboard's worst-case cockpit tile.
now() is only available in the formula editor; the basic
single-metric monitor will reject it. The ≥ 600 threshold is the
conservative floor: 2–3× monitor.backoff.max_interval_s, where
the default is 300 s. Adaptive backoff widens the gap between
successful cycles during sustained errors, so anything tighter than
the maximum backoff will alert on a healthy exporter that is
pacing itself. If you also wire Tier 1 #2 (breaker not closed) to
catch the sustained-error case, a tighter threshold like
5–10× check_interval_s (default 30 s, so 150–300 s) detects
total Click-Dog failure much faster. The breaker monitor limits
the false-positive risk that the wider floor exists to avoid.
-
Circuit breaker not closed: exporter is shedding load against ClickHouse.
0=closed,1=half_open,2=open. Half-open is a normal-but-noisy recovery state; tighten to>= 2if you only want hard-open pages. -
Cycle error ratio: sustained errors even before the breaker trips.
cycle_resultstags every cycle as exactly one ofsuccess,error,skipped, so this is a clean ratio against the summed total.sum:click_dog.cycle_results.count{result:error}.as_rate() / sum:click_dog.cycle_results.count{*}.as_rate() > 0.1as_rate()resolves per evaluation bucket, so a single noisy bucket can page. Evaluate over a multi-bucket window (e.g. ≥ 10 min) with an at-least-N-of-M condition (e.g. 3 of 5 buckets over threshold) to avoid single-bucket flap. -
OTLP self-metrics are missing: backstop in case the self-metrics exporter or collector path stops running (and therefore Tier 1 #1–#3 stop firing). Create a metric monitor on
avg:click_dog.cycle_results.count{*}and enable "Notify if data is missing for the last 5 minutes" in the monitor options panel.no datais a monitor condition, not a query function, so it lives in the options checkbox rather than the query field. -
Stale span_log source: Click-Dog is exporting, but ClickHouse has stopped producing spans for click-dog to read. This is the "click-dog looks healthy, dashboards are empty" failure mode that Tier 1 #1 misses:
Pair with #1 (stale exports) to triangulate: #1 alone means click-dog stopped; #5 alone means ClickHouse stopped emitting spans; both firing means click-dog crashed while ClickHouse kept producing. Threshold islast_success_timestampkeeps advancing on every poll (including zero-row polls), so the staleness signal lives on the data-plane gauge introduced for #183:> 600 sto match the dashboard tile's red band; tune for your span retention window. Note0is ambiguous on this gauge ("no observation yet" OR "the newest span arrived just now"), so always alert on>rather than==(see Observability: Available Metrics). -
Optional push-side webhook events: if the
webhookblock is configured (see Observability: Webhook Notifications), click-dog itself firescircuit_breaker_opened,circuit_breaker_closed,backfill_complete,backfill_failed,error_spike,startup, andshutdownstraight to Slack or a Datadog webhook integration without waiting on a scrape window. PagerDuty Events API v2 requires an adapter for click-dog's Slack-shaped payload.
Tier 2: Query analytics (from traces)¶
These monitors fire on the traces click-dog exports. They catch ClickHouse-side regressions, not click-dog-side problems.
The monitor examples below focus on the metric and tag names are the load-bearing part. Tune your team's exact monitor query (window, threshold, group-by, anomaly bounds) against them; what's worth paging on differs by workload, and Tier 1 already covers exporter health, so Tier 2 deliberately stays at the "here is the right signal" level rather than prescribing thresholds.
-
Slow query spike: P95 query duration regression.
-
Trace error rate (backfill mode only): alert when >5% of spans have
error:true. Theerrorattribute only exists on backfill spans fromquery_log; scheduled-mode spans fromopentelemetry_span_logalways have status OK. -
Memory hog queries: use
@query_log.memory_usageon enriched live query-root spans or@db.memory_usageon backfill spans. -
Read amplification: watch P95 of
@query_log.read_byteson enriched live query-root spans or@db.read_byteson backfill spans. -
Unusual query volume: anomaly detection on span count; catches both traffic spikes and unexpected drops.
-
Top-N queries by duration: weekly digest of the slowest query families, grouped by
@query_log.normalized_query_hash. Useful for regression hunting without an alerting threshold. -
Per-user query patterns: track which users generate load, grouped by
@query_log.user. Useful for capacity planning and noisy-neighbour detection. -
Cross-host skew: alert if one ClickHouse node's P95 latency drifts well above its peers, grouped by
@hostname. Catches single-node degradation that aggregate latency monitors mask.
Query-analysis Event Monitor¶
When the optional datadog_events destination
is enabled, create a Datadog Event Monitor under Monitors → New Monitor →
Event. Use Event Explorer search syntax and count matching events over a
window that covers the cadence of your external analysis job:
Trigger when the event count is above or equal to 1. For a paging-only
monitor, narrow the query to critical summaries:
event_type:analysis_findings environment:production service:click-dog highest_eligible_severity:critical
To page only when the comparison proves a newly critical condition, include
the numeric custom attribute (Events Explorer attribute searches use @):
event_type:analysis_findings environment:production service:click-dog highest_eligible_severity:critical @new_critical:>0
For a warning workflow use highest_eligible_severity:warning
@new_warning:>0. Confirm the new_critical and new_warning attribute names
from one ingested event's Attributes side panel before saving the monitor;
an organization-level event processing pipeline can remap custom attributes.
Use environment and service as group facets if one monitor covers multiple
deployments. A concise notification template can safely reference the event's
bounded text and tags:
Click-Dog query analysis found an eligible condition.
Service: {{event.tags.service}}
Environment: {{event.tags.environment}}
Severity: {{event.tags.highest_eligible_severity}}
{{event.text}}
event.text includes the total and new critical/warning/info counts and points
to the local report; the underlying numeric new_* custom attributes remain
available for the monitor query.
The event intentionally carries condition identities and counts, not full
finding evidence. Keep the local JSON report from the same run and drill down
with click-dog analyze trace --normalized-query-hash <hash>. An Events API
2xx confirms intake acceptance only; configure and verify this monitor
separately before depending on downstream paging.
Tier 3: Fallback signals¶
Complementary backstops for setups where Tier 1 and Tier 2 aren't fully wired yet, or for catching failures the higher tiers can't observe:
- Log-based monitors: ship Click-Dog logs to Datadog Logs and alert on
[ERROR]lines or circuit-breaker warnings. Useful when/metricsis not (yet) scraped. - Process check: Datadog Agent
processcheck on theclick-dogbinary. The one failure mode/metricscannot report on is the binary crashing because the process serving/metricsis the one that died. - No data on traces: fires if no spans land in the trace pipeline. Less precise than Tier 1 #1 because it conflates "click-dog is down" with "ClickHouse has no traffic", but it works without any extra wiring.