Skip to content

Datadog Monitors

This guide is part of the Datadog integration. Configure these alerts after the Click-Dog health metrics and packaged dashboards are reporting data.

The canonical setup is the Click-Dog: Health dashboard wired to OTLP self-metrics. Build alerts on top of it in this order:

Tier 1: Exporter health (canonical)

These monitors fire on click-dog itself and use the default OTLP self-metric names (Datadog adds .count to monotonic sums). They are the primary signal that click-dog is doing its job.

  1. Stale exports: last successful cycle is too old. The single most important alert: if click-dog stops, you lose visibility into ClickHouse. Configure this as a formula metric monitor (Datadog → Monitors → New Monitor → Metric → "Use a formula"), not a single-metric monitor:

  2. Query a: min:click_dog.last_success_timestamp_seconds{role:active}

  3. Formula: now() - a
  4. Alert threshold: > 600

last_success_timestamp_seconds is absent until the first successful cycle, so a fresh or failed-start instance reads as no-data rather than epoch zero. In HA mode, the role:active filter excludes leader-gated standbys from stale-export paging while keeping the active exporter covered. The min: aggregation is intentional: it pages when any active exporter is stale, matching the Health dashboard's worst-case cockpit tile. now() is only available in the formula editor; the basic single-metric monitor will reject it. The ≥ 600 threshold is the conservative floor: 2–3× monitor.backoff.max_interval_s, where the default is 300 s. Adaptive backoff widens the gap between successful cycles during sustained errors, so anything tighter than the maximum backoff will alert on a healthy exporter that is pacing itself. If you also wire Tier 1 #2 (breaker not closed) to catch the sustained-error case, a tighter threshold like 5–10× check_interval_s (default 30 s, so 150–300 s) detects total Click-Dog failure much faster. The breaker monitor limits the false-positive risk that the wider floor exists to avoid.

  1. Circuit breaker not closed: exporter is shedding load against ClickHouse.

    max:click_dog.circuit_breaker_state{*} >= 1
    
    0=closed, 1=half_open, 2=open. Half-open is a normal-but-noisy recovery state; tighten to >= 2 if you only want hard-open pages.

  2. Cycle error ratio: sustained errors even before the breaker trips. cycle_results tags every cycle as exactly one of success, error, skipped, so this is a clean ratio against the summed total.

    sum:click_dog.cycle_results.count{result:error}.as_rate() /
      sum:click_dog.cycle_results.count{*}.as_rate() > 0.1
    
    as_rate() resolves per evaluation bucket, so a single noisy bucket can page. Evaluate over a multi-bucket window (e.g. ≥ 10 min) with an at-least-N-of-M condition (e.g. 3 of 5 buckets over threshold) to avoid single-bucket flap.

  3. OTLP self-metrics are missing: backstop in case the self-metrics exporter or collector path stops running (and therefore Tier 1 #1–#3 stop firing). Create a metric monitor on avg:click_dog.cycle_results.count{*} and enable "Notify if data is missing for the last 5 minutes" in the monitor options panel. no data is a monitor condition, not a query function, so it lives in the options checkbox rather than the query field.

  4. Stale span_log source: Click-Dog is exporting, but ClickHouse has stopped producing spans for click-dog to read. This is the "click-dog looks healthy, dashboards are empty" failure mode that Tier 1 #1 misses: last_success_timestamp keeps advancing on every poll (including zero-row polls), so the staleness signal lives on the data-plane gauge introduced for #183:

    max:click_dog.span_log_newest_row_age_seconds{*} > 600
    
    Pair with #1 (stale exports) to triangulate: #1 alone means click-dog stopped; #5 alone means ClickHouse stopped emitting spans; both firing means click-dog crashed while ClickHouse kept producing. Threshold is > 600 s to match the dashboard tile's red band; tune for your span retention window. Note 0 is ambiguous on this gauge ("no observation yet" OR "the newest span arrived just now"), so always alert on > rather than == (see Observability: Available Metrics).

  5. Optional push-side webhook events: if the webhook block is configured (see Observability: Webhook Notifications), click-dog itself fires circuit_breaker_opened, circuit_breaker_closed, backfill_complete, backfill_failed, error_spike, startup, and shutdown straight to Slack or a Datadog webhook integration without waiting on a scrape window. PagerDuty Events API v2 requires an adapter for click-dog's Slack-shaped payload.

Tier 2: Query analytics (from traces)

These monitors fire on the traces click-dog exports. They catch ClickHouse-side regressions, not click-dog-side problems.

The monitor examples below focus on the metric and tag names are the load-bearing part. Tune your team's exact monitor query (window, threshold, group-by, anomaly bounds) against them; what's worth paging on differs by workload, and Tier 1 already covers exporter health, so Tier 2 deliberately stays at the "here is the right signal" level rather than prescribing thresholds.

  1. Slow query spike: P95 query duration regression.

    avg:trace.duration{service:click-dog-monitor} > 5000000000  # 5e9 ns = 5 s
    

  2. Trace error rate (backfill mode only): alert when >5% of spans have error:true. The error attribute only exists on backfill spans from query_log; scheduled-mode spans from opentelemetry_span_log always have status OK.

  3. Memory hog queries: use @query_log.memory_usage on enriched live query-root spans or @db.memory_usage on backfill spans.

  4. Read amplification: watch P95 of @query_log.read_bytes on enriched live query-root spans or @db.read_bytes on backfill spans.

  5. Unusual query volume: anomaly detection on span count; catches both traffic spikes and unexpected drops.

  6. Top-N queries by duration: weekly digest of the slowest query families, grouped by @query_log.normalized_query_hash. Useful for regression hunting without an alerting threshold.

  7. Per-user query patterns: track which users generate load, grouped by @query_log.user. Useful for capacity planning and noisy-neighbour detection.

  8. Cross-host skew: alert if one ClickHouse node's P95 latency drifts well above its peers, grouped by @hostname. Catches single-node degradation that aggregate latency monitors mask.

Query-analysis Event Monitor

When the optional datadog_events destination is enabled, create a Datadog Event Monitor under Monitors → New Monitor → Event. Use Event Explorer search syntax and count matching events over a window that covers the cadence of your external analysis job:

event_type:analysis_findings environment:production service:click-dog

Trigger when the event count is above or equal to 1. For a paging-only monitor, narrow the query to critical summaries:

event_type:analysis_findings environment:production service:click-dog highest_eligible_severity:critical

To page only when the comparison proves a newly critical condition, include the numeric custom attribute (Events Explorer attribute searches use @):

event_type:analysis_findings environment:production service:click-dog highest_eligible_severity:critical @new_critical:>0

For a warning workflow use highest_eligible_severity:warning @new_warning:>0. Confirm the new_critical and new_warning attribute names from one ingested event's Attributes side panel before saving the monitor; an organization-level event processing pipeline can remap custom attributes.

Use environment and service as group facets if one monitor covers multiple deployments. A concise notification template can safely reference the event's bounded text and tags:

Click-Dog query analysis found an eligible condition.
Service: {{event.tags.service}}
Environment: {{event.tags.environment}}
Severity: {{event.tags.highest_eligible_severity}}
{{event.text}}

event.text includes the total and new critical/warning/info counts and points to the local report; the underlying numeric new_* custom attributes remain available for the monitor query.

The event intentionally carries condition identities and counts, not full finding evidence. Keep the local JSON report from the same run and drill down with click-dog analyze trace --normalized-query-hash <hash>. An Events API 2xx confirms intake acceptance only; configure and verify this monitor separately before depending on downstream paging.

Tier 3: Fallback signals

Complementary backstops for setups where Tier 1 and Tier 2 aren't fully wired yet, or for catching failures the higher tiers can't observe:

  • Log-based monitors: ship Click-Dog logs to Datadog Logs and alert on [ERROR] lines or circuit-breaker warnings. Useful when /metrics is not (yet) scraped.
  • Process check: Datadog Agent process check on the click-dog binary. The one failure mode /metrics cannot report on is the binary crashing because the process serving /metrics is the one that died.
  • No data on traces: fires if no spans land in the trace pipeline. Less precise than Tier 1 #1 because it conflates "click-dog is down" with "ClickHouse has no traffic", but it works without any extra wiring.