Datadog Self-Monitoring¶
This guide is part of the Datadog integration. It covers the Click-Dog health metrics consumed by the packaged Health dashboard, including the default OTLP path and the legacy OpenMetrics scrape path.
Click-dog can push its own low-cardinality exporter health metrics over OTLP:
cycle counts and outcomes, export totals, errors, circuit-breaker state,
backoff interval, last-success timestamp, and last-cycle duration / spans.
Enable metrics.otlp.enabled: true; by default click-dog reuses
exporters.otel[0] and sends self-metrics to the same collector as spans. The
production init profile (the default for click-dog init) writes this block.
A guided install already feeds the Health dashboard. Delete the block to opt out.
Set metrics.otlp.host in containers if you want a stable host.name dashboard
variable instead of the pod/container hostname.
Standalone metrics.otlp connections currently expose plaintext/TLS settings
only; use inherit_otel_connection: true when the collector requires the
span exporter's CA or mTLS client certificate settings.
The Prometheus/OpenMetrics-compatible /metrics endpoint is still available
(default :9090/metrics, opt-in via metrics.enabled: true; see
Observability) for Prometheus, Grafana Alloy, the OTel
Collector's prometheus receiver, or Datadog users who cannot enable OTLP
self-metrics.
Traces configured through the
Datadog Agent or OpenTelemetry Collector carry query data such
as db.statement, latency, user, app, and normalized query family. The
Click-Dog: Application Query Analysis dashboard reads these traces.
Self-metrics show whether Click-Dog is running, scraping, and exporting. The
Click-Dog: Health dashboard shows current action state first: last success
age, breaker state, current error rate, export throughput, and backoff interval.
Trends and lifetime counters follow.
Legacy Datadog Agent OpenMetrics scrape config¶
Prometheus-only users can still scrape /metrics with Datadog. Drop this in
/etc/datadog-agent/conf.d/openmetrics.d/conf.yaml and restart the Agent. The
explicit metrics: rename list is required for the scrape path; a generic
wildcard such as - click_dog_* will not produce usable Datadog names.
dashboards/README.md and dashboards/datadog-clickdog-health.json document
the same legacy mapping.
init_config:
instances:
- openmetrics_endpoint: http://localhost:9090/metrics
namespace: click_dog
metrics:
# Counters: Datadog appends `.count` and submits as a monotonic rate.
- click_dog_spans_exported_total: spans_exported
- click_dog_spans_filtered_total: spans_filtered
- click_dog_spans_duplicates_total: spans_duplicates
- click_dog_export_attempts_total: export.attempts
- click_dog_export_accepted_total: export.accepted
- click_dog_export_errors_total: export.errors
- click_dog_cycle_results_total: cycle_results
# Gauges: submitted as-is under the namespace.
- click_dog_circuit_breaker_state: circuit_breaker.state
- click_dog_leader: leader
- click_dog_backoff_interval_seconds: backoff_interval.seconds
- click_dog_last_success_timestamp_seconds: last_success_timestamp.seconds
- click_dog_uptime_seconds: uptime.seconds
- click_dog_last_cycle_duration_seconds: last_cycle.duration_seconds
- click_dog_last_cycle_exported_spans: last_cycle.exported_spans
- click_dog_last_cycle_filtered_spans: last_cycle.filtered_spans
- click_dog_last_cycle_duplicate_spans: last_cycle.duplicate_spans
# ClickHouse data-plane health: span_log freshness, enrichment
# health, and query_id coverage. Lets an operator distinguish
# "click-dog is down" from "ClickHouse is not emitting useful data".
- click_dog_span_log_last_poll_timestamp_seconds: span_log.last_poll_timestamp.seconds
- click_dog_span_log_newest_row_age_seconds: span_log.newest_row_age.seconds
- click_dog_span_log_rows_last_cycle: span_log.rows_last_cycle
- click_dog_query_log_enrichment_attempts_total: query_log.enrichment.attempts
- click_dog_query_log_enrichment_successes_total: query_log.enrichment.successes
- click_dog_query_log_enrichment_failures_total: query_log.enrichment.failures
- click_dog_query_log_enrichment_match_ratio: query_log.enrichment.match_ratio
- click_dog_spans_with_query_id_ratio: spans_with_query_id_ratio
- click_dog_normalized_query_supported: normalized_query_supported
- click_dog_query_operation_supported: query_operation_supported
# Topology self-audit: detects the sidecar + use_cluster_queries anti-pattern.
- click_dog_topology_warning: topology_warning
Mapped metric names¶
Prometheus metric (/metrics) |
Type | Datadog metric |
|---|---|---|
click_dog_spans_exported_total |
counter | click_dog.spans_exported.count |
click_dog_spans_filtered_total |
counter | click_dog.spans_filtered.count |
click_dog_spans_duplicates_total |
counter | click_dog.spans_duplicates.count |
click_dog_export_attempts_total |
counter | click_dog.export.attempts.count (tag: sink) |
click_dog_export_accepted_total |
counter | click_dog.export.accepted.count (tag: sink) |
click_dog_export_errors_total |
counter | click_dog.export.errors.count (tag: sink) |
click_dog_cycle_results_total |
counter | click_dog.cycle_results.count (tag: result) |
click_dog_circuit_breaker_state |
gauge | click_dog.circuit_breaker.state |
click_dog_leader |
gauge | click_dog.leader |
click_dog_backoff_interval_seconds |
gauge | click_dog.backoff_interval.seconds |
click_dog_last_success_timestamp_seconds |
gauge | click_dog.last_success_timestamp.seconds (tag: role) |
click_dog_uptime_seconds |
gauge | click_dog.uptime.seconds |
click_dog_last_cycle_duration_seconds |
gauge | click_dog.last_cycle.duration_seconds |
click_dog_last_cycle_exported_spans |
gauge | click_dog.last_cycle.exported_spans |
click_dog_last_cycle_filtered_spans |
gauge | click_dog.last_cycle.filtered_spans |
click_dog_last_cycle_duplicate_spans |
gauge | click_dog.last_cycle.duplicate_spans |
click_dog_span_log_last_poll_timestamp_seconds |
gauge | click_dog.span_log.last_poll_timestamp.seconds |
click_dog_span_log_newest_row_age_seconds |
gauge | click_dog.span_log.newest_row_age.seconds |
click_dog_span_log_rows_last_cycle |
gauge | click_dog.span_log.rows_last_cycle |
click_dog_query_log_enrichment_attempts_total |
counter | click_dog.query_log.enrichment.attempts.count |
click_dog_query_log_enrichment_successes_total |
counter | click_dog.query_log.enrichment.successes.count |
click_dog_query_log_enrichment_failures_total |
counter | click_dog.query_log.enrichment.failures.count |
click_dog_query_log_enrichment_match_ratio |
gauge | click_dog.query_log.enrichment.match_ratio |
click_dog_spans_with_query_id_ratio |
gauge | click_dog.spans_with_query_id_ratio |
click_dog_normalized_query_supported |
gauge | click_dog.normalized_query_supported |
click_dog_query_operation_supported |
gauge | click_dog.query_operation_supported |
click_dog_topology_warning |
gauge | click_dog.topology_warning (tag: reason) |
ClickHouse data-plane health signals¶
The metrics in the bottom block of the table above show ClickHouse data-plane health. They answer whether ClickHouse is producing the data Click-Dog needs. Use them when dashboards are empty but Click-Dog looks healthy:
click_dog_span_log_last_poll_timestamp_seconds: Unix timestamp of the most recent span-log fetch that completed without error. An empty fetch (zero rows returned) still advances this gauge; absence of the metric family from/metricsmeans click-dog is down.click_dog_span_log_newest_row_age_seconds: age in seconds of the newest span observed in the most recent non-empty fetch. Empty cycles do NOT reset this; the gauge keeps growing while no new spans arrive, so monotonic growth is the staleness signal.0means no observation yet.click_dog_span_log_rows_last_cycle: raw row count returned by the most recent span-log fetch, before click-dog's in-process filter/dedup.click_dog_query_log_enrichment_{attempts,successes,failures}_total: cycle-level counters tracking query_log join health. An attempt is recorded only when at least one span carried aclickhouse.query_idto look up, so non-enrichment configurations cost nothing.click_dog_query_log_enrichment_match_ratio: last cycle's matched / requested ratio. Persists across failed attempts (a failed lookup tells us nothing about how well a future join would match).0before the first successful enrichment.click_dog_spans_with_query_id_ratio: last cycle's fraction of fetched spans carryingclickhouse.query_id. Counted on the raw (pre-filter) span set so the gauge reflects ClickHouse's output, not click-dog's choices. A value below 1.0 means downstream query-analysis dashboards lose join context for some spans.click_dog_normalized_query_supported:1ifsystem.query_log.normalized_query_hashis available on the connected ClickHouse and click-dog is using it;0otherwise. In cluster query mode,1means every replica passed theclusterAllReplicas(..., system.columns)compatibility probe. Set once at startup from the capability probe ininternal/clickhouse/reader.go.click_dog_query_operation_supported:1ifsystem.query_log.query_kindis available and click-dog is exportingquery_log.operationplusquery_log.access_type;0explains why those activity dimensions are absent. In cluster query mode every replica must pass the startup compatibility probe.click_dog_topology_warning:1when the topology self-audit detects more than one instance running whole-cluster span reads (use_cluster_queries), which duplicates exports;0when clean. Emitted perreason(sidecar_cluster_queries|multi_instance_cluster_queries) and always present for both reasons, so a green0is distinguishable from no-data. It affects observability only and never changes/readyz. A red1can also reflect bounded duplication during a sustained Keeper outage, when instances fail open and export by design, or a long-running backfill from another host. It clears on its own once resolved. The gauge is always emitted (stable0for both reasons from startup, even when the auditor is skipped); the auditor only runs itssystem.query_logprobe and can raise the gauge to1when this instance hasuse_cluster_queries: true. See Operating · Cluster topology.
The recommended pairing on the Health dashboard:
| Operator question | Read these together |
|---|---|
| "Is click-dog alive?" | uptime_seconds (cockpit) + span_log_last_poll_timestamp_seconds |
| "Is ClickHouse emitting spans?" | span_log_rows_last_cycle + span_log_newest_row_age_seconds |
| "Are query_log joins still working?" | query_log_enrichment_failures.count rate + query_log_enrichment_match_ratio |
| "Are downstream query dashboards complete?" | spans_with_query_id_ratio + normalized_query_supported + query_operation_supported |
| "Is exactly one reader doing cluster reads?" | topology_warning (per reason) + the Per-host health table |
Troubleshooting "No data" on the health dashboard¶
- Confirm
metrics.otlp.enabled: true. - Confirm the Datadog Agent or collector OTLP gRPC receiver is reachable from click-dog.
- In containers, set
metrics.otlp.hostto the stable node/instance identity you want to use for thehost.nametemplate variable. - On the legacy scrape path, search Metrics Explorer for
click_dog.click_dog_spans_exported_total; that means the explicit OpenMetricsmetrics:rename list above is not being applied.
Health endpoints (HTTP / Synthetic check alternative)¶
Alongside /metrics, click-dog exposes dedicated HTTP endpoints suitable for
Datadog HTTP checks, Synthetic checks, or Kubernetes probes. Full reference:
Observability: Health Endpoints.
| Endpoint | Default port | Use from Datadog |
|---|---|---|
GET /healthz |
:8686 |
HTTP liveness check; always 200 while the process is up |
GET /readyz |
:8686 |
HTTP readiness check; 200 only when ClickHouse is reachable AND the breaker is closed |
GET /status |
:8686 |
Synthetic or scripted check with a JSON body containing version, clickhouse_healthy, circuit_breaker, backoff_interval_s, and last_cycle.{exported,filtered,duplicates,duration_ms,at,error}. Always returns 200; assert on the JSON body, not the HTTP status |
GET /clusterz |
:8686 (leader only, HA) |
Aggregate /readyz across configured peers. Only the leader serves this path; followers return 404 {"status":"not_leader"}. Returns 200 only when every node is healthy; 503 when any node is degraded or unreachable |
These overlap with Tier 1 metrics on purpose. Pick endpoints when you want a
probe Datadog already speaks (HTTP / Synthetic) without standing up an
OpenMetrics scrape; pick metrics when you want trends, ratios, and dashboard
widgets. Health endpoints are opt-in via the health: block. See the
Observability doc for the binding rules and the /clusterz HA configuration.