Filtering¶
Click-Dog provides multiple filtering layers to control which spans get exported.
Filter Evaluation Order¶
- Operation blacklist (SQL-level):
blacklist_operationspatterns are pushed down to ClickHouse asNOT LIKEclauses. Matching spans are never fetched. Use this primary volume control to drop high-volume internal pipeline spans such asMergeTreeIndex. - User filter:
whitelist_usersis pushed down to thequery_logenrichment WHERE clause,blacklist_usersis kept for the Go-layer span re-check, and both lists are pushed down for the slow-query / backfill path. - IP whitelist: if configured, reject spans from IPs not in the list
- Operation whitelist: if configured, reject spans with non-matching operation names
- Query blacklist: reject spans whose SQL text matches any blacklist pattern
- Query-text shaping: apply
query_text_modeafter enrichment and immediately before export; this does not affect whether a span is exported
Step 1 runs entirely at the SQL level (before data transfer). Step 2 is mixed: the enrichment leg is SQL-level, the per-span re-check is post-fetch. Steps 3–5 run in Go after spans are fetched. A span must pass all checks in steps 1–5 to be exported. Step 6 is the final emission boundary before MultiExporter and every sink. If a whitelist is not configured (empty list), that check is skipped.
The list above describes enforcement order (SQL first, then Go filtering, then query-text shaping). Filtering may inspect raw SQL in process, but exporters receive only the selected representation. Within the Go-level loop the per-span user re-check runs alongside the IP / operation / query checks; the relative order of those Go-level checks is an implementation detail and shouldn't be relied on for debugging filter behavior.
Note: blacklist_operations (SQL-level, step 1) and whitelist_operations (post-fetch, step 4) are independent. The blacklist always wins because matching spans are never fetched: they cannot reach the whitelist check.
Operation Blacklist (SQL-level)¶
Drop high-volume internal ClickHouse spans before they are fetched. Patterns are matched as substrings against operation_name using SQL NOT LIKE '%pattern%'.
filters:
blacklist_operations:
- "MergeTreeSource"
- "MergeTreeMarksLoader"
- "MergeTreeIndex"
- "MergeTreeSequentialSource"
- "VFSWrite"
- "WriteBufferFromS3"
- "ConcurrentJoin"
- "QueryPipelineEx"
This is the default for new installs. It cuts span volume by ~90-99% without affecting query-level spans or their sub-queries.
Operation Whitelist¶
Include only spans with matching operation names. Supports * as a wildcard character.
filters:
whitelist_operations:
- "DB::InterpreterSelectQuery::execute()"
- "DB::Interpreter*::execute()"
- "HTTPHandler::*"
Wildcard Behavior¶
The * character matches any sequence of characters. Internally, it's converted to .* in a regex anchored with ^...$.
| Pattern | Matches | Doesn't Match |
|---|---|---|
DB::InterpreterSelectQuery::execute() |
Exact match only | Any other operation |
DB::Interpreter*::execute() |
DB::InterpreterSelectQuery::execute(), DB::InterpreterInsertQuery::execute() |
HTTPHandler::handleRequest() |
*::execute() |
Any operation ending in ::execute() |
DB::merge() |
* |
Everything | None |
When Not Configured¶
If whitelist_operations is empty or not set, all operation names are allowed through.
IP Whitelist¶
Include only spans from specific client IP addresses.
Entries without / match by exact IP string. Entries in CIDR notation match any
address in that range. The value is the originating client: ClickHouse's
initial_address in the query log (falling back to address when a row has
none), or the client.address attribute in the span log when nothing in the
query log answers. A distributed query fans out to the other servers as
secondary queries whose own address is the initiating server; every row of
the query carries the same initial_address, so the whole trace, including
its shard work, follows the client's decision rather than the servers'. Backfill
applies the same rule to each query-log row.
In scheduled mode the address always comes from query-log enrichment, because
ClickHouse does not put a client address on system.opentelemetry_span_log
rows. Only the query root carries clickhouse.query_id, so the filter resolves
one address per trace and applies it to every span in that trace — a trace is
admitted or excluded whole, rather than reduced to its root span. That holds
when monitor.max_spans_per_cycle splits a trace across polling cycles: a
cycle that holds a trace's children but not its root looks the trace's query
IDs up in the span log and resolves it through the query log all the same.
Two consequences follow. Enrichment must be able to reach system.query_log,
so whitelist_ips with monitor.enrich_from_query_log: false is rejected at
config load, and a runtime enrichment failure logs a warning saying every span
will be filtered. Confirm click-dog check reports the enrichment join before
relying on this filter. And a trace whose address cannot be resolved at all is
dropped, which is strict in the same way whitelist_users is.
When Not Configured¶
If whitelist_ips is empty or not set, all IP addresses are allowed through.
User Filter¶
Restrict export by originating ClickHouse user. Two independent lists:
filters:
whitelist_users: # If set, only these users' spans are exported
- "app_frontend"
- "app_analytics"
blacklist_users: # Never export spans from these users
- "patient_records"
- "ml_training"
The user is sourced from system.query_log.user. The whitelist is pushed down to the enrichment query as AND user IN (...). The blacklist is not pushed down to the enrichment query: doing so would strip blacklisted users' rows from the enrichment map and leave the per-span resolver unable to identify the user, which the Go-level blacklist (permissive on unknown users by design) would then silently let through. Blacklist is enforced exclusively at the Go layer for the span path. For the slow-query / backfill path, both lists are pushed to SQL because that query returns rows with user populated.
| Filter | Enrichment SQL (query_log by query_id) |
Slow-query SQL (backfill / flush) | Go-layer re-check |
|---|---|---|---|
whitelist_users |
AND user IN ? |
AND user IN ? |
yes |
blacklist_users |
(not pushed. See rationale above) | AND user NOT IN ? |
yes |
Requires
monitor.enrich_from_query_log: true(the default). With enrichment disabled the per-span resolver has noquery_logrow to consult, every span resolves to an empty user, and the filter degenerates: a whitelist drops everything; a blacklist passes everything. Click-dog's config validation rejects this combination at startup.
Matches are case-sensitive (exact string equality), mirroring ClickHouse's username handling: App_Frontend will not match app_frontend. When both lists are configured, a user must appear in whitelist_users AND not appear in blacklist_users to pass.
query_log.user is the user that initiated the outermost query, not a sub-query author or a UDF runner. Multi-tenant deployments where one CH service account runs queries on behalf of many end-users cannot be separated at this layer. Add the end-user identity as a log_comment attribute or use distinct ClickHouse users per tenant.
If a future ClickHouse version starts populating a clickhouse.user attribute on system.opentelemetry_span_log, click-dog will use that attribute in preference to the query_log.user enrichment lookup. The two may legitimately differ (e.g. an executing user after SET ROLE versus the session user); operators relying on the user filter for compliance should verify which identity their CH version emits before upgrading.
Use cases¶
- Compliance / PII segregation. Pair with minimum-privilege ClickHouse grants
and the span-attribute privacy notes
to ensure data from a sensitive user, such as
patient_records, is never forwarded to an external observability backend. - Noise reduction. Exclude background users (replication, backup, maintenance bots) that generate span volume without analytical value.
Whitelist semantics¶
whitelist_users is strict, mirroring whitelist_ips: any span whose user cannot be determined (typically internal child spans without a clickhouse.query_id) is dropped. If most of your traffic is single-step queries this is fine; if you rely on multi-span trace fan-out, prefer blacklist_users or pair with the ClickHouse-side user grants for hard segregation.
When not configured¶
If both lists are empty, the user filter is a no-op and no enrichment SQL is added.
Debug-log note¶
log_level: debug writes the matching/non-matching ClickHouse user name to the log when a span is filtered (e.g. Filtering span from blacklisted user: "patient_records"). For PII-segregation deployments where usernames themselves are sensitive, keep log_level: info or higher in production.
Query Blacklist¶
Exclude spans whose SQL query text matches any of the provided regex patterns.
filters:
blacklist_queries:
- "^SELECT \\* FROM system\\." # System table queries
- "SHOW TABLES" # SHOW commands
- "(?i)healthcheck" # Health checks (case-insensitive)
- "INSERT INTO.*_staging" # Staging table writes
Regex Syntax¶
Patterns use Go's regexp syntax (RE2). Common patterns:
| Pattern | Description |
|---|---|
^SELECT |
Queries starting with SELECT |
(?i)pattern |
Case-insensitive matching |
system\\. |
Literal dot (escaped) |
table1\|table2 |
Match either table |
.* |
Match anything |
blacklist_queries expressions are not compiled by configuration loading.
That means both click-dog validate and click-dog check can succeed with an
invalid expression. The daemon/backfill path compiles them when it initializes
the query filter and exits before processing with Failed to initialize query
filter if any pattern is invalid. Test new blacklist expressions by starting
the target binary in a safe environment; redact_queries[].pattern does not
have this gap and is compile-checked during configuration validation.
YAML Escaping¶
Remember that backslashes need to be doubled in YAML strings:
# Correct - double backslash
blacklist_queries:
- "^SELECT \\* FROM system\\."
# Wrong - single backslash (YAML will consume it)
blacklist_queries:
- "^SELECT \* FROM system\."
Query Text Export Modes¶
Raw SQL is exported by default
query_text_mode defaults to raw for compatibility. Raw query text can
include tokens, email addresses, card numbers, or other literals. Use
normalized_only or none for privacy-sensitive production environments.
| Mode | Export behavior |
|---|---|
raw |
Retains the original db.statement / query-log query, subject to existing length limits. |
redacted |
Applies redact_queries. At least one valid rule is required. If no rule matches a statement, that statement is omitted instead of falling back to raw text. URI attributes carrying an HTTP query parameter and exception messages are dropped. |
normalized_only |
Removes raw statement fields, URI attributes carrying an HTTP query parameter, and exception messages, then retains a bounded, literal-stripped normalized preview only when ClickHouse safely provides one. If normalization is unavailable, no query text is exported. |
none |
Removes raw and normalized query text, URI attributes carrying an HTTP query parameter, and exception messages, while retaining query IDs, normalized hashes, tables, databases, timings, status, exception codes, users, clients, and resource attributes. |
Exception messages are dropped because ClickHouse quotes the failing statement,
literals included, in the message it records on the query span
(clickhouse.exception). Every privacy-restricting mode removes any span
attribute naming an exception except its code.
Normalized text stays under query_log.normalized_query,
db.normalized_query, or Splunk's normalized_query; it is never mislabeled
as raw db.statement. Scheduled native spans and query-log backfill use the
same boundary before OTLP and Splunk HEC fan-out.
For scheduled/native spans, normalized_only depends on query-log enrichment
to add query_log.normalized_query. Disabling
monitor.enrich_from_query_log produces a configuration warning. If the
ClickHouse normalized-query capability probe is unsupported or fails,
click-dog emits a mode-specific startup warning. Both cases remain fail closed:
query text is omitted.
Query Redaction¶
Rewrite sensitive fragments in SQL text before export without dropping the span.
Redaction rules use Go's RE2 regexp syntax, run after query blacklist filtering,
and are applied sequentially in query_text_mode: redacted. If replacement
is omitted, matches are replaced with [REDACTED].
filters:
query_text_mode: redacted
redact_queries:
- pattern: |
(?i)identified\s+by\s+'[^']*'
replacement: "IDENTIFIED BY '[REDACTED]'"
- pattern: "(?i)token\\s*=\\s*'[^']*'"
On the scheduled span path, successful redaction rewrites exactly one exported attribute:
db.statement. On the backfill path the query text is redacted before export
and surfaces per backend: as db.statement in OTLP spans, and as the top-level
query field in Splunk HEC events. Normalized attributes and log_comment.*
are not processed by redaction rules. Pass-through URI values containing an
HTTP query parameter are dropped in every privacy-restricting mode so a
URL-encoded raw copy cannot bypass the boundary. Arbitrary custom attributes
are not parsed as SQL. The mode does not change ClickHouse data and it does not
affect filter decisions. If none of the rules match, the statement is omitted. If a query
matches blacklist_queries, it is dropped entirely and no redacted copy is
exported.
For upgrade safety, an older config that has redact_queries but omits
query_text_mode migrates to fail-closed redacted mode and logs a warning
that unmatched statements are newly omitted. This is intentionally stricter
than the legacy redact-and-pass-through behavior. Set the mode explicitly.
Rules under raw remain a hard error because they would falsely imply that
redaction is active. Stale rules under normalized_only or none are inert
and produce a warning instead of preventing a privacy-tightening startup.
Duration Filtering¶
Duration thresholds are applied at the SQL level in ClickHouse, not in application code. This minimizes data transfer.
monitor:
# Trace-level, which traces to look at
min_trace_duration_ms: 1000 # Traces with at least one span >= 1s
max_trace_duration_ms: 300000 # Qualifying span must be <= 5 minutes
# Span-level, which spans to export from matching traces
min_span_duration_ms: 500 # Only export spans >= 500ms
max_span_duration_ms: 60000 # Skip spans > 1 minute
Two-Tier Approach¶
Step 1, trace selection: Find trace IDs where at least one span falls within [min_trace_duration_ms, max_trace_duration_ms].
Step 2, span export: From those traces, export individual spans within [min_span_duration_ms, max_span_duration_ms].
This finds traces that contain slow operations (step 1) while controlling which spans from those traces are exported (step 2).
Examples¶
Export all spans from traces containing a slow span:
min_trace_duration_ms: 5000 # Traces with 5s+ spans
min_span_duration_ms: 0 # But export ALL spans from those traces
Export only the slow spans themselves:
min_trace_duration_ms: 5000 # Traces with 5s+ spans
min_span_duration_ms: 5000 # Only export the 5s+ spans
Skip outlier traces (likely system operations):
min_trace_duration_ms: 1000
max_trace_duration_ms: 600000 # Qualifying span must be <= 10 minutes
# Note: this bounds the span that QUALIFIES the trace, not every span in it. A
# trace containing both a 2-second and a 15-minute span still qualifies via the
# 2-second span; use max_span_duration_ms to keep the long span from exporting.
Query Length Filtering¶
Bound excessively long SQL text. There are two separate stages, and they cover different paths:
monitor:
max_query_length: 100000 # Skip query_log rows with SQL > 100k chars
# (backfill and `analyze trace` candidates only)
exporters:
otel:
- collector_address: localhost:4317
max_query_length: 100000 # Truncate SQL > 100k chars (at export time)
splunk_hec:
- endpoint: https://splunk.internal:8088
token: ${SPLUNK_HEC_TOKEN}
max_query_length: 100000 # Truncate SQL > 100k chars (at export time)
| Setting | When Applied | Behavior |
|---|---|---|
monitor.max_query_length |
During the query_log fetch (backfill, analyze trace candidate search) |
Rows with SQL text longer than this are excluded entirely |
exporters.otel[].max_query_length |
During OTEL export | SQL text is truncated to this length with ... appended |
exporters.splunk_hec[].max_query_length |
During Splunk HEC export | SQL text is truncated to this length with ... appended |
Scheduled mode does not exclude long queries at fetch time
The length(query) predicate is built only by the shared query_log query
builder, which backfill and the analyze trace candidate search use. The
scheduled span-log fetch has no equivalent, so monitor.max_query_length
never drops a live span. Its only effect on that path is truncating the
enriched query_log.normalized_query attribute. The analyze queries
family rollups do not take this key. They bound preview text with
their own limit instead. What bounds outbound SQL in scheduled mode is the
exporter-level max_query_length, which truncates db.statement on both
the live and backfill paths.
Combining Filters¶
All filters can be used together. Here's a production example:
monitor:
min_trace_duration_ms: 1000
max_trace_duration_ms: 600000
min_span_duration_ms: 500
max_query_length: 50000
filters:
# Only care about query execution
whitelist_operations:
- "DB::Interpreter*::execute()"
# Only from application servers
whitelist_ips:
- "10.0.1.50"
- "10.0.1.51"
- "10.0.1.52"
# Exclude noise
blacklist_queries:
- "^SELECT \\* FROM system\\."
- "(?i)healthcheck"
- "SHOW TABLES"
- "INSERT INTO.*_tmp_"
Debugging Filters¶
Set log_level: debug to see which spans are being filtered and why:
[DEBUG] Filtering span from IP not in whitelist: 172.16.0.5
[DEBUG] Filtering operation not in whitelist: TCPHandler::runImpl()
[DEBUG] Filtering query matching blacklist pattern: ^SELECT \* FROM system\.
[DEBUG] Filtering span from user not in whitelist: "patient_records"
[DEBUG] Filtering span from blacklisted user: "ml_training"