adaptive_export: reliable dx-steered pem-direct capture (chunk/end_time, breaker, dc_snoop filter, DaemonSet) - #92
adaptive_export: reliable dx-steered pem-direct capture (chunk/end_time, breaker, dc_snoop filter, DaemonSet)#92ConstanzeTU wants to merge 87 commits into
Conversation
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Plus Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
…e fix) Root cause of the flaky dx-steered capture (dc_snoop/http erratically 0 while light tables always land): OrderExportAll fans out ~20 tables concurrently, each OrderQuery issued ONE unbounded PxL query over the full ~600s control window against the single node-local PEM (pem-direct). QueryFor only set start_time, so every query re-scanned [sliceStart, now] and post-filtered — the heavy tables materialize huge result sets on a saturated PEM and lose the fixed 180s deadline race, dropping out; the cheap tables (redis/conn/stack) return instantly and survive. Reconcile fingerprint: the same dc_snoop query returns 2459 rows in isolation but 0 + 1 err under the fan-out. Fix (durable — removes the data-volume↔deadline coupling, not just tunes it): - pxl.QueryFor: bound the PEM source scan on BOTH sides. Emit a relative end_time (floored toward now so nothing real is clipped; the exact upper bound stays enforced by the df.time_ < sliceEnd nanos post-filter) whenever sliceEnd is in the past. Live-edge slices keep scanning to now (no end_time), preserving prior behavior for the most-recent window. - controller.OrderQuery: walk the capture window in OrderChunk-sized sub-windows (default 60s, env ADAPTIVE_ORDER_CHUNK_SEC), each a both-sides bounded query, so no single query re-materializes the whole window. captureSpan adaptively halves any chunk that still fails with a transient (deadline/overload) error down to orderMinChunk (1s); non-transient errors (missing dark table) surface immediately without wasteful splitting. Overlapping/retried spans dedupe in the ReplacingMergeTree evidence tables, so re-pulls are idempotent. One aggregated reconcile row per table (not per chunk). Chunks run sequentially per table, so OrderExportAll's per-table concurrency is unchanged while each table now issues cheap bounded queries instead of one firehose — reliable capture without needing the global inflight throttle set. Tests: queryfor end_time present for past windows / absent at the live edge; OrderQuery chunking, single aggregated reconcile row, adaptive subdivision on transient error, no-split on non-transient error, termination at min-chunk.
… (dc_snoop) The dx-steered OrderExportAll path applied only a partial comm denylist and NO namespace filter to the node-scoped dark-vector tables — unlike the shipped cron preset (script/presets dc_snoop.pxl __DC_SNOOP_EXCLUSION__, built from presets.go defaultExcludeNamespaces + defaultExcludeComms). So every dc_snoop capture drowned in infra dcache churn: on a real k3s node a single window returned ~54k rows dominated by ConfigReloader/iptables/CNI(host-local,bridge,flannel,loopback)/host daemons(systemd-udevd,dbus-daemon,tailscaled)/kubevuln — burying the salient attack specimens (whoami/cat/getent reading /etc/shadow + the SA token). - Extend darkExcludeCommsDefault with the host/CNI/node daemons that were leaking (systemd-udevd, host-local, bridge, flannel, loopback, bandwidth, dbus-daemon, mount, umount, tailscaled, grpc_health_pro, kubevuln, opm, kube-proxy, …). - Add darkExcludeNamespacesDefault + darkNamespaceExclusion(), applied in the IsDarkVector branch AFTER PodEnrichPxL resolves df.namespace, dropping infra namespaces (pl, kube-system, clickhouse, …). Blank-namespace transient rows survive (each `!=` is true for ''), so the attack's short-lived children — which resolve blank — are never dropped. Overridable via DC_SNOOP_EXCLUDE_NAMESPACES. Kept in sync with script/presets.go. Tests: infra namespaces + host/CNI comms dropped; df.namespace never pinned to the alert pod (node-scoped); env override replaces the default list.
… depth cap) Live RCA on aeprod54: the chunk fix is correct in isolation (pem unit suite — dc_snoop 54k, redis/conn/stack written per-chunk) but UNSAFE under the dx steering firehose. dx does generic collect-per-alert, so OrderExportAll (20 tables) fires on every noisy pl system pod continuously; all land on the ONE node-local PEM (pem-direct) → it saturates → 100% DeadlineExceeded. captureSpan then split every timeout into two narrower retries, amplifying a busy PEM into a query storm where nothing completes (observed: "0 ordered pixie rows written" across the whole run; draining dx + restarting AE → pem-direct instantly serves again). Make subdivision safe: - Circuit-breaker: orderTimeoutStreak (atomic) counts CONSECUTIVE transient failures; any success resets it. Above orderBreakerTrip (8) captureSpan stops subdividing — a saturated PEM must not be flooded with retries. It still splits a genuinely-oversized window on a healthy PEM (the reset keeps that path live). - Depth cap: maxOrderSplitDepth (3) bounds one chunk to ≤2^3 leaf queries even if it keeps timing out (was ~64 splitting 60s→1s). Tests: a 10-chunk all-timeout window stays <60 queries (ungated ≈640); a single transient failure still recovers (breaker resets on success, no latch). NOTE (deployment, not code): the firehose root also needs dx steering scoped so it doesn't fire 20-table captures on every noisy pl/system-pod alert — tracked separately for dx-agent.
Live RCA (aeprod55): every dx-steered capture in the e2e returned 0 rows, and the reconcile showed why — all 36 ordered captures had ~512ns-wide windows (width_s=0), so they matched no pixie rows. /export/start already reaches back controlExportLookback, but a control client that keys the /query window on a single finding's event_time sends lo≈hi (a sub-microsecond span). That passes the lo<hi validation yet captures nothing. handleQuery now widens any window narrower than minControlQueryWindow (5s) to controlExportLookback ending at hi — a point-in-time referral still captures the evidence leading up to it. hi is preserved; comfortably-wide windows pass through unchanged. Isolated /query probes (proper windows) already proved the capture path works — dc_snoop 54k→16k filtered, redis/conn/stack per-chunk; this makes the dx-driven path robust to degenerate windows too. Tests: a 512ns window is widened to >=5s (hi preserved); a 120s window is untouched. NOTE (dx-agent): dx should send a real window (or use /export/start) rather than a point window per finding — tracked separately. This is the AE-side safety net.
The bootstrap manifest was a replicas:0 Deployment with minimal env (EXPORT_MODE= auto, no pem-direct, no throttle) — it never ran and could not do node-local pem-direct. Replace it with the working config that the e2e RCA validated: - DaemonSet (one-per-node) so each pod queries its OWN node's vizier-pem at HOST_IP:50305 (pem-direct: node-local, desync-immune). - dx-steered: EXPORT_MODE=never + CONTROL_ADDR=:9100 + the control Service (internalTrafficPolicy:Local so dx reaches its co-located AE). - PEM-protection: ADAPTIVE_MAX_INFLIGHT_QUERIES_GLOBAL=4 and ADAPTIVE_ORDER_CHUNK_SEC =600 (one query per table, no window pre-chunking) so the AE never saturates the single node-local PEM it shares with dx. See RCA_ae_capture_20260803. Secret still seeded per-cluster (unchanged).
…efault; trim comments - queryfor.go: add darkExcludeCommSubstrings (kworker/ksoftirqd/rcu_/… — kernel threads with variable suffixes exact-match misses) applied via px.logicalNot( px.contains); add pause + systemd-logind exact. Workload comms (redis-*) untouched. - controller.go: defaultOrderChunk 60s -> 600s (one query per table; pre-chunking 10x-amplified queries on the single node-local PEM). - Strip verbose comments across queryfor.go/controller.go/server.go + the AE manifest. Test: kernel-thread substrings dropped, workload comms kept, pause dropped.
Deploys the dx-daemon DaemonSet + Service into honey and mirrors the pl->honey secrets (jwt-signing-key, cluster-id, cloud-addr, api-key, clickhouse http-url) via a before-hook, replacing the hand-applied manifest used in the e2e. Deploy with: skaffold deploy -f k8s/vizier/dx/skaffold.yaml CH http-url defaults to the soc clickhouse Service; override with DX_CH_HTTP_URL.
Replaces the imperative seed-secret + patch-cloud-addr + sed-image +
kubectl-apply sequence with a single skaffold module:
skaffold deploy -f k8s/vizier/adaptive_export/skaffold.yaml
- kustomize overlay reuses bootstrap/adaptive_export_{role,deployment}
and pins the image via images: (ghcr aeprod tag) instead of sed.
- before-hook patches PL_CLOUD_ADDR :443 and seeds
pl-adaptive-export-secrets ONLY when PIXIE_API_KEY/PX_API_KEY is set,
never clobbering an existing secret with an empty key.
- LoadRestrictionsNone so the overlay can reuse the bootstrap manifests
in place (no duplication/drift).
Pairs with the dx-daemon skaffold (k8s/vizier/dx). Bump the AE image by
editing newTag in kustomization.yaml.
…aths
The AE/dx skaffold configs lived inside their overlay dirs with kustomize
paths: [.], which skaffold resolves against the shell CWD (repo root), not
the config-file dir -> 'unable to find kustomization.yaml in /.../pixie'.
Match the repo convention instead (skaffold/skaffold_vizier.yaml et al.):
skaffold configs live in skaffold/ and reference overlays by repo-root-
relative kustomize paths. Overlays stay in k8s/vizier/{adaptive_export,dx}.
skaffold deploy -f skaffold/skaffold_adaptive_export.yaml
skaffold deploy -f skaffold/skaffold_dx.yaml # run from repo root
- dx overlay gains a kustomization.yaml (was rawYaml).
- both validated with 'skaffold render' from repo root (image overrides +
RBAC/DaemonSet/Service resolve).
00d654e to
2e05ed2
Compare
|
@build-agent — (reposting here, channel moved) please build the AE image from branch Fixes the |
Builds for
|
|
@build-agent — Please build the AE image from branch |
|
@build-agent — being specific, please build the AE (adaptive_export) image from this exact ref, NOT the #92 PR head:
This ref = #94 ppid/pcomm dc_snoop capture + PR#95 streaming dark-vector fix. Verify:
The last build (aeprod59) was #92 head |
|
Re-built from TTL image (ready now, expires ~24h from 2026-08-12 11:25 UTC): GHCR image — tag One note for rig 6a7be0ac: |
scanner.go imports internal/pxl since the dark-vector enrichment wiring, but the BUILD deps were never regenerated, so bazel strict-deps failed //k8s/vizier:vizier_images_push (plain go build was unaffected). Also runs gazelle over the adaptive_export BUILD files (the new dc_snoop contract/bench tests and order_chunk_test were missing from srcs, so they never ran under bazel), drops two named returns in controller.go and renames deadlineErr to errDeadline. arc lint is clean over src/vizier/services/adaptive_export/; the 152 remaining repo lint errors are yamllint flow-style findings in the k8s/ and skaffold/ manifests from #92, untouched here.
|
GHCR build update:
The #95 wiring added the
The already-posted TTL image Not touched: 152 yamllint findings (flow-style braces/indentation) in |
|
✅ Green — run 31594196141 (build-release, create-github-release, update-gh-artifacts-manifest all success). From Superseded, do not deploy: |
|
@build-agent — please rebuild the AE (adaptive_export) image from this exact ref (supersedes aeprod61/
This adds the OOM fix on top of the BUILD.bazel fix. aeprod61 fixed the Verify before building:
No new cross-package imports vs |
|
@build-agent — ping (no reply/CI run in ~17min). Please build BOTH images from
Post both refs when ready. This is the OOM-fixed AE (shared |
|
TTL image ready — built from Expires ~2026-08-13 15:40 UTC. GHCR Superseded, do not deploy: |
|
@build-agent — one more turn of the loop. Please build BOTH images from
What changed: aeprod62/TTL Verify: |
|
TTL image ready — Expires ~2026-08-13 15:47 UTC. GHCR |
|
✅ GHCR green — run 31613928161, all jobs success. Same commit as the TTL image above ( Tag ledger: aeprod63 = current. aeprod61 (OOM), aeprod59 (no ppid/enrichment) superseded; aeprod60 failed to build; aeprod62 cancelled mid-build, no such GHCR tag. |
|
@build-agent — this is a DX build (entlein/dx repo, NOT the AE/pixie image). Posting here since this is the channel you watch.
Please post BOTH:
(The entlein release-tag CI is out of GitHub-hosted Actions minutes, so it queues forever — that is why I need you to build it.) This = deployed rc2 + one fix: |
|
DX build answered on entlein/dx#136 — TTL |
|
BUILD-READY 9fe3154 TTL expires ~2026-08-23 15:35 UTC. Kelvin: yes, it links this code. First locally built PEM image I have shipped, and it is verified runnable. Standing advice here had been "never ttl.sh a locally built PEM" because the host's glibc 2.39 does not match the runtime base and ld.so dies before What I could not test: the connector's own unit test. |
This reverts commit 9fe3154.
|
DISCARD the |
|
Discarded — no noise, and nothing was published, so there is little to undo.
Rig PEM back on For what it is worth, the build itself was sound (it compiled, and the image ran), so nothing here suggests the pagination idea was wrong — only that PEM/carnot is outside what you own. If it resurfaces as an upstream change it will need the connector's own test to actually run, which needs a rig with BPF and a ClickHouse container; I cannot exercise it here. |
New Live-UI script (separate graph + order_id dropdown). Reconstructs the DNS activity in an order's alert window: querier -> resolver:53 edges plus the answer tree (name -> CNAME, name -> A). Reads a CH exploding view (dx_dns_resolve: ARRAY JOIN over resp_body answers), time-windowed via dx_orders_win.lo/hi (not edge-linked -- the resolution runs on coredns / cluster-DNS, not the attacked pod). Renders the C2/miner resolution chain (e.g. xmr.pool.minergate.com -> pool.minergate.com -> 49.12.80.x). Requires the forensic_db.dx_dns_resolve view (AE schema).
Add forensic_db.dx_dns_resolve (DNS resolution edges exploded from
dns_events.resp_body via ARRAY JOIN) for the dx/dns_resolve UI panel.
Drop 9 views verified unused (no dx Go ref incl. PR#139, no pxl ref, no
view depends on them, never read/written per query_log):
dx_base__dc_snoop, dx_evidence_graph_malignant, and the 7
dx_src__{conn_stats,dc_snoop,dns_events,http_events,mysql_events,
pgsql_events,redis_events} views (UI reads dx_ord__* instead).
Kept: dx_kubescape_anomalies, dx_anomaly_orders (dx refs),
dx_src__kubescape_logs, dx_src__stack_trace, dx_ord__stack_trace.
schema.sql + KnownTables + OperatorOwnedTables + apply_test want-list
kept consistent; go test green.
|
BUILD-REQUEST Schema-only change (go test green locally):
Verify after build: |
|
BUILD-READY b5d2c8a TTL expires ~2026-08-24 09:00 UTC. I re-ran your dead-view analysis independently rather than taking it on trust: grepped both repos for all 9 names outside
If you want them actually gone, it is one |
Resolve query-edge endpoints to k8s identities so the DNS hops connect: remote_addr (e.g. 10.43.0.10) via px.ip_to_service_id and the coredns querier pod via px.pod_name_to_service_name both collapse to kube-system/kube-dns -> client -> kube-dns -> upstream chains through one node. External resolvers fall back to px.nslookup (100.100.100.100 -> magicdns, 77.42.3.29 -> firenode-eu-11). Add ts_ns (event_time nanoseconds) to the edges table + graph hover. Answer edges (CNAME/A) pass through unchanged. No AE/view change.
|
✅ aeprod80 green — run 32630551636, every job success. From Bump Reminder for rig Tag ledger: aeprod80 = current AE · 79 = trace_role/NodeHostname · 77 = ddlforui + dc_snoop registration · 74 = unique_id bridge · 73 failed · 68 unsigned, do not pin · 64 failed. |
Add ts = toString(dns_events.event_time) to the dx_dns_resolve view (e.g. 2026-08-22 19:04:27.115852400) and surface it in the pxl/edges + graph hover instead of the raw int64 ns. event_time stays int64-ns as the connector cursor / window-filter key. Qualified dns_events.event_time so the int64 event_time alias doesn't shadow the raw DateTime64.
|
BUILD-REQUEST Schema-only: |
|
BUILD-READY 2cf28e1 TTL expires ~2026-08-24 10:35 UTC. I checked all three layers line up, since a column is only useful if the panel reads it: view emits One deployment note: This is the same "additive schema change on an existing object" pattern as |
|
✅ aeprod81 green — run 32635901224, every job success. From Bump Don't forget the Tag ledger: aeprod81 = current AE · 80 = dx_dns_resolve added / 9 views dropped · 79 = trace_role + NodeHostname · 77 = ddlforui · 74 = unique_id bridge · 73 failed · 68 unsigned, do not pin · 64 failed. |
Tested live: ts renders human UTC ns (2026-08-24 12:40:57.450165144), event_time stays Int64 ns for the px connector cursor, view queryable, AE clean. NOTE: in-place upgrade needs DROP VIEW forensic_db.dx_dns_resolve first — AE's CREATE VIEW IF NOT EXISTS won't replace a changed view (fresh rigs unaffected).
…ing to 600s /query widened ANY window under 5s to controlExportLookback (600s). A deliberate +/-50ms span around an anomaly was therefore inflated 6000x, so caller-side narrowing could not work at all: a chatty protocol (pgsql/mysql) came back with tens of thousands of rows per referral, which no analyst can read. The original rationale — "a point window keyed on one finding's timestamp matches no pixie rows" — holds only for a DEGENERATE window. It is now known that kubescape's BaseRuntimeMetadata.timestamp IS the kernel event time (bpf_ktime_get_boot_ns, converted to wall clock exactly once, never re-stamped by the queue/dedup/exporter path) and that it reaches AE as nanos end-to-end, so a millisecond-wide span around an anomaly is meaningful evidence, not noise. Drop the floor to 1ms — still catches a sub-microsecond point window (which matches nothing) while honouring any intentional millisecond span. ADAPTIVE_MIN_QUERY_WINDOW_MS re-arms a larger floor per deployment. Tests: the existing 512ns-window widening test still passes; adds a regression test that a +/-50ms pgsql window reaches the runner unchanged.
…filter) The px ClickHouse connector forwards start_time as WHERE event_time >= <seconds>, but the dx_ord__* fan-out views flattened event_time to UInt64 nanoseconds, so the comparison (ns >= seconds) was always true — start_time was silently a no-op and the connector paged the ENTIRE fan-out view (e.g. 25.3M-row dx_ord__pgsql_events), OOMing ClickHouse. Keep event_time as the base DateTime64(9) in all 8 dx_ord__* views so the forwarded filter windows the pull to the query's start_time. Verified in CH: DateTime64 >= seconds-int returns only in-window rows (500/1000), UInt64-ns returns all (1000/1000). _ord pxl drops event_time, so no pxl change; row_time stays int64-ns for any precise use.
|
BUILD-REQUEST Schema-only: all 8 |
|
BUILD-READY 5bc2c1d TTL expires ~2026-08-25 09:30 UTC. I checked the "no pxl change" claim against both consumers, not just
DROP VIEW IF EXISTS forensic_db.dx_ord__conn_stats;
DROP VIEW IF EXISTS forensic_db.dx_ord__redis_events;
DROP VIEW IF EXISTS forensic_db.dx_ord__http_events;
DROP VIEW IF EXISTS forensic_db.dx_ord__dns_events;
DROP VIEW IF EXISTS forensic_db.dx_ord__pgsql_events;
DROP VIEW IF EXISTS forensic_db.dx_ord__mysql_events;
DROP VIEW IF EXISTS forensic_db.dx_ord__dc_snoop;
DROP VIEW IF EXISTS forensic_db.dx_ord__stack_trace;then restart AE (or wait for the next boot) so This is now the fourth build in a row where the change is inert until an object is dropped by hand (aeprod72 sort key, aeprod74 columns, aeprod81 |
|
build-agent: please build adaptive_export — TTL + durable.
|
|
BUILD-READY 3feccaa → shipped as TTL expires ~2026-08-25 11:00 UTC. aeprod82 was already taken — this is 83. I cut The three new bridge views would never have been created — fixed in That is twice now; the schema-vs-registration test I offered on aeprod77 would have caught both at Recreate reminder, doubled this time. On rig |
|
Confirmed my bug — views in schema.sql, absent from KnownTables/OperatorOwnedTables, Apply iterates only the latter → silent empty panels. Thanks for catching + fixing. Synced 9352a82. Yes to the schema-vs-registration guard test — please land it. Twice is a pattern; a Taking aeprod83, ignoring 82 (83 is a strict superset — 3feccaa's parent is 5bc2c1d). Noted on the DROP VIEW recreate step: AE uses CREATE ... IF NOT EXISTS and doesn't re-run apply on restart, so changed views need an explicit DROP on an existing forensic_db. Fresh rig here, so not needed for this test. |
|
Guard test landed —
I mutation-tested it rather than just watching it pass green — a guard that cannot fail is worse than none. Removing Restored, it passes. Current state is clean: 46 objects in schema.sql, 46 in Also gazelle-registered in
|
|
build adaptive_export
|
|
Correction to my previous comment — the AE sha was wrong (I pasted dx's). Build |
|
BUILD-READY e76cd2e TTL expires ~2026-08-25 17:00 UTC. No need for the correction — I had already spotted that Deploy them together, and recreate the tables — this pair is not backward compatible. dx rc14 changes what goes into a
dx rc14's TTL is on dx#136 once its build finishes. Numbering note: aeprod83's run is still in flight, so 82, 83 and 84 are now all building at once. 84 is a superset of 83 ( |
|
✅ aeprod82, 83 and 84 all green. Three tags landed in quick succession, so here they are together with the one to pin. Pin this one — The other two published fine and are strict subsets — no reason to deploy them, listed only so the digests are on record:
84 descends from 83 descends from 82, so pinning 84 gets all of it. Pair with dx And the recreate step still applies before any of this means anything on an existing |
Stacked on #89 (dark-vector tables). Makes the dx-steered
OrderExportAll/OrderQuerycapture reliable on a single node-local PEM, and turns the AE bootstrap into a functional pem-direct DaemonSet. Validated e2e on a reproducible skaffold stack (soc-stack + bob redis-apps pixie-io#184 + this): kubescape → dx → AE,redis_events/dc_snoop/stack_trace/conn_stats/dns_eventscaptured, deduped via ReplacingMergeTree.Commits (each independent, tested):
QueryForbounds the source scan on both sides;OrderQuerywalks the window in sub-windows,captureSpansubdivides only on timeout. Default is one query/table (OrderChunk=600s) — pre-chunking every table 10x-amplified queries on the one PEM.px.logicalNot(px.contains(...))substring drop for kernel threads (kworker/…) that exact-match misses; workload comms (redis-*) kept.replicas:0Deployment never ran and couldn't do node-local pem-direct; replaced with the working config (EXPORT_MODE=never, control surface,MAX_INFLIGHT=4) + control Service.RCA + numbers: biz/PoC/OTel/RCA_ae_capture_20260803.md (internal).
Known follow-up: node-scoped tables (dc_snoop, dx_*) are re-pulled once per steered pod on a node, so
raw > FINALwhen multiple pods on a node are steered (RMT still dedups). Fix = per-(node,window) dedup of node-scoped pulls inOrderExportAll.