fix(infra): upgrade otel-collector to 0.149.0 and move logs to OTLP - #449
Merged
shmsr merged 1 commit intoSep 22, 2026
Merged
Conversation
The lab collector was pinned at 0.120.0, whose elasticsearch exporter drops ~43% of metric data points on a TSDS target. Over identical 10-minute windows against the same node_exporter scrape, Prometheus recorded 41 samples of node_memory_Dirty_bytes at a clean 15s cadence while Elasticsearch received 22-23, with surviving timestamps alternating 15s/30s. Isolated to the exporter, not the scrape: the metrics pipeline fans out to a prometheus exporter and an elasticsearch exporter behind the same receiver and the same batch processor, and Prometheus scrapes the former as job otel-prometheus-exporter. Counting distinct value changes, that path retained 40/40 while the elasticsearch path retained 23, so the receiver and batch processor are lossless. Elasticsearch was not rejecting (write thread pool rejected=0) and nothing was overwritten (every recent document _version=1). Upgrading to 0.149.0 raises retention to ~95% and leaves the target field layout byte-identical (metrics.*, resource.attributes.*, data_stream.*), so no re-migration or --field-profile change is needed. The dedicated loki exporter was removed from collector-contrib after v0.120, so the bump alone crash-loops the collector with 'exporters unknown type: "loki"'. Loki 3.4.2 ingests OTLP natively, so the logs pipeline now uses otlphttp against http://loki:3100/otlp. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
8 tasks
This was referenced Sep 22, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #453
Summary
0.120.0→0.149.0. The 0.120elasticsearchexporter drops ~43% of metric data points on a TSDS target.lokiexporter withotlphttp/loki, becauselokiwas removed from collector-contrib after v0.120 and the bump alone crash-loops the collector.Evidence for the data loss
Same node_exporter scrape, identical 10-minute windows,
node_memory_Dirty_bytes:Isolated to the exporter rather than the scrape. The metrics pipeline fans out to a
prometheusexporter and anelasticsearchexporter behind the same receiver and same batch processor, and Prometheus scrapes the former as jobotel-prometheus-exporter. Counting distinct value changes:One exporter is lossless off identical upstream input, so the receiver and batch processor are exonerated. Elasticsearch was not applying backpressure (
_cat/thread_pool/write→rejected=0) and nothing was overwritten (every recent doc_version=1).After the bump: retention ~95%.
Why the second change is required
0.149.0alone fails at startup:Loki 3.4.2 ingests OTLP natively, so the logs pipeline now uses
otlphttpagainsthttp://loki:3100/otlp. Logs verified flowing afterwards.Field layout is unchanged
metrics.*,resource.attributes.*anddata_stream.*are byte-identical before and after, so no re-migration or--field-profilechange is needed and existing uploaded dashboards keep working.Not fully resolved
version_conflict_engine_exceptionstill appears after the bump, but now lands on the cadvisor firehose rather than node_exporter (which measures 19/20 samples). Worth revisiting if cadvisor-backed panels start mattering.Test plan
Everything is ready, health endpoint 200,:8889serving)🤖 Generated with Claude Code