Skip to content

fix(infra): upgrade otel-collector to 0.149.0 and move logs to OTLP - #449

Merged
shmsr merged 1 commit into
elastic:mainfrom
stefans-elastic:fix/otel-collector-0149-loki-exporter
Sep 22, 2026
Merged

shmsr merged 1 commit into
elastic:mainfrom
stefans-elastic:fix/otel-collector-0149-loki-exporter

Conversation

@stefans-elastic

@stefans-elastic stefans-elastic commented Sep 18, 2026 •

Copy link
Copy Markdown
Collaborator

Closes #453

Summary

  • Bumps the lab collector 0.120.0 → 0.149.0. The 0.120 elasticsearch exporter drops ~43% of metric data points on a TSDS target.
  • Replaces the loki exporter with otlphttp/loki, because loki was removed from collector-contrib after v0.120 and the bump alone crash-loops the collector.

Evidence for the data loss

Same node_exporter scrape, identical 10-minute windows, node_memory_Dirty_bytes:

source samples / 10 min cadence
Prometheus 41 clean 15s
Elasticsearch (via OTel) 22–23 alternating 15s/30s

Isolated to the exporter rather than the scrape. The metrics pipeline fans out to a prometheus exporter and an elasticsearch exporter behind the same receiver and same batch processor, and Prometheus scrapes the former as job otel-prometheus-exporter. Counting distinct value changes:

path retained
Prometheus direct (ground truth) 40 changes
receiver → batch → prometheus exporter 40 — 100%
receiver → batch → elasticsearch exporter 23 — 57%

One exporter is lossless off identical upstream input, so the receiver and batch processor are exonerated. Elasticsearch was not applying backpressure (_cat/thread_pool/write → rejected=0) and nothing was overwritten (every recent doc _version=1).

After the bump: retention ~95%.

Why the second change is required

0.149.0 alone fails at startup:

Error: failed to get config: cannot unmarshal the configuration:
'exporters' unknown type: "loki" for id: "loki"

Loki 3.4.2 ingests OTLP natively, so the logs pipeline now uses otlphttp against http://loki:3100/otlp. Logs verified flowing afterwards.

Field layout is unchanged

metrics.*, resource.attributes.* and data_stream.* are byte-identical before and after, so no re-migration or --field-profile change is needed and existing uploaded dashboards keep working.

Not fully resolved

version_conflict_engine_exception still appears after the bump, but now lands on the cadvisor firehose rather than node_exporter (which measures 19/20 samples). Worth revisiting if cadvisor-backed panels start mattering.

Test plan

  • Collector starts clean at 0.149.0 (Everything is ready, health endpoint 200, :8889 serving)
  • node_exporter metric retention measured at ~95% vs ~57% before
  • Target field layout confirmed byte-identical
  • Logs pipeline confirmed working over OTLP
  • Grafana and Kibana both still serving; a previously-uploaded dashboard unaffected

🤖 Generated with Claude Code

The lab collector was pinned at 0.120.0, whose elasticsearch exporter drops
~43% of metric data points on a TSDS target. Over identical 10-minute windows
against the same node_exporter scrape, Prometheus recorded 41 samples of
node_memory_Dirty_bytes at a clean 15s cadence while Elasticsearch received
22-23, with surviving timestamps alternating 15s/30s.

Isolated to the exporter, not the scrape: the metrics pipeline fans out to a
prometheus exporter and an elasticsearch exporter behind the same receiver and
the same batch processor, and Prometheus scrapes the former as job
otel-prometheus-exporter. Counting distinct value changes, that path retained
40/40 while the elasticsearch path retained 23, so the receiver and batch
processor are lossless. Elasticsearch was not rejecting (write thread pool
rejected=0) and nothing was overwritten (every recent document _version=1).

Upgrading to 0.149.0 raises retention to ~95% and leaves the target field
layout byte-identical (metrics.*, resource.attributes.*, data_stream.*), so no
re-migration or --field-profile change is needed.

The dedicated loki exporter was removed from collector-contrib after v0.120, so
the bump alone crash-loops the collector with 'exporters unknown type: "loki"'.
Loki 3.4.2 ingests OTLP natively, so the logs pipeline now uses otlphttp
against http://loki:3100/otlp.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

infra lab: OTel collector 0.120.0 elasticsearch exporter drops ~43% of metric samples, making dashboard comparison unfair

2 participants