feat(metrics): use native request latency histogram - #11251
Conversation
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
This breaking metrics migration has broad exporter and dashboard compatibility implications requiring human review.
Review tier: Lite
Findings: None
What changed in this PR
Replaces custom request-latency counters with a native Histogram<double> using fractional millisecond observations and explicit bucket boundaries.
Changes:
- Adds native histogram instrumentation and comprehensive behavior tests.
- Removes obsolete aggregators and related tests.
- Updates monitoring documentation and migration guidance.
| File | Description |
|---|---|
test/Orleans.Runtime.Tests/Orleans.Runtime.Tests.csproj |
Adds histogram test dependencies. |
test/Orleans.Runtime.Tests/Diagnostics/ApplicationRequestInstrumentsTests.cs |
Tests histogram behavior and boundaries. |
test/Orleans.Runtime.Tests/CallbackDataTests.cs |
Verifies fractional callback latency. |
test/Orleans.Core.Tests/General/HistogramAggregatorTests.cs |
Removes obsolete aggregator tests. |
src/Orleans.Core/Runtime/CallbackData.cs |
Preserves fractional elapsed milliseconds. |
src/Orleans.Core/Diagnostics/Metrics/ApplicationRequestInstruments.cs |
Implements native latency histogram recording. |
src/Orleans.Core/Diagnostics/Metrics/Aggregators/HistogramBucketAggregator.cs |
Removes obsolete bucket aggregation. |
src/Orleans.Core/Diagnostics/Metrics/Aggregators/HistogramAggregator.cs |
Removes obsolete aggregation. |
docs/site/src/content/docs/host/monitoring/troubleshooting-symptom-catalog.md |
Updates latency troubleshooting guidance. |
docs/site/src/content/docs/host/monitoring/metrics.md |
Documents the new metric and migration. |
docs/site/src/content/docs/host/monitoring/metrics-catalog.md |
Updates the metrics catalog. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Replace the custom request-latency observable counters with a standard Histogram<double>, preserve fractional milliseconds, and document the breaking exporter contract. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Align runtime tests, exporters, instrumentation, documentation snippets, and samples with the OpenTelemetry release used for the request histogram performance comparison. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Remove Orleans-defined latency boundaries and document deployment-level Views with aligned, application-specific buckets. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
968f605 to
4ac5043
Compare
Code coverage
Report-only conclusion: mixed. The current-main baseline is commit Coverage combines every CI test matrix job, including providers, CodeGen, .NET 8/10, Linux, Windows, and macOS, using canonical physical source and branch identities. The comparison remains report-only while normal line and branch variance is calibrated. Coverage details |

Problem
Application request latency is currently preaggregated into three observable counters (
-bucket,-count, and-sum). This was introduced in #7814 in 2022 because native OpenTelemetry histogram aggregation was prohibitively expensive at the time, and the concern was reiterated in #9609.The custom representation has non-standard semantics: bucket series are disjoint ranges labeled with
duration, collection is process-lifetime cumulative, and the sum and bucket selection truncate elapsed time to integer milliseconds. It also prevents exporters from treating request latency as a normal histogram distribution.Change
Replace the custom aggregation with one
System.Diagnostics.Metrics.Histogram<double>namedorleans-app-requests-latency, with unitms, no metric attributes, and untruncatedTimeSpan.TotalMillisecondsobservations.Orleans does not prescribe bucket boundaries. Latency objectives and useful long-tail coverage are application-specific, so deployments should configure this instrument using an exact-name OpenTelemetry View. Instances whose distributions are aggregated together should use the same boundaries. A coarse general-purpose millisecond profile could be:
This is an example, not an Orleans default. Applications should add thresholds corresponding to their own SLOs and timeout policy.
The custom aggregator and its tests are removed. Correctness coverage verifies fractional timing through
CallbackData, disabled instrumentation, application-configured boundaries, exact-boundary and overflow behavior, floating-point count/sum/MinMax, cumulative and delta readers, and concurrent recording.This also updates the repository's OpenTelemetry packages, exporters, instrumentation, documentation snippets, and affected samples from 1.16 to 1.18. OpenTelemetry 1.18 contains the histogram synchronization optimization from open-telemetry/opentelemetry-dotnet#7601 and is the version used for the sustained-load comparison below. On .NET 10,
System.Diagnostics.DiagnosticSourceis supplied by the shared runtime (10.0.12 in the test environment), and OpenTelemetry 1.18 does not add a separate DiagnosticSource package dependency for net10, so this PR does not add a redundant framework-package reference.Compatibility
This is an intentional breaking metrics change with no dual-publication period:
orleans-app-requests-latency-bucket,orleans-app-requests-latency-count, andorleans-app-requests-latency-sumas separate Orleans instruments.durationattribute and its mutually exclusive range semantics.lebuckets plus count and sum series.Exporter Views, dashboards, alerts, recording rules, and queries must be updated together with the Orleans upgrade. Use the same explicit boundaries across instances or regions which must be aggregated exactly. Differing explicit histograms can be coarsened only where their boundaries are compatible; finer resolution cannot be reconstructed after collection.
Performance evidence
Benchmarks were run as a bounded prototype and are intentionally not included in this commit. Evidence and validation are focused on .NET 10.
Environment: BenchmarkDotNet 0.15.6, macOS 26.6.2, Apple M4 Max (14 physical/logical cores), .NET SDK 10.0.401, .NET 10.0.12, Server GC, and OpenTelemetry 1.18. Each sustained-load run used one client and one silo in the same process, 100 workers, a 150,000-request warmup, ten 500,000-request rounds, one reader, cumulative temporality, one-second export cadence, MinMax enabled, and no request tags. Three alternating runs were collected per implementation.
The native benchmark used 22 explicit boundaries, more than the coarse example above, so it is conservative with respect to explicit-bucket lookup and retained bucket state.
The native histogram is within 1% of custom and slightly ahead under sustained Orleans request load. This is the relevant comparison for the production migration.
The focused .NET 10 ShortRun microbenchmark shows the local tradeoff:
Native recording remains slower in isolation, while collection is about 14x faster and allocates about one fifth as much. Under the end-to-end request workload, the recording difference was not measurable. This is why the migration does not add grain type or local/remote tags; those dimensions can be evaluated separately.
OpenTelemetry 1.16 already contained the explicit-bound lookup improvements from open-telemetry/opentelemetry-dotnet#7165. Moving the repository to 1.18 also brings in open-telemetry/opentelemetry-dotnet#7601, which replaced an interlocked lock-release operation with
Volatile.Writeand reported 2-5.6% hot-path histogram improvements upstream.Bucket and percentile policy
A native histogram does not improve percentile approximation if boundaries are unchanged. Its benefits here are standard distribution semantics, fractional observations, MinMax, native temporality, and standard exporter handling.
Bucket selection belongs to the application or deployment because the useful thresholds depend on its SLOs, timeout, expected latency range, and backend capabilities. Submillisecond buckets are not enabled by Orleans. A 1 ms lower boundary is sufficient for the example profile, while 30-second and 60-second boundaries retain useful long-tail visibility. Deployments which need finer low-latency resolution can opt into it using their View.
The local benchmarks are directional rather than publication-grade. Dedicated Linux/CI hardware, multiple readers, delta-temporality load, precise retained memory per series, and remote exporter serialization remain unmeasured.
Microsoft Reviewers: Open in CodeFlow