Skip to content

Fix Tensor correctness and accelerate contiguous operations - #134914

Merged
tannergooding merged 6 commits into
dotnet:mainfrom
tannergooding:tannergooding-tensor-correctness-audit
Sep 30, 2026
Merged

tannergooding merged 6 commits into
dotnet:mainfrom
tannergooding:tannergooding-tensor-correctness-audit

Conversation

@tannergooding

@tannergooding tannergooding commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Fix Tensor shape, storage, broadcasting, overlap, resize, and enumeration correctness; cover the behavior with regression tests and document intentional differences from NumPy. Empty shapes use zero automatically computed strides so unreachable products cannot overflow before the zero-length dimension.
  • Dispatch elementwise Tensor operations to TensorPrimitives for dense layouts and contiguous trailing slices, including compatible outer-dimension broadcasts. Preserve indexed traversal for incompatible layouts and ordering-sensitive reductions.
  • Accelerate strided flattening, equality, non-dense random fills, and non-dense Resize without changing logical iteration order.
  • Correct vector tangent range reduction discovered by CI: use guaranteed fused multiply-add on FMA hardware and an input-dependent scalar fallback for cancellation-sensitive vectors on other hardware, retaining vector evaluation for safe inputs.

Bug reports and fuzzing findings

Measurements

Local BenchmarkDotNet on Windows x64 (Ryzen 9 7950X, .NET 11). Representative results below compare the prior implementation to this branch except where noted:

Scenario Prior Current
Dense Add (64x16) 12.69 us 0.121 us
Gapped-row Add (64x16) 11.58 us 2.29 us
Broadcast-row Add (64x16) 12.58 us 1.91 us
Transposed Add (64x16) 12.54 us 11.60 us
Gapped-row FlattenTo 5.66 us 0.79 us
Gapped-row SequenceEqual 6.41 us 0.57 us
#121463 cropped FlattenTo repro (3x512x512 bytes) 2.030 ms 19.71 us

The #121463 row uses the issue's original 3x2000x2048 input cropped to 3x512x512, run locally against a saved pre-optimization assembly and this PR's Release assembly on the same .NET 11 host (approximately 103x faster). Additional local comparisons against equivalent indexed loops: non-dense Resize 1,909 -> 289 ns; uniform fill 5.56 -> 2.41 us; Gaussian fill 13.97 -> 10.88 us. Unary float Sin versus equivalent scalar iteration: dense 2.88 -> 0.63 us, gapped rows 8.41 -> 3.15 us. No additional managed allocation was measured in these samples. The transposed case is not addressed by the trailing-slice optimization.

Following review, sliced elementwise dispatch is gated centrally at 32 elements, or 16 for tensor/tensor binary operations. Against the previously published PR head, the threshold follow-up improved four-element gapped Add from 237 to 135 ns and Abs from 134 to 68 ns; sixteen-element Add measured 235 to 243 ns, and retained sliced paths had small differences in both directions. Normal-path pool returns are retained without exceptional-path finally cleanup. Removing the reinstated finally blocks changed permutation codegen: rank-three permutation measured about 45 to 59 ns in the focused comparison, with unchanged allocations. These follow-up changes are not claimed as universally performance-neutral.

The empty-stride correction was also compared locally with the preceding PR head across 12 dense-construction and reshape cases (ranks 1, 3, and 6, empty and nonempty). Allocation counts were unchanged; throughput measurements were noisy across repeated runs, so no performance improvement is claimed for this correctness fix.

CI exposed tangent range-reduction errors in the newly activated vector path. Without fused multiply-add, Tan(-32.986717f) returned about -148998.98 instead of the scalar/reference result -177349.88. The correction is in TensorPrimitives.Tan, preserving Tensor's vector dispatch and existing tangent forwarding tolerances. FMA hardware uses guaranteed fused reduction rather than MultiplyAddEstimate, whose fused behavior is not guaranteed across runtimes. Other hardware retains non-fused vector evaluation when abs(f) >= dn / 256; otherwise the affected vector uses the existing scalar fallback. The threshold scales with the reduction count and has a documented rounding-error bound.

With FMA unavailable, dense 4x1024 float Tan measured 4.31 -> 4.70 us on random inputs in [-1, 1] and 4.42 -> 8.08 us on [-50, 50], comparing the original vector implementation with the accuracy correction. Against a local global-FMA-gate implementation that scalarizes every non-FMA input, the corresponding results were 11.03 -> 4.70 us and 13.73 -> 8.08 us. Allocation counts remain unchanged. On FMA hardware, float and double kernels at all three widths retain the original instruction counts and code sizes; register allocation differs. FMA throughput measurements varied with code placement, so no precise performance-neutrality claim is made.

Validation

  • System.Numerics.Tensors.Tests: 6,100 net11.0 and 200 net481 tests passed in each of five local Windows x64 configurations: default intrinsics, FMA-enabled 256-bit vectors, FMA-enabled 128-bit vectors, non-FMA 128-bit vectors, and intrinsics disabled. ARM64/Mono and Linux CI have not yet validated the tangent correction.
  • System.Numerics.Tensors Release analyzer build succeeded with zero warnings or errors.
  • Direct float/double vector-operator probes covered 1,966,080 checks per FMA configuration across all three widths, small and large random inputs, broad exponents, and homogeneous cancellation-boundary vectors. No tolerance failures occurred with or without FMA; non-FMA 256/512-bit operators were exercised through their software vector implementations on this host.
  • Independent deterministic layout oracle: 28,049 checks passed across 1,500 cases covering sliced, permuted, gapped, broadcast, and empty layouts.

Resolves #128555
Resolves #133459
Resolves #134691
Resolves #121463

Note

This pull request description was generated by GitHub Copilot.

tannergooding and others added 2 commits September 29, 2026 19:28
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 4 pipeline(s).
12 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-numerics-tensors
See info in area-owners.md if you want to be subscribed.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

A rank-zero permutation regression, exception-unsafe pooled rentals, and an undocumented breaking overlap change remain.

Review effort: Balanced
Findings: 3 Medium severity · 1 Low severity

Open (4)
What changed in this PR

Improves System.Numerics.Tensors correctness and accelerates dense and contiguous-strided operations.

Changes:

  • Fixes broadcasting, resizing, overlap, reshaping, enumeration, pinning, and exception behavior.
  • Adds dense/contiguous fast paths backed by TensorPrimitives.
  • Expands regression tests and documents intentional NumPy differences.
File Description
tests/​TensorTests.cs Adds broad correctness and optimization coverage.
tests/​TensorSpanTests.cs Tests mutable spans, broadcasting, reshaping, and arrays.
tests/​TensorDimensionSpanTests.cs Tests dimension views and overflow handling.
tests/​ReadOnlyTensorSpanTests.cs Tests overlap, arrays, enumeration, and exceptions.
src/​System/​ThrowHelper.cs Adds validation and exception helpers.
src/​.../​TensorSpan_1.cs Fixes slicing, array validation, and enumeration.
src/​.../​TensorShape.cs Revises shape arithmetic, compatibility, and overlap logic.
src/​.../​TensorOperation.cs Adds contiguous fast paths and broadcast-aware dispatch.
src/​.../​TensorDimensionSpan_1.cs Fixes dimension-length and conversion logic.
src/​.../​Tensor.op_ExclusiveOr.cs Allocates broadcast-compatible results.
src/​.../​Tensor.op_BitwiseOr.cs Allocates broadcast-compatible results.
src/​.../​Tensor.op_BitwiseAnd.cs Allocates broadcast-compatible results.
src/​.../​Tensor.cs Implements the main correctness and performance changes.
src/​.../​Tensor_1.cs Fixes pinning and enumeration behavior.
src/​.../​ReadOnlyTensorSpan_1.cs Adds safe array validation and overlap handling.
src/​.../​ReadOnlyTensorDimensionSpan_1.cs Uses checked dimension products.
src/​.../​IReadOnlyTensor_1.cs Updates overlap contract documentation.
src/​Resources/​Strings.resx Adds new exception messages.
README.md Documents storage and NumPy differences.

@tannergooding tannergooding added the needs-breaking-change-doc-created Breaking changes need an issue opened with https://github.com/dotnet/docs/issues/new?template=dotnet label Sep 30, 2026
Reject nonexistent axes for rank-zero permutations and return temporary buffers on exceptional paths.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Valid zero-sized reshapes can still throw, and the new optimization and pooling paths retain performance and resource-management issues.

Review effort: Balanced
Findings: 1 Medium severity · 3 Low severity

Open (4)
Resolved since last review (3)

Keep normal-path buffer returns without exception-path finally blocks.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Use zero automatic strides for empty shapes so unreachable products cannot overflow before the zero-length dimension. Cover every zero position in inline and heap-backed shapes across construction and reshape.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Non-dense enumeration now performs rank-wide division-based offset calculation for every element, introducing an unvalidated asymptotic performance regression.

Review effort: Balanced
Findings: 1 Low severity

Open (1)
Resolved since last review (1)
Previously missed (3)

In code that hasn't changed since last review

Medium severity Read-only span enumeration incurs per-dimension divisions

src/​libraries/​System.Numerics.Tensors/​src/​System/​Numerics/​Tensors/​netcore/​ReadOnlyTensorSpan_1.cs:582

For non-dense read-only spans, every MoveNext now derives the offset with a division per dimension instead of incrementing a carried index, making enumeration O(element count × rank). This can substantially regress direct enumeration of high-rank strided tensors. Please preserve independent copied-enumerator state without full coordinate recomputation, or validate this tradeoff with representative benchmarks.

Medium severity Non-dense span enumeration regresses to O(rank) per element

src/​libraries/​System.Numerics.Tensors/​src/​System/​Numerics/​Tensors/​netcore/​TensorSpan_1.cs:503

For non-dense spans, each MoveNext now performs a full rank walk with DivRem operations. The prior carry-based traversal normally advanced in constant time, so iterating a high-rank strided span now costs O(rank) arithmetic per element. Please use copy-independent value-owned carry state or add measurements demonstrating that this regression is acceptable.

Medium severity MoveNext recomputes strided offsets per dimension

src/​libraries/​System.Numerics.Tensors/​src/​System/​Numerics/​Tensors/​netcore/​Tensor_1.cs:419

For non-dense tensors, MoveNext now recomputes the offset from scratch with one DivRem per dimension for every element. The previous carry-based traversal usually advanced the innermost dimension with one addition, so public enumeration changes from amortized O(1) work per item to O(rank) divisions. Please retain copy-independent value-owned state without recomputing every coordinate, or provide enumeration benchmarks that justify this regression for high-rank strided tensors.

Use guaranteed fused range reduction on FMA hardware. On other hardware, retain vector evaluation when the reduced remainder bounds cancellation error, and fall back to scalar tangent for affected vectors. Cover floating-point types, reduction boundaries, layouts, special values, and in-place operations.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@tannergooding
tannergooding force-pushed the tannergooding-tensor-correctness-audit branch from a17d3a4 to c431dcd Compare September 30, 2026 14:42
Comment thread src/libraries/System.Numerics.Tensors/README.md
@tannergooding

Copy link
Copy Markdown
Member Author

The breaking-change issue is open in dotnet/docs#56290, and the compatibility article and Tensor/NumPy guide are in dotnet/docs#56291. The docs PR is ready for review and intended to merge alongside this PR.

The breaking-change issue template also requests an email with the issue link to .NET Breaking Change Notifications.

Note

This comment was generated with GitHub Copilot.

@tannergooding

Copy link
Copy Markdown
Member Author

/ba-g build analysis leg is stuck, failure is unrelated

@tannergooding
tannergooding merged commit 137a670 into dotnet:main Sep 30, 2026
85 of 87 checks passed
@tannergooding
tannergooding deployed to copilot-pat-pool September 30, 2026 21:26 — with GitHub Actions Active
@tannergooding
tannergooding deployed to copilot-pat-pool September 30, 2026 21:26 — with GitHub Actions Active
@tannergooding
tannergooding deleted the tannergooding-tensor-correctness-audit branch September 30, 2026 21:26
@tannergooding
tannergooding deployed to copilot-pat-pool September 30, 2026 21:27 — with GitHub Actions Active
@tannergooding
tannergooding deployed to copilot-pat-pool September 30, 2026 21:29 — with GitHub Actions Active
@tannergooding
tannergooding deployed to copilot-pat-pool September 30, 2026 21:30 — with GitHub Actions Active
@tannergooding
tannergooding deployed to copilot-pat-pool September 30, 2026 21:31 — with GitHub Actions Active
@dotnet-milestone-bot dotnet-milestone-bot Bot added this to the 12.0-preview1 milestone Oct 1, 2026
@lilinus lilinus mentioned this pull request Oct 1, 2026
@jeffhandley

Copy link
Copy Markdown
Member

@tannergooding I would be supportive of proactively backporting this to .NET 11 GA. While there's the niche breaking change, I think it's a valuable change to make and I do think all of the other changes here meet our servicing bar pretty clearly for correctness.

/cc @artl93 heads-up that we will be bringing a .NET 11 GA backport request for this soon.

tannergooding added a commit that referenced this pull request Oct 6, 2026
#135281)

Fixes Issue <!-- Issue Number -->
[#133459](#133459),
[#134691](#134691),
[#128555](#128555),
and
[#121463](#121463).

main PR <!-- Link to PR if any that fixed this in the main branch. -->
[#134914](#134914)
and
[#135060](#135060).

# Description

<!-- Give a brief summary of the issue and how the pull request is
fixing it. -->

Backport both merged Tensor fixes to `release/11.0` in their original
order:

- `137a6701515dd7427e78bc46bcdd520c8d6e3a76`: fix array element-type
validation, slice pinning, shape/storage validation, broadcasting,
overlap handling, resize zero-fill, enumeration, and numerical
operations. Restore efficient dense and contiguous-slice dispatch,
including the tangent range-reduction accuracy correction needed by that
dispatch.
- `be7b579e4230520b5c47cd956513147e13f737a6`: make shape alignment
consistent across operations, treat default rank-zero empties as
effective shape `[0]` without changing their stored metadata, fix all
ten internal `Any` span kernels, and preserve native-width counts,
overlap-safe copy ordering, and valid empty-view origins.

Both cherry-picks applied without conflicts or backport-specific source
changes. The combined implementation, regression tests, README, and API
remarks match the merged changes; release-specific compatibility
suppressions are left unchanged. No public API is added.

# Customer Impact

<!-- What is the impact to customers of not taking this fix? -->

Without these fixes, incompatible `System.Array` storage can be
interpreted as the wrong element type and read beyond its bounds.
Pinning a tensor slice can pass the parent's data, rather than the
slice's data, to native consumers such as inference engines. Other
affected operations can produce incorrect results, mishandle empty or
broadcast shapes, omit the required resize zero-fill, or corrupt
overlapping output.

The existing cropped/gapped `FlattenTo` performance regression and
per-element overhead in dense and contiguous-slice operations also
remain. Native-backed spans with more than `int.MaxValue` logical
elements retain paths that prematurely narrow counts or offsets. Taking
both fixes together keeps the initial correctness/dispatch changes
paired with the subsequent shape-consistency and native-width
corrections.

# Regression

<!-- Is this fixing a problem that was introduced in the most recent
release, ie., fixing a regression? -->

Yes, in part, but these are not all newly introduced .NET 11
regressions. The array type-safety issue and cropped `FlattenTo`
regression were introduced by the .NET 10 Tensor rewrite in
[#114927](#114927)
and remain in `release/11.0`. Slice pinning was reproduced with the
10.0.9 package. The remaining changes address existing correctness and
consistency defects rather than one recent servicing regression.

# Testing

<!-- What kind of testing has been done with the fix. -->

Local Windows x64 validation on the actual `release/11.0` backport:

- `.\build.cmd clr+libs -rc release -lc release` on the unmodified
release baseline: passed, zero warnings/errors.
- `.\dotnet.cmd build
src\libraries\System.Numerics.Tensors\System.Numerics.Tensors.slnx -c
Release /p:RuntimeConfiguration=Release` after both cherry-picks:
passed, zero warnings/errors. This built the configured `net11.0`,
`net10.0`, `netstandard2.0`, and `net462` implementations and both test
targets.
- `.\dotnet.cmd build
src\libraries\System.Numerics.Tensors\tests\System.Numerics.Tensors.Tests.csproj
-t:Test -c Release /p:RuntimeConfiguration=Release`: passed in the
default configuration.
- The same full test command with `/p:testnobuild=true --no-restore` was
rerun under each environment override below: all passed. Fresh XML
results were checked for the full counts, with no failures, errors, or
skips.

| Configuration | Environment override | net11.0 passed | net481 passed
|
| --- | --- | ---: | ---: |
| Default intrinsics | None | 6,266 | 200 |
| 256-bit vectors | `DOTNET_PreferredVectorBitWidth=256` | 6,266 | 200 |
| 128-bit vectors | `DOTNET_PreferredVectorBitWidth=128` | 6,266 | 200 |
| AVX2/FMA disabled | `DOTNET_EnableAVX2=0` | 6,266 | 200 |
| All intrinsics disabled | `DOTNET_EnableHWIntrinsic=0` | 6,266 | 200 |

Coverage includes incompatible arrays, pinning, dense/strided/broadcast
layouts, rejection before writes, empty shapes/views, high ranks,
native-width logical counts, forced copy/reduction chunk boundaries,
NaNs/ties/signed zero, and protected-memory boundaries.

Local BenchmarkDotNet ShortRun comparisons used the same built Release
runtime with distinct baseline/backport Tensor assemblies on a Ryzen 9
7950X. Assembly loading and the active vector/FMA configuration were
verified. The 10 default-intrinsic jobs and four non-FMA tangent jobs
all produced measured results:

| Scenario | Release baseline | Combined backport |
| --- | ---: | ---: |
| Dense float Add, 64x16 | 16.37 us | 0.114 us |
| Gapped-row float Add, 64x16, strides `[32, 1]` | 16.25 us | 2.05 us |
| Gapped-row float FlattenTo, same layout | 7.92 us | 0.875 us |
| Float Tan, 4,096 values in `[-1, 1]`, FMA | 1.80 us | 1.72 us |
| Float Tan, 4,096 values in `[-50, 50]`, FMA | 1.63 us | 1.53 us |
| Float Tan, same small range, no FMA | 4.98 us | 5.08 us |
| Float Tan, same wide range, no FMA | 5.60 us | 9.19 us |

No managed allocation was measured in these samples. Small timing
differences are not claimed as precise gains or proof of performance
neutrality; the non-FMA wide-range slowdown is an accuracy tradeoff,
discussed below. These checks do not replace the more extensive
measurements in the main PRs.

Linux/macOS, Arm64, Mono, NativeAOT, and browser/WASM were not locally
validated.

# Risk

<!-- Please assess the risk of taking this fix. Provide details backing
up your assessment. -->

Low-to-moderate risk. The changes are consistent backports of two fixes
already merged on main, applied without conflicts or release-specific
source adaptations. The extensive regression tests passed in all five
local configurations, and the scope is confined to
`System.Numerics.Tensors`. Remaining risk is in the intentional behavior
changes and edge cases described below, rather than the patch size.

Intentional behavior changes include accepting/rejecting shape
combinations consistently, interpreting default empties as effective
shape `[0]`, rejecting incompatible array element types and unsupported
overlapping layouts before writes, and zero-filling newly exposed
resized elements. Consumers relying on the previous inconsistent
behavior can be affected. There is no compatibility switch.

On hardware without FMA, cancellation-sensitive tangent vectors now use
scalar evaluation to avoid inaccurate results. The wide-range local
sample increased from 5.60 to 9.19 us, approximately 64%; this is an
acknowledged throughput cost for correctness, not a performance-neutral
change. FMA-capable hardware retains fused vector range reduction.

Breaking-change documentation is already tracked by
[dotnet/docs#56290](dotnet/docs#56290) and
[dotnet/docs#56305](dotnet/docs#56305). Both
currently name .NET 12 Preview 1; their shipment-version metadata needs
to be aligned with the approved .NET 11 delivery if this backport is
accepted.

# Package authoring no longer needed in .NET 9

IMPORTANT: Starting with .NET 9, you no longer need to edit a NuGet
package's csproj to enable building and bump the version.
Keep in mind that we still need package authoring in .NET 8 and older
versions.

> [!NOTE]
> This pull request description was generated with assistance from
GitHub Copilot.

---------

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

This branch was successfully deployed

1 active deployment
copilot-pat-pool — c431dcd8 Deployed Sep 30, 2026 by tannergooding via conclusion #9411
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-System.Numerics.Tensors needs-breaking-change-doc-created Breaking changes need an issue opened with https://github.com/dotnet/docs/issues/new?template=dotnet

Projects

None yet

6 participants