You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Fix Tensor correctness and accelerate contiguous operations - #134914
Fix Tensor shape, storage, broadcasting, overlap, resize, and enumeration correctness; cover the behavior with regression tests and document intentional differences from NumPy. Empty shapes use zero automatically computed strides so unreachable products cannot overflow before the zero-length dimension.
Dispatch elementwise Tensor operations to TensorPrimitives for dense layouts and contiguous trailing slices, including compatible outer-dimension broadcasts. Preserve indexed traversal for incompatible layouts and ordering-sensitive reductions.
Accelerate strided flattening, equality, non-dense random fills, and non-dense Resize without changing logical iteration order.
Correct vector tangent range reduction discovered by CI: use guaranteed fused multiply-add on FMA hardware and an input-dependent scalar fallback for cancellation-sensitive vectors on other hardware, retaining vector evaluation for safe inputs.
ResizeTo fills unwritten logical destination elements with default(T) and rejects ambiguous overlapping destination layouts; the behavior is documented (Questions / issues with Tensor.ResizeTo #128555).
This PR addresses the Tensor findings 3, 6, 7, 8, 9, and 13 in the community SharpFuzz report: strided ToString, broadcast pairwise comparisons, empty *Any comparisons, preserving the element when squeezing all singleton dimensions, ResizeTo zero-fill, and constructor/slice exception contracts. Squeezing retains a rank-one [1] shape rather than NumPy's rank-zero shape; this intentional difference is documented in the library README. Findings 4, 5, 10, 11, and 12 concern TensorPrimitives and are not addressed here.
Local BenchmarkDotNet on Windows x64 (Ryzen 9 7950X, .NET 11). Representative results below compare the prior implementation to this branch except where noted:
The #121463 row uses the issue's original 3x2000x2048 input cropped to 3x512x512, run locally against a saved pre-optimization assembly and this PR's Release assembly on the same .NET 11 host (approximately 103x faster). Additional local comparisons against equivalent indexed loops: non-dense Resize 1,909 -> 289 ns; uniform fill 5.56 -> 2.41 us; Gaussian fill 13.97 -> 10.88 us. Unary float Sin versus equivalent scalar iteration: dense 2.88 -> 0.63 us, gapped rows 8.41 -> 3.15 us. No additional managed allocation was measured in these samples. The transposed case is not addressed by the trailing-slice optimization.
Following review, sliced elementwise dispatch is gated centrally at 32 elements, or 16 for tensor/tensor binary operations. Against the previously published PR head, the threshold follow-up improved four-element gapped Add from 237 to 135 ns and Abs from 134 to 68 ns; sixteen-element Add measured 235 to 243 ns, and retained sliced paths had small differences in both directions. Normal-path pool returns are retained without exceptional-path finally cleanup. Removing the reinstated finally blocks changed permutation codegen: rank-three permutation measured about 45 to 59 ns in the focused comparison, with unchanged allocations. These follow-up changes are not claimed as universally performance-neutral.
The empty-stride correction was also compared locally with the preceding PR head across 12 dense-construction and reshape cases (ranks 1, 3, and 6, empty and nonempty). Allocation counts were unchanged; throughput measurements were noisy across repeated runs, so no performance improvement is claimed for this correctness fix.
CI exposed tangent range-reduction errors in the newly activated vector path. Without fused multiply-add, Tan(-32.986717f) returned about -148998.98 instead of the scalar/reference result -177349.88. The correction is in TensorPrimitives.Tan, preserving Tensor's vector dispatch and existing tangent forwarding tolerances. FMA hardware uses guaranteed fused reduction rather than MultiplyAddEstimate, whose fused behavior is not guaranteed across runtimes. Other hardware retains non-fused vector evaluation when abs(f) >= dn / 256; otherwise the affected vector uses the existing scalar fallback. The threshold scales with the reduction count and has a documented rounding-error bound.
With FMA unavailable, dense 4x1024 float Tan measured 4.31 -> 4.70 us on random inputs in [-1, 1] and 4.42 -> 8.08 us on [-50, 50], comparing the original vector implementation with the accuracy correction. Against a local global-FMA-gate implementation that scalarizes every non-FMA input, the corresponding results were 11.03 -> 4.70 us and 13.73 -> 8.08 us. Allocation counts remain unchanged. On FMA hardware, float and double kernels at all three widths retain the original instruction counts and code sizes; register allocation differs. FMA throughput measurements varied with code placement, so no precise performance-neutrality claim is made.
Validation
System.Numerics.Tensors.Tests: 6,100 net11.0 and 200 net481 tests passed in each of five local Windows x64 configurations: default intrinsics, FMA-enabled 256-bit vectors, FMA-enabled 128-bit vectors, non-FMA 128-bit vectors, and intrinsics disabled. ARM64/Mono and Linux CI have not yet validated the tangent correction.
System.Numerics.Tensors Release analyzer build succeeded with zero warnings or errors.
Direct float/double vector-operator probes covered 1,966,080 checks per FMA configuration across all three widths, small and large random inputs, broad exponents, and homogeneous cancellation-boundary vectors. No tolerance failures occurred with or without FMA; non-FMA 256/512-bit operators were exercised through their software vector implementations on this host.
Independent deterministic layout oracle: 28,049 checks passed across 1,500 cases covering sliced, permuted, gapped, broadcast, and empty layouts.
Azure Pipelines:
Successfully started running 4 pipeline(s).
12 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.
Use zero automatic strides for empty shapes so unreachable products cannot overflow before the zero-length dimension. Cover every zero position in inline and heap-backed shapes across construction and reshape.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
The reason will be displayed to describe this comment to others. Learn more.
Copilot review overview
🔵 Needs a closer look
Non-dense enumeration now performs rank-wide division-based offset calculation for every element, introducing an unvalidated asymptotic performance regression.
For non-dense read-only spans, every MoveNext now derives the offset with a division per dimension instead of incrementing a carried index, making enumeration O(element count × rank). This can substantially regress direct enumeration of high-rank strided tensors. Please preserve independent copied-enumerator state without full coordinate recomputation, or validate this tradeoff with representative benchmarks.
Non-dense span enumeration regresses to O(rank) per element
For non-dense spans, each MoveNext now performs a full rank walk with DivRem operations. The prior carry-based traversal normally advanced in constant time, so iterating a high-rank strided span now costs O(rank) arithmetic per element. Please use copy-independent value-owned carry state or add measurements demonstrating that this regression is acceptable.
For non-dense tensors, MoveNext now recomputes the offset from scratch with one DivRem per dimension for every element. The previous carry-based traversal usually advanced the innermost dimension with one addition, so public enumeration changes from amortized O(1) work per item to O(rank) divisions. Please retain copy-independent value-owned state without recomputing every coordinate, or provide enumeration benchmarks that justify this regression for high-rank strided tensors.
Use guaranteed fused range reduction on FMA hardware. On other hardware, retain vector evaluation when the reduced remainder bounds cancellation error, and fall back to scalar tangent for affected vectors. Cover floating-point types, reduction boundaries, layouts, special values, and in-place operations.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
The breaking-change issue is open in dotnet/docs#56290, and the compatibility article and Tensor/NumPy guide are in dotnet/docs#56291. The docs PR is ready for review and intended to merge alongside this PR.
@tannergooding I would be supportive of proactively backporting this to .NET 11 GA. While there's the niche breaking change, I think it's a valuable change to make and I do think all of the other changes here meet our servicing bar pretty clearly for correctness.
/cc @artl93 heads-up that we will be bringing a .NET 11 GA backport request for this soon.
#135281)
Fixes Issue <!-- Issue Number -->
[#133459](#133459),
[#134691](#134691),
[#128555](#128555),
and
[#121463](#121463).
main PR <!-- Link to PR if any that fixed this in the main branch. -->
[#134914](#134914)
and
[#135060](#135060).
# Description
<!-- Give a brief summary of the issue and how the pull request is
fixing it. -->
Backport both merged Tensor fixes to `release/11.0` in their original
order:
- `137a6701515dd7427e78bc46bcdd520c8d6e3a76`: fix array element-type
validation, slice pinning, shape/storage validation, broadcasting,
overlap handling, resize zero-fill, enumeration, and numerical
operations. Restore efficient dense and contiguous-slice dispatch,
including the tangent range-reduction accuracy correction needed by that
dispatch.
- `be7b579e4230520b5c47cd956513147e13f737a6`: make shape alignment
consistent across operations, treat default rank-zero empties as
effective shape `[0]` without changing their stored metadata, fix all
ten internal `Any` span kernels, and preserve native-width counts,
overlap-safe copy ordering, and valid empty-view origins.
Both cherry-picks applied without conflicts or backport-specific source
changes. The combined implementation, regression tests, README, and API
remarks match the merged changes; release-specific compatibility
suppressions are left unchanged. No public API is added.
# Customer Impact
<!-- What is the impact to customers of not taking this fix? -->
Without these fixes, incompatible `System.Array` storage can be
interpreted as the wrong element type and read beyond its bounds.
Pinning a tensor slice can pass the parent's data, rather than the
slice's data, to native consumers such as inference engines. Other
affected operations can produce incorrect results, mishandle empty or
broadcast shapes, omit the required resize zero-fill, or corrupt
overlapping output.
The existing cropped/gapped `FlattenTo` performance regression and
per-element overhead in dense and contiguous-slice operations also
remain. Native-backed spans with more than `int.MaxValue` logical
elements retain paths that prematurely narrow counts or offsets. Taking
both fixes together keeps the initial correctness/dispatch changes
paired with the subsequent shape-consistency and native-width
corrections.
# Regression
<!-- Is this fixing a problem that was introduced in the most recent
release, ie., fixing a regression? -->
Yes, in part, but these are not all newly introduced .NET 11
regressions. The array type-safety issue and cropped `FlattenTo`
regression were introduced by the .NET 10 Tensor rewrite in
[#114927](#114927)
and remain in `release/11.0`. Slice pinning was reproduced with the
10.0.9 package. The remaining changes address existing correctness and
consistency defects rather than one recent servicing regression.
# Testing
<!-- What kind of testing has been done with the fix. -->
Local Windows x64 validation on the actual `release/11.0` backport:
- `.\build.cmd clr+libs -rc release -lc release` on the unmodified
release baseline: passed, zero warnings/errors.
- `.\dotnet.cmd build
src\libraries\System.Numerics.Tensors\System.Numerics.Tensors.slnx -c
Release /p:RuntimeConfiguration=Release` after both cherry-picks:
passed, zero warnings/errors. This built the configured `net11.0`,
`net10.0`, `netstandard2.0`, and `net462` implementations and both test
targets.
- `.\dotnet.cmd build
src\libraries\System.Numerics.Tensors\tests\System.Numerics.Tensors.Tests.csproj
-t:Test -c Release /p:RuntimeConfiguration=Release`: passed in the
default configuration.
- The same full test command with `/p:testnobuild=true --no-restore` was
rerun under each environment override below: all passed. Fresh XML
results were checked for the full counts, with no failures, errors, or
skips.
| Configuration | Environment override | net11.0 passed | net481 passed
|
| --- | --- | ---: | ---: |
| Default intrinsics | None | 6,266 | 200 |
| 256-bit vectors | `DOTNET_PreferredVectorBitWidth=256` | 6,266 | 200 |
| 128-bit vectors | `DOTNET_PreferredVectorBitWidth=128` | 6,266 | 200 |
| AVX2/FMA disabled | `DOTNET_EnableAVX2=0` | 6,266 | 200 |
| All intrinsics disabled | `DOTNET_EnableHWIntrinsic=0` | 6,266 | 200 |
Coverage includes incompatible arrays, pinning, dense/strided/broadcast
layouts, rejection before writes, empty shapes/views, high ranks,
native-width logical counts, forced copy/reduction chunk boundaries,
NaNs/ties/signed zero, and protected-memory boundaries.
Local BenchmarkDotNet ShortRun comparisons used the same built Release
runtime with distinct baseline/backport Tensor assemblies on a Ryzen 9
7950X. Assembly loading and the active vector/FMA configuration were
verified. The 10 default-intrinsic jobs and four non-FMA tangent jobs
all produced measured results:
| Scenario | Release baseline | Combined backport |
| --- | ---: | ---: |
| Dense float Add, 64x16 | 16.37 us | 0.114 us |
| Gapped-row float Add, 64x16, strides `[32, 1]` | 16.25 us | 2.05 us |
| Gapped-row float FlattenTo, same layout | 7.92 us | 0.875 us |
| Float Tan, 4,096 values in `[-1, 1]`, FMA | 1.80 us | 1.72 us |
| Float Tan, 4,096 values in `[-50, 50]`, FMA | 1.63 us | 1.53 us |
| Float Tan, same small range, no FMA | 4.98 us | 5.08 us |
| Float Tan, same wide range, no FMA | 5.60 us | 9.19 us |
No managed allocation was measured in these samples. Small timing
differences are not claimed as precise gains or proof of performance
neutrality; the non-FMA wide-range slowdown is an accuracy tradeoff,
discussed below. These checks do not replace the more extensive
measurements in the main PRs.
Linux/macOS, Arm64, Mono, NativeAOT, and browser/WASM were not locally
validated.
# Risk
<!-- Please assess the risk of taking this fix. Provide details backing
up your assessment. -->
Low-to-moderate risk. The changes are consistent backports of two fixes
already merged on main, applied without conflicts or release-specific
source adaptations. The extensive regression tests passed in all five
local configurations, and the scope is confined to
`System.Numerics.Tensors`. Remaining risk is in the intentional behavior
changes and edge cases described below, rather than the patch size.
Intentional behavior changes include accepting/rejecting shape
combinations consistently, interpreting default empties as effective
shape `[0]`, rejecting incompatible array element types and unsupported
overlapping layouts before writes, and zero-filling newly exposed
resized elements. Consumers relying on the previous inconsistent
behavior can be affected. There is no compatibility switch.
On hardware without FMA, cancellation-sensitive tangent vectors now use
scalar evaluation to avoid inaccurate results. The wide-range local
sample increased from 5.60 to 9.19 us, approximately 64%; this is an
acknowledged throughput cost for correctness, not a performance-neutral
change. FMA-capable hardware retains fused vector range reduction.
Breaking-change documentation is already tracked by
[dotnet/docs#56290](dotnet/docs#56290) and
[dotnet/docs#56305](dotnet/docs#56305). Both
currently name .NET 12 Preview 1; their shipment-version metadata needs
to be aligned with the approved .NET 11 delivery if this backport is
accepted.
# Package authoring no longer needed in .NET 9
IMPORTANT: Starting with .NET 9, you no longer need to edit a NuGet
package's csproj to enable building and bump the version.
Keep in mind that we still need package authoring in .NET 8 and older
versions.
> [!NOTE]
> This pull request description was generated with assistance from
GitHub Copilot.
---------
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Bug reports and fuzzing findings
ReadOnlyTensorSpan<T>now rejects incompatibleSystem.Arrayelement types (ReadOnlyTensorSpan<T>can read outside bounds of array #133459), andTensor<T>slice pinning starts at the slice offset (Tensor<T> slice pinning starts at the parent array instead of the slice #134691).ResizeTofills unwritten logical destination elements withdefault(T)and rejects ambiguous overlapping destination layouts; the behavior is documented (Questions / issues withTensor.ResizeTo#128555).TensorSpan<T>.FlattenTonow copies contiguous rows rather than individual elements, addressing the reported regression (TensorSpan<T>.FlattenTois 80x slower with the .NET 10 package #121463); the exact issue repro is measured below.Tensorfindings 3, 6, 7, 8, 9, and 13 in the community SharpFuzz report: stridedToString, broadcast pairwise comparisons, empty*Anycomparisons, preserving the element when squeezing all singleton dimensions,ResizeTozero-fill, and constructor/slice exception contracts. Squeezing retains a rank-one[1]shape rather than NumPy's rank-zero shape; this intentional difference is documented in the library README. Findings 4, 5, 10, 11, and 12 concernTensorPrimitivesand are not addressed here.Measurements
Local BenchmarkDotNet on Windows x64 (Ryzen 9 7950X, .NET 11). Representative results below compare the prior implementation to this branch except where noted:
The #121463 row uses the issue's original 3x2000x2048 input cropped to 3x512x512, run locally against a saved pre-optimization assembly and this PR's Release assembly on the same .NET 11 host (approximately 103x faster). Additional local comparisons against equivalent indexed loops: non-dense Resize 1,909 -> 289 ns; uniform fill 5.56 -> 2.41 us; Gaussian fill 13.97 -> 10.88 us. Unary float Sin versus equivalent scalar iteration: dense 2.88 -> 0.63 us, gapped rows 8.41 -> 3.15 us. No additional managed allocation was measured in these samples. The transposed case is not addressed by the trailing-slice optimization.
Following review, sliced elementwise dispatch is gated centrally at 32 elements, or 16 for tensor/tensor binary operations. Against the previously published PR head, the threshold follow-up improved four-element gapped Add from 237 to 135 ns and Abs from 134 to 68 ns; sixteen-element Add measured 235 to 243 ns, and retained sliced paths had small differences in both directions. Normal-path pool returns are retained without exceptional-path
finallycleanup. Removing the reinstatedfinallyblocks changed permutation codegen: rank-three permutation measured about 45 to 59 ns in the focused comparison, with unchanged allocations. These follow-up changes are not claimed as universally performance-neutral.The empty-stride correction was also compared locally with the preceding PR head across 12 dense-construction and reshape cases (ranks 1, 3, and 6, empty and nonempty). Allocation counts were unchanged; throughput measurements were noisy across repeated runs, so no performance improvement is claimed for this correctness fix.
CI exposed tangent range-reduction errors in the newly activated vector path. Without fused multiply-add,
Tan(-32.986717f)returned about -148998.98 instead of the scalar/reference result -177349.88. The correction is inTensorPrimitives.Tan, preserving Tensor's vector dispatch and existing tangent forwarding tolerances. FMA hardware uses guaranteed fused reduction rather thanMultiplyAddEstimate, whose fused behavior is not guaranteed across runtimes. Other hardware retains non-fused vector evaluation whenabs(f) >= dn / 256; otherwise the affected vector uses the existing scalar fallback. The threshold scales with the reduction count and has a documented rounding-error bound.With FMA unavailable, dense 4x1024 float Tan measured 4.31 -> 4.70 us on random inputs in [-1, 1] and 4.42 -> 8.08 us on [-50, 50], comparing the original vector implementation with the accuracy correction. Against a local global-FMA-gate implementation that scalarizes every non-FMA input, the corresponding results were 11.03 -> 4.70 us and 13.73 -> 8.08 us. Allocation counts remain unchanged. On FMA hardware, float and double kernels at all three widths retain the original instruction counts and code sizes; register allocation differs. FMA throughput measurements varied with code placement, so no precise performance-neutrality claim is made.
Validation
Resolves #128555
Resolves #133459
Resolves #134691
Resolves #121463
Note
This pull request description was generated by GitHub Copilot.