Repository navigation
Conversation
_slice_scale_for_dimension maps a data slice onto the scale with floor/ceil division by the block size, but never resolved negative indices first. A negative `end` with |end| < block size ceil-divides to 0, so the scale is sliced to length zero. For per-row quantization the block spans the whole row, so any negative `end` on the last dim hits this: w[:, :-32] returns a Float8Tensor/Int8Tensor with a (N, 0) scale and no error, and it only fails later in dequantize (shape error for Int8, ZeroDivisionError for Float8). Resolve negative start/end against the data size before the block arithmetic. Covers Float8Tensor, Int8Tensor (scale and zero_point) and the prototype static float8 tensor, which all share this helper.
Kaif10
requested review from
andrewor14,
jerryzh168 and
vkuzo
as code owners
September 28, 2026 19:13
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/ao/4958
Note: Links to docs will display an error until the docs builds have been completed. This comment was automatically generated by Dr. CI and updates every 15 minutes. |
Author
|
Could a maintainer approve the CI workflows for this first-time-contributor PR? Suggested label: |
Author
|
@andrewor14 friendly ping. This is a 4-line fix to |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
_slice_scale_for_dimensionmaps a data-tensor slice onto the scale with floor/ceil division by the block size, but never resolves negative indices first. A negativeendwith|end| < block_sizeceil-divides to0, so the scale is sliced to length zero.For per-row quantization the block spans the whole row, so any negative
endon the last dim triggers it:Same with
Float8WeightOnlyConfig()/Float8DynamicActivationFloat8WeightConfig(granularity=PerRow()), which fails later withZeroDivisionError. The slice itself succeeds and returns a tensor with a plausible shape, so the error surfaces far from its cause.Negative
starthappens to give the right block because floor division lands on the same block when the dim is block-aligned, anddim=0is unaffected because a block size of 1 delegates toaten.slice, which handles negatives natively. Only the ceil-dividedendbreaks.Fix
Resolve negative
start/endagainstdata_shape[dim](clamped at 0, same as Python slice semantics) before the block arithmetic. SinceFloat8Tensor,Int8Tensor(scale andzero_point) and the prototype static float8 tensor all call this helper, they are all fixed.Tests
test_int8_tensor_cpu.py::test_slice_negative_indices: CPU. PerRow weight-only, PerGroup(32) weight-only and PerRow dynamic, over(None, -32),(-32, None),(-96, -32),(-1000, None)on both dims. Each negative slice is compared against the equivalent non-negative slice onqdata,scaleanddequantize(). The PerRow + negative-endcases fail on main and pass with this change; the rest pass on both and guard the unaffected paths.test_float8_tensor.py::test_slice_negative_indices: the same check forFloat8Tensorwith PerRow, next to the existingtest_slice(GPU-gated like the rest of that class). I validated its body on CPU by substitutingFloat8WeightOnlyConfig: it fails on main with a(64, 0)scale and passes with this change.ruff checkandruff formatare clean.test_int8_tensor_cpu.pygains 12 passing tests. Its existingcompile=Truecases fail identically before and after this change on my machine (no C++ compiler for Inductor on Windows), so they're unrelated.Suggested labels:
module: inference,topic: bug fix.