Skip to content

[BugFix] Keep the first element of the discount series in _geom_series_like when gamma * lmbda is 0 or below the threshold - #4502

Draft
Nicholas022400701 wants to merge 3 commits into
pytorch:mainfrom
Nicholas022400701:fix/vec-gae-zero-lmbda
Draft

Nicholas022400701 wants to merge 3 commits into
pytorch:mainfrom
Nicholas022400701:fix/vec-gae-zero-lmbda

Conversation

@Nicholas022400701

Copy link
Copy Markdown
Contributor

Description

_geom_series_like (torchrl/objectives/value/functional.py) returned a 1D zeros_like(t) for r == 0.0, skipping the rs[0] = 1.0 and the unsqueeze(-1) below it, so _custom_conv1d got a 1D filter and raised IndexError: tuple index out of range; for 0 < r < thr it computed lim = int(log(thr) / log(r)) == 0, so t[:lim] was empty and rs[0] = 1.0 raised IndexError: index 0 is out of bounds for dimension 0 with size 0. r is gamma * lmbda in _fast_vec_gae and _fast_td_lambda_return_estimate and gamma in reward2go, so vec_generalized_advantage_estimate(gamma, 0.0, ...), GAE(lmbda=0.0, vectorized=True), TDLambdaEstimator(lmbda=0.0) (vectorized by default) and reward2go(gamma=0.0) all crashed, while the non-vectorized generalized_advantage_estimate returned the TD(0) advantage.

The series now always keeps its first element: r ** 0 is 1 whatever r, so lim = max(int(log(thr) / log(r)), 1), and r == 0 is lim = 1 (the series is [1, 0, 0, ...], a single-tap filter that makes _custom_conv1d the identity). The r >= 1.0 branch is unchanged.

Checked on main b8b6f0c (the patch applies unchanged to a4592ec) with 3 trajectories of 8 steps, float64, gamma 0.9, random done with a terminated subset: vectorized and non-vectorized GAE match exactly at lmbda=0 and to 3.3e-9 at lmbda=1e-9; vec_td_lambda_return_estimate matches td_lambda_return_estimate to 6e-8 (the same gap as at lmbda=0.5, from the float32 gamma_tensor there); reward2go(gamma=0) equals the reward. pytest test/objectives/test_values.py -k "test_gae or tdlambda or reward2go": 1833 passed. The new test_vec_estimates_zero_lmbda (lmbda 0.0 and 1e-9) fails on main with the two errors above and passes here.

Draft until a maintainer confirms the direction in #4489; I will mark it ready for review then.

Motivation and Context

Fixes #4489. lmbda=0 is the TD(0) advantage and gamma=0 reward-to-go is the reward itself; the vectorized paths should return what the non-vectorized ones return.

  • I have raised an issue to propose this change (required for new features and bug fixes)

Types of changes

  • Bug fix (non-breaking change which fixes an issue)

Checklist

  • I have read the CONTRIBUTION guide (required)
  • My change requires a change to the documentation.
  • I have updated the tests accordingly (required for a bug fix or a new feature).
  • I have updated the documentation accordingly.

AI disclosure: I used an AI coding agent to help write this patch, the tests and this description. I have read the change and the tests myself and I will answer review comments personally.

@pytorch-bot

pytorch-bot Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/rl/4502

Note: Links to docs will display an error until the docs builds have been completed.

⚠️ 16 Awaiting Approval

As of commit 59a0ffd with merge base a4592ec (image):

AWAITING APPROVAL - The following workflows need approval before CI can run:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 30, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. Objectives

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG] Vectorized GAE, TD(lambda) and reward2go raise IndexError when gamma * lmbda is 0 or below the 1e-7 threshold

1 participant