Repository navigation
[RFC] Introducing transform primatives for sim2real #4422
Description
Activity
A data point for ActuatorModel and ActionDelay defaults, from a public recording: in lerobot/svla_so101_pickplace (SO-101 arm, Feetech STS3215 servos, 30 Hz loop) the measured joint position trails the commanded target by 4 frames (133 ms) on the most-moving joint, with a steady mean offset of −0.98 units (action − state) on the gravity-loaded shoulder_lift, in dataset units (−100..100); reproducible from episode 0's action and observation.state columns (303 frames, best lag by cross-correlation over 0–14 frames). A one-parameter first-order servo fit gives k ≈ 0.25 per tick on every arm joint, which is the same as 133 ms.
On the design question: option A maps cleanly onto MuJoCo's mjModel arrays (body_mass, dof_damping, geom_friction, actuator_gainprm). Would physics_param_spec be a Composite so that sampled values are loggable per reset?
@sattyamjjain I think that physics param will be a composite in my view so far. for now, i am just looking at sampled values per reset as a property just of DomainRandomization. I don't see the point in having it for EnvBase reset right now
This might be related to the streaming TensorDict tutorial: https://docs.pytorch.org/tensordict/stable/tutorials/streamed_tensordict.html
There we use LazyStackedTensorDict to represent asynchronous streams sampled at different frequencies, bucketize them over time, and optionally densify them afterwards.
I think this could be relevant in particular for ObservationDelay. On real systems, delay often comes together with observations arriving asynchronously / at different frequencies, rather than being purely an N-env-steps FIFO. It may be useful to make sure the abstraction we choose for delayed observations can eventually accommodate that case too.
Potentially the buffering/stream representation and the transform deciding what constitutes the current observation could remain separate concerns.
Agree, and my data point has exactly that limitation, so worth saying it out loud.
In lerobot/svla_so101_pickplace the meta/info.json carries one timestamp and one frame_index shared by action, observation.state, and both camera streams, all at 30 fps. So the 4-frame lag I measured is a lag inside an already resampled uniform grid. That recording cannot show the async case at all; it was densified before it was stored. Most public LeRobot datasets are shaped the same way, which means a pure N-step FIFO will look correct against every one of them and still be wrong on hardware.
The split you are describing also matches what the number actually is. The 4 frames is a property of the servo bus, constant and estimable from a recording. Camera arrival jitter is a property of the stream. A single global N conflates the two, and then neither can be
calibrated. If ObservationDelay took per-key delays instead of one N, the measured value drops straight in as observation.state: 4 at 30 Hz, and the image keys keep their own./assign
Reacted by github-actions
##Summary:
ObservationDelay,ActionDelay,ObservationNoise,ActuatorModel,SymmetryAugmentation,DomainRandomization.DomainRandomization##RFC decision for domainrandomization interaction with sim