A goal that is close in space can be far away in time. Obstacles, terrain, and the agent's capabilities determine the actual time needed to reach the goal. Existing contrastive and survival reinforcement‑learning critics do not measure distances in representation space using temporal units. ChronoSRL addresses this by giving the critic's embeddings an explicit temporal geometry. The distance between state‑action and goal embeddings is trained to match the observed goal‑reaching time, while unreachable goals and goals from other trajectories are pushed beyond at least one discount horizon.
A single fast reach does not guarantee reliable reachability, so the policy should not follow the temporal distance directly. Building on survival reinforcement learning, we predict from the temporal embeddings not only the full distribution of goal‑reaching times but also the time spent near the goal. Consequently, the policy is encouraged to reach the goal sooner, more reliably, and to stay close to it.
Experiments on seven standard locomotion and navigation benchmarks show that ChronoSRL learns faster and achieves higher performance than contrastive, action‑chunked contrastive, and survival RL baselines, even with much smaller networks. To probe the limits of self‑supervised RL, we add velocity tracking, goal‑position reaching, and box‑climbing tasks for a quadruped robot in a realistic sim‑to‑real setup, demonstrating that typical robotics shaping terms can be naturally incorporated. ChronoSRL is the only tested method that stays at commanded velocities, reaches precise goal positions, and climbs the highest boxes.
Review: By embedding temporal metrics into the representation space, ChronoSRL resolves the mismatch between spatial proximity and actual time, and its prediction of time distributions improves policy reliability and robustness, highlighting the promise of self‑supervised reinforcement learning for complex robotic tasks.