NeFut Logo NeFut
中 Admin Login

[CS.AI] Learning Multiple Timescales for Goal-Conditioned Reinforcement Learning

Published at: 2026-10-02 22:00 Last updated: 2026-10-06 12:11
#algorithm #Machine Learning #optimization

Existing offline goal‑conditioned reinforcement learning (GCRL) methods struggle with long‑horizon tasks. Discounting shrinks value differences between distant states until they fall below the function‑approximation error, leaving the agent without a ranking signal. Temporal abstraction treats $k$ environment steps as a single transition, restoring long‑range value differences, but a single fixed $k$ cannot suit all state‑goal distances: large $k$ preserves distant value gaps while collapsing distinctions between nearby states, and small $k$ does the opposite. To make this trade‑off explicit, we introduce Generalized Implicit Temporal Abstraction (GITA), which conditions a single value function on $k$. GITA trains one policy by aggregating advantage‑weighted supervision across multiple $k$ values, giving larger positive advantages to state‑goal pairs a stronger influence on the update. Thus GITA does not need to choose between local resolution and long‑range signal; it retains both without committing to a single $k$. Experiments on the OGBench benchmark show that GITA outperforms a broad range of offline GCRL baselines, raising the average success rate across all tasks by 25 percentage points (73% relative improvement) over HIQL and improving over the strongest fixed‑$k$ method OTA by 7 percentage points (14% relative).

Review: GITA’s explicit conditioning on $k$ provides a unified multi‑scale value estimation framework, markedly boosting offline learning performance on long‑horizon goal‑directed tasks and offering a fresh perspective on leveraging temporal abstraction in reinforcement learning.

Original Source: https://arxiv.org/abs/2610.00849

[h] Back to Home