The paper formulates anticipation as goal inference from a partially observed multimodal episode together with structured prediction of the remaining behavior, rather than exact motor forecasting. The architecture consists of a frozen neuro‑symbolic recognition encoder and a compact Hierarchical Planning Decoder (HPD). HPD jointly predicts at four ontological levels: the next actions, the remaining activities, low‑level intentions, and the episode‑level high‑level intention (HLI). Training employs soft neuro‑symbolic regularization that combines transition‑coherence and hierarchical‑continuity losses; inference uses hard reachability masks to enforce ontological validity. Experiments are conducted on a compositional four‑level benchmark of 15,002 multimodal episodes built from NTU RGB+D 120 features. HPD outperforms the strongest sequential baseline by +1.7 % Top‑5 at step 1 and +7.3 % at step 3; under compositional generalization the advantage widens to +4.9 % at step 1. Overall, 96.8 % of anticipated trajectories satisfy joint logic constraints, exceeding the baseline’s 88.1 % and the ground‑truth floor of 73.9 %. Soft logic terms alone reduce HLI‑reachability violations by 59.8 %–71.1 %, and hard masks eliminate them completely. Neural generation supplies predictive ranking, symbolic constraints ensure logical validity, and their combination yields coherent hierarchical anticipation while exposing remaining challenges in compositional goal generalization and unordered set prediction.
Review