This work proposes using category theory as a practical scaffold for structuring reinforcement learning in high‑dimensional partially observable environments. By partitioning the state space into equivalence classes induced by symmetry orbits and equipping each class with a groupoid centred on a canonical representative, the agent can share knowledge across many similar states instead of treating every orientation or position as a brand‑new problem. Consequently learning proceeds on a symmetry‑reduced state space where each orbit appears only once, preserving environmental structure while eliminating redundancy and boosting sample efficiency.
We embed the approach into standard reinforcement‑learning pipelines and evaluate it on two partially observable benchmarks, comparing orbit‑based partitioning with conventional baselines. Results show consistent performance gains in environments that exhibit latent symmetry. More fundamentally, the categorical construction offers a principled bridge between abstract reinforcement‑learning formulations and concrete computational implementations, pointing toward more structured and scalable learning systems.
Review