This paper introduces a data‑driven self‑learning control approach for highly flexible modular manufacturing systems. The key idea is a model‑based reinforcement learning framework that incorporates approximate inverse process models during policy training. The inverse models decouple actuation dynamics from state‑space dynamics, allowing reinforcement learning to operate solely in the task space. We propose a lightweight feedforward network to realize the approximate inverse models and embed them into the policy network of standard RL algorithms. The approach is evaluated on a laboratory modular production testbed with heterogeneous modules. Results show notable gains in both performance and training speed, particularly for off‑policy algorithms.
Review: The integration of inverse models simplifies the learning problem and enhances adaptability and efficiency of modular production control.