NeFut Logo NeFut
中 Admin Login

[CS.AI] PORL: Pretrained Offline Reinforcement Learning for the Job Shop Scheduling Problem

Published at: 2026-09-29 22:00 Last updated: 2026-09-30 01:41
#AI #Machine Learning #optimization

The Job Shop Scheduling Problem (JSSP) is a fundamental combinatorial optimization challenge in industry. This paper introduces PORL (Pretrained Offline Reinforcement Learning), a hybrid framework that first conducts online pretraining in a simulation environment and then fine‑tunes the policy offline using production‑specific data. Online interaction explores general scheduling strategies but suffers from the simulation‑to‑reality gap; offline reinforcement learning avoids direct interaction by learning from historical logs, yet its performance heavily depends on data quality and coverage. PORL bridges the two by applying a KL‑divergence‑based constraint during offline fine‑tuning to limit deviation from the pretrained policy. Experiments on JSSP instances with distribution shift use offline datasets generated by heuristic, noisy‑expert, and random behavioral policies. The results show that PORL consistently yields smaller optimality gaps than standalone offline RL and standard baselines, and its advantage grows as dataset quality deteriorates, indicating reduced sensitivity to offline data quality. These findings suggest that offline adaptation of pretrained policies is a promising solution for industrial scheduling environments where online exploration is impractical.

Review

Original Source: https://arxiv.org/abs/2609.30948

[h] Back to Home