Programmable Logic Controllers (PLCs) coordinate industrial equipment by reading sensor inputs and issuing control commands. To assess whether large language model (LLM) generated PLC programs meet task requirements and safety constraints, one must observe how their commands affect device and workpiece states. We therefore introduce PLCWorld, a closed‑loop execution environment and benchmark that couples Structured Text (ST) execution with simulated plant responses and sensor feedback.
Grounded in control relations extracted from industrial PLC code and engineering documentation, PLCWorld comprises 100 synthetic tasks and 473 task‑condition pairs spanning Motion Control and Material Handling. Task difficulty is defined by the scope of control dependencies. A unified protocol reports Task Success and Safety Violation separately.
Validation combines practitioner review, comparison of reference and alternative programs, targeted counterexample testing, specification‑evaluator alignment checks, and cross‑validation with independent ST runtimes. Both reference and alternative programs satisfy their applicable cases, while all 542 targeted counterexamples trigger their designated evaluator rules under at least one registered condition.
The Execution Gap metric relates acceptance of a submission to subsequent task failure or observed safety violation. Results show that direct GPT‑5.5 achieves 82.70% task success on Easy cases but only 25.10% on Hard cases. Evaluations of six LLMs and four adapted generation‑and‑verification workflows further expose trade‑offs among completion, safety, and generation cost.
The code, simulation environment, benchmark tasks, and baseline implementations are publicly available at https://yunji0516.github.io/PLCWorld/.
Review