Behavioral cloning trains a policy on offline expert demonstrations, yet deployment is closed‑loop: each action influences the next observations the policy receives. We investigate how much closed‑loop driving competence a compact multimodal policy can acquire from offline demonstrations in the CARLA simulator. The policy consumes five‑frame histories of RGB images, LiDAR, vehicle telemetry and lane waypoints, and predicts throttle, brake and steering at 20 Hz. Demonstrations were gathered in three stages, ending with a systematic route‑generation pipeline that enumerates spawn points, feasible maneuvers and verifies completed autopilot routes. The released model has 1.36 million parameters and was trained on 236 882 windows, representing roughly 3.3 hours of driving from 448 captures. After training, the policy drives autonomously for hours on both training and held‑out routes without collisions and qualitatively transfers to an unseen CARLA town with different road geometry. Recovery from large trajectory deviations was observed, though systematic recovery was not evaluated. We report offline metrics and separate them from qualitative closed‑loop observations. Code, trained checkpoint, ONNX model, data sample and an evidence audit are publicly released.
Review