PolyBridgeBench is an executable benchmark that measures how well multimodal large language models (MLLMs) can produce fully load‑bearing bridge structures. A model receives a visual scene of a bridge site together with engineering constraints such as span and material budget, then generates a complete node‑member‑material topology. The system first runs deterministic legality checks to verify geometric and material rules; only legal designs are passed to a native dynamic physics simulator for functional validation. When a simulation fails, the benchmark returns temporal visual evidence of the rollout and asks the model to repair the design within a fixed interaction budget. By separately reporting deterministic validity, dynamic functional success, and post‑failure recovery, the benchmark pinpoints the stage at which a design breaks down. Experiments on 189 levels with six representative MLLMs reveal a large gap between deterministic validity and dynamic success, strong sensitivity to material budgets, and limited recovery ability under the primary strict‑budget setting.
Review