Grasping real-world physics demands more than object recognition, scene description, or short-term visual prediction, because actual systems involve multiple continuous fields, hidden causal mechanisms, partial observability, and dynamics that are highly sensitive to interventions.
We introduce FireWorldBench, a benchmark for multimodal large language models and agents that uses coupled-field fire dynamics as a canonical stress‑test environment. The combustion process naturally intertwines heat, concentration, and flow fields, allowing simultaneous assessment of state perception, temporal reasoning, causal explanation, and intervention planning.
The benchmark is organized along two complementary axes: a physical capability axis that focuses on understanding physical states, temporal evolution, and causal mechanisms; and a fire‑scenario task axis that designs tasks around various intervention contexts. Together they cover state inference, dynamic prediction, mechanism explanation, and intervention impact.
FireWorldBench comprises 520 fire‑world entries, including 494 controlled simulations and 26 real‑world‑aligned event groups, spanning 47 scene archetypes across 7 environment families. Each entry provides structured textual observations, multiple 2D physical‑field visualizations, and 3D event‑level scene modeling, yielding 9,074 interleaved text‑image Q&A pairs in both multiple‑choice and open‑ended report formats.
The benchmark evaluates whether models can infer latent physical states, explain underlying combustion mechanisms, forecast the evolution of coupled fields, and assess the consequences of various interventions, offering a demanding testbed for complex physical world intelligence.
Review: By embedding multi‑field coupling into the evaluation pipeline, FireWorldBench sets a new bar for physical reasoning in LLMs and agents, presenting a rich resource for advancing multimodal causal learning.