Scientific figures are meant to convey information visually, yet most multimodal large language models (MLLMs) merely translate the visual content back into text, forcing readers to manually map the explanation onto the figure. To address this gap, we introduce FigAct, a framework that acts directly on existing graphical elements to turn static scientific figures into question‑conditioned visual presentations. FigAct mimics a human presenter:
- Generates a sequence of short narrations;
- Grounds each narration in the corresponding visual evidence;
- Applies visual actions (e.g., highlight, zoom, overlay) to guide the viewer’s attention.
For efficient element localization, FigAct employs a hierarchical search strategy that cuts token usage by roughly 40\times. The model, FigAct‑8B, is trained with three task‑specific rewards:
- Grounding accuracy (ensuring narration matches visual evidence);
- Search efficiency (encouraging fast localization);
- Rendering quality (producing natural, readable visual actions).
We also construct a human‑verified benchmark from figures in real scientific papers to evaluate MLLM’s ability to generate grounded visual explanations. Experiments demonstrate that FigAct markedly improves the clarity and traceability of explanations, confirming the benefit of treating scientific figures as presentation canvases.
Review: FigAct’s “speak‑see‑point” loop tightly couples textual explanation with interactive visual cues, offering a more intuitive and followable way to understand scientific graphics.