NeFut Logo NeFut
中 Admin Login

[CS.AI] FigAct: Turning Scientific Figures into Active Canvases for Explanation

Published at: 2026-09-30 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

Scientific figures are meant to convey information visually, yet most multimodal large language models (MLLMs) merely translate the visual content back into text, forcing readers to manually map the explanation onto the figure. To address this gap, we introduce FigAct, a framework that acts directly on existing graphical elements to turn static scientific figures into question‑conditioned visual presentations. FigAct mimics a human presenter:

  1. Generates a sequence of short narrations;
  2. Grounds each narration in the corresponding visual evidence;
  3. Applies visual actions (e.g., highlight, zoom, overlay) to guide the viewer’s attention.

For efficient element localization, FigAct employs a hierarchical search strategy that cuts token usage by roughly 40\times. The model, FigAct‑8B, is trained with three task‑specific rewards:

We also construct a human‑verified benchmark from figures in real scientific papers to evaluate MLLM’s ability to generate grounded visual explanations. Experiments demonstrate that FigAct markedly improves the clarity and traceability of explanations, confirming the benefit of treating scientific figures as presentation canvases.

Review: FigAct’s “speak‑see‑point” loop tightly couples textual explanation with interactive visual cues, offering a more intuitive and followable way to understand scientific graphics.

Original Source: https://arxiv.org/abs/2609.36190

[h] Back to Home