Chart editing requires cross‑modal grounding: a visual change requested by the user must be realized directly in the drawing code, with all related updates applied and unrelated parts left untouched. Existing benchmarks either stress code executability or chart quality, but their metrics cannot clearly separate request fulfillment, missed coupled updates, and gratuitous modifications. ChartRevise addresses this gap by providing a structured dataset and a reference‑free evaluation protocol focused on exact program‑grounded chart editing.
The dataset is built on the Grammar of Graphics, systematically covering all chart‑editing operations. Source‑program checks verify that each operation is applicable across 20 chart types, three plotting libraries (e.g., Matplotlib, ggplot2, Plotly), and 344 edit types. The editing pipeline validates each requirement individually, triggering repair or exclusion when a requirement is unmet, thereby improving edit exactness. The final collection contains 92,438 records.
The evaluation protocol measures atomic requirement completion, identifies gratuitous changes, and detects missed coupled updates. These checks are combined with successful code execution and rendering to determine exact‑edit success. Fine‑tuning across five models and four external benchmarks yields a relative gain of 16% in mean requirement recall and 22% in mean exact‑edit rate.
Review