The Abstraction and Reasoning Corpus (ARC) has become a leading benchmark for measuring general abstract reasoning and fluid intelligence in AI models. Conventional ARC evaluation, however, only checks whether a model produces the correct output grid for a given test input, overlooking the range of abilities that true abstract skill acquisition should entail. PotARCin extends ARC by introducing five evaluation dimensions: Definition, Classification, Constrained Generation, Editing, and Inversion.
PotARCin uses programmatic techniques to synthesize new task instances and transform inputs for any ARC task, enabling dynamic generative sampling beyond fixed input‑output pairs. We evaluated five state‑of‑the‑art models on the ARC‑AGI‑1 training set and observed a 25%–52% performance gap between standard ARC accuracy and PotARCin’s multi‑dimensional scores, with the latter reshuffling models that appear tied under single‑metric ranking.
Additional studies examined the effect of generative sampling, the difficulty of various corruption types, and model self‑consistency. Models often contradicted their own formalized rule even after stating it correctly. We also introduced P‑ARC, a held‑out hand‑crafted test set, where all models achieved only 1%–8% accuracy across the five dimensions, underscoring the importance of holistic evaluation of abstract reasoning capabilities.
Review