NeFut Logo NeFut
中 Admin Login

[CS.AI] Frozen Scenes, Shifting Winners: Configuration Fragility in Text-to-3D Evaluation

Published at: 2026-10-02 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #Evaluation

We investigate whether a text‑to‑3D leaderboard can shift when every generated scene is kept fixed, focusing on rendered‑image evaluation where camera parameters and caption wording are part of the measurement protocol. Using 300 frozen scenes from six generators, we vary eight rendering and caption factors, evaluate with 19 alignment metrics plus one perceptual‑quality control, and apply four targeted scene degradations.

The peak configuration variance exceeds the between‑generator variance for 17 of the 19 alignment metrics, and prompt‑bootstrap lower bounds are above 1 for 11 metrics. Rankings are more stable than raw scores, yet 18 of the 19 evaluators change their point‑estimate winner under some configuration. Pairwise protocol margin envelopes reveal which comparisons retain their direction across all tested settings. Selected pairs show opposite pointwise intervals, but no reversal survives simultaneous inference over the full search, meaning observed winner changes are descriptive rather than confirmed superiority shifts.

Sensitivity remains separate: even the prompt‑free control does not surpass a 67% tie‑adjusted directional discrimination on layout scrambling, which serves as a diagnostic rather than a human‑validated ground truth. The audit separates score stability, decision uncertainty, and targeted sensitivity, and recommends reporting (generator, score, card ID) together with protocol‑dependent comparisons and selection‑aware uncertainty.

Review

Original Source: https://arxiv.org/abs/2610.00447

[h] Back to Home