NeFut Logo NeFut
中 Admin Login

[CS.AI] FARE: Forensic Acceptance Region Estimation for Detecting Bait‑and‑Switch Image Generators

Published at: 2026-09-29 22:00 Last updated: 2026-09-30 01:41
#algorithm #AI #Machine Learning

Modern AI image generators are often offered as opaque APIs: users can query the service but cannot inspect model weights or architecture. This creates a concrete risk—after passing governance certification, a provider might silently switch to a cheaper, lower‑quality generator, undermining public trust and potentially endangering high‑stakes applications.

To address integrity auditing at deployment time, we introduce FARE (Forensic Acceptance Region Estimation). During certification, a large set of images sampled from the approved generator is used to train FARE, establishing an acceptance region for that specific model. After deployment, FARE can decide whether a single generated image is consistent with the enrolled generator, using only the image itself.

FARE’s features are derived from generator‑specific artifacts identified in prior forensic work. The training process deliberately seeks hard samples—images that lie near the boundary of the acceptance region—to tighten that region and boost sensitivity to subtle changes in the certified generator.

Across a variety of generator‑swap scenarios, including swaps between similar model versions and different architectural variants, FARE reliably detects the replacement. It consistently outperforms existing baselines at strict operating points and remains effective against the exact‑model and decision‑only attacks evaluated in this study.

Review

Original Source: https://arxiv.org/abs/2609.30982

[h] Back to Home