NeFut Logo NeFut
中 Admin Login

[CS.AI] Generation Provenance Before Behavior Attribution: Auditing Synthetic Speech Research Objects

Published at: 2026-10-03 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning

Attributing model behavior to synthetic training data requires first knowing the generation source of each sample before estimating its effect. A simple waveform‑label pair does not retain this information. We propose a generation‑provenance substrate that binds source specification, generated content, waveform, target, fact requirements, quality signals, review lineage, and an immutable manifest identity. The producer and selection mechanism determine evidentiary meaning; storage location and variable name do not.

We audited this substrate in a private Japanese hand‑off pipeline. The audit population contains 113 assets, totaling 1.552 hours of synthetic speech across six scenario families; each item links audio, transcript, candidate notes, and fact‑check checklists, but human evidence is selective and source‑specific. Two fidelity‑only manifests are scenario‑seed disjoint and immutably versioned, while exact upstream attribution is blocked by floating generator aliases, missing per‑clip TTS and code stamps, and an unversioned checking prompt.

We argue that generation provenance is a necessary precondition for behavior attribution but not sufficient for causal attribution: it defines the candidate causal graph and audit units, whereas true contributive attribution still requires frozen training runs and intervention or influence evidence. The paper contributes a compact provenance contract, an audit protocol, and a bounded case study for synthetic‑data attribution; controlled research access may be offered, but we do not claim causal training‑data attribution, clinical validity, or unrestricted public release.

Review

Original Source: https://arxiv.org/abs/2610.01378

[h] Back to Home