NeFut Logo NeFut
Admin Login

[CS.AI] DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation

Published at: 2026-09-03 22:00 Last updated: 2026-09-04 02:14
#AI #Machine Learning #optimization

Commercial short‑drama production follows five stages: script, storyboard, keyframe, shot‑level video, and the final cut. Existing benchmarks evaluate only the video‑generation stage with pre‑authored inputs, leaving two questions unanswered: whether each stage respects the original script intent and whether assembled shots remain coherent. DramaChain Bench is the first benchmark that covers the entire chain, built on three in‑house systems that share a common DramaChain Dimensions framework. The framework defines five evaluation axes at every stage, instantiated into 63 leaf dimensions.

DramaChain Agent is calibrated against commercial short‑drama platforms in both workflow and final‑cut quality, enabling fair stage‑wise comparison across models. The DramaChain Labeling System scores 5,785 items independently by three professional annotators; defects are localized spatio‑temporally and drawn from a predefined defect list, yielding 17,488 valid scores and 255,925 traceable attribution records. Human annotations reveal that upstream defects cascade through the pipeline, so final quality is not governed by video generation alone.

DramaChain Agentic Judge automatically scores each leaf dimension, gathering evidence over multiple agentic rounds before judging against a per‑item checklist. Its rankings correlate with human rankings at a mean PLCC of 0.918, sufficient to admit new models without any annotation cost.

Review

Original Source: https://arxiv.org/abs/2609.00646

[h] Back to Home