Sustainable protein discovery lacks fast computational proxies comparable to molecular docking or density functional theory, so expensive human sensory panels are needed to judge whether a novel food tastes like its animal‑based counterpart, creating a bottleneck in the design‑build‑test loop.
To address this, we introduce TasteBench, a multimodal benchmark coupled with a privacy‑preserving competition, comprising two tasks: a food‑level ranking task built on over 21K human evaluations covering 215 plant‑based foods across 24 product categories, yielding 935 within‑category ranking pairs; and a molecular‑level taste classification task involving 15K flavor molecules.
For rigorous model assessment we characterize ground‑truth reliability: inter‑rater agreement is low (Krippendorff's $\alpha = .077$), while the split‑half reliability ceiling of panel‑aggregated rankings is $.825$, defining the performance range models should aim for.
Baseline models were evaluated across four input modalities. On the same pairs rated by panelists, the best model achieves a pairwise accuracy of $.661$, comparable to the median individual panelist ($.650$), and reaches $.683$ across all within‑category pairs.
TasteBench provides the evaluation infrastructure and baselines to quantify progress in computational screening for sustainable protein discovery.
Review