Protein modification must explore an astronomically large sequence space, yet wet‑lab validation is costly and low‑throughput. Recent computational approaches—protein language models (PLMs), large language models (LLMs), and LLM‑based agents—show promise, but their relative effectiveness in realistic experimental decision‑making remains unclear. To address this, we introduce the PFArena benchmark, which provides four controlled task interfaces covering single‑mutant generation and multi‑mutant ranking. By supplying varying levels of mutation fitness data, PFArena mimics four representative research scenarios with differing amounts of prior experimental context.
We evaluate six PLMs, six LLMs, and five LLM‑based agents using complementary metrics that capture both peak and overall protein‑modification performance. Results reveal a systematic shift in model behavior as target‑specific experimental evidence becomes available: PLMs excel at open‑ended single‑mutant generation by leveraging protein‑specific priors, whereas LLMs and agents achieve strong performance in multi‑mutant ranking, particularly when fitness data are provided. Nonetheless, all model families encounter fundamental difficulties as the search space grows and mutation depth increases.
Our code and benchmark suite are released to enable reproducible research in model‑assisted protein modification. Review