Long‑horizon agentic workflows demand that models keep performing state‑dependent actions while the context expands, sub‑task difficulty shifts, and new data arrives. Each condition forms an independent axis where an agent can fail. For instance, an agent reconciling a lengthy ledger must repeatedly read its state, update the correct entry, and stay aligned across thousands of outputs. A model may ingest the whole ledger yet lose its place or stop applying the operation consistently as generation proceeds. To probe this, we introduce Long‑Transduction, a controlled diagnostic that evaluates a model's ability to stay on task during extended generation while continuously reading, mutating, and outputting input‑context‑dependent operations such as arithmetic, sorting, variable lookups, and table transformations. Long‑Transduction independently varies local task complexity, input data formatting, and context length to isolate failures along each axis. Evaluating seven open‑weight models reveals a $62.8\%$ drop when scaling context from 4K to 128K, a $36.5\%$ drop with format changes, and a $39.9\%$ drop as local task complexity increases. Together, these shortcomings represent critical liabilities in long‑horizon agentic workflows.
Review