Professional work often begins with a brief request, after which the professional must determine the required information, relevant documents, and whether the request’s premise is valid. We introduce DAYJOB, a benchmark comprising 130 tasks created by practitioners in healthcare (50) and finance (80). The tasks are estimated to take on average 13.6 hours in healthcare and 16.6 hours in finance. Each task is delivered as a containerized Harbor environment together with an expert‑crafted binary rubric (median of 47.5 criteria for healthcare and 57.5 for finance) that an agentic judge must satisfy completely to pass. We evaluated 30 model configurations from 13 developers; the best performer, Claude Opus 5.5, achieved pass rates of 24.7% on healthcare and 23.9% on finance, while the median configuration passed only 0.6% and 2.5% respectively. Case studies reveal that agents sometimes accept premises contradicted by the record and propagate incorrect inputs through otherwise consistent analyses. We release all healthcare tasks, 50 of the 80 finance tasks, the evaluation harness, and the leaderboard.
Review