NeFut Logo NeFut
中 Admin Login

[CS.AI] MoMHa: Multi-Objective Optimization of LLM Harnesses over Accuracy, Safety, and Tokens

Published at: 2026-09-29 22:00 Last updated: 2026-09-30 01:41
#AI #optimization #LLM

Most work on improving large language models treats accuracy as the sole objective. We argue that the surrounding Python harness—code that builds prompts, routes calls, and parses outputs—is itself a first‑class design surface that must balance multiple objectives: a harness that is accurate but never refuses unsafe requests, or that consumes an order of magnitude more tokens, is not useful. To address this, we introduce Meta‑Harness, which casts harness design as a search over three per‑domain objectives (accuracy, behavioral safety, and token cost) and solves it with an agentic proposer, Claude Code, that has full filesystem access to prior harness source, execution traces, and scoring artifacts. Our central finding is that a single‑phase joint‑reward proposer (MoMHa) outperforms every alternative, including a two‑phase “accuracy then tokens” ablation, scalar‑only feedback, and an accuracy‑only baseline. We evaluate on seventeen domains: seven synthetic capability suites, seven real‑world public benchmarks (HumanEval, MBPP, Spider, FEVER, MMLU‑Pro, LawBench, NuminaMath), and three user‑specific safety domains derived from U‑SafeBench, using a fleet of twelve models spanning four families. On the synthetic track MoMHa achieves a joint mean of 0.482 versus 0.198‑0.422 for ten baselines, winning 7 out of 10 per‑domain columns; on the real‑world track it scores 0.461 versus 0.377 for the strongest baseline (DSPy), winning 5/7 columns, demonstrating that harness strategies transfer to unseen benchmarks without retraining on 8 of 12 target models. MoMHa attains the highest measured behavioral safety composite (U‑SafeBench, 0.781) and uses 95 fewer tokens per example than the two‑phase alternative. We will release all harness code, evaluation infrastructure, and cross‑model logs.

Review

Original Source: https://arxiv.org/abs/2609.30967

[h] Back to Home