NeFut Logo NeFut
Admin Login

[CS.AI] Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting Agents

Published at: 2026-09-15 22:00 Last updated: 2026-09-16 00:22
#Machine Learning #optimization #LLM

Vibe Patenting is an end-to-end patent‑drafting testbed designed to evaluate AI agents in a professional patent workflow.

The evaluation uses a separately‑invoked LLM judge that reads the generated draft, returns structured feedback, and guides the agent through iterative revisions.

Across several inventions and drafting‑agent configurations, judge‑guided revisions consistently raise the judge‑assessed quality, while unguided revisions quickly plateau.

Key findings are that judge feedback enables a low‑reasoning agent to approach the performance of a much more expensive high‑reasoning agent; stronger models and more reasoning steps generally improve drafting quality; and domain‑specific agentic workflows add further gains.

We validated the judge against an independent assessment by a professional patent attorney and observed meaningful agreement that varies strongly with the chosen metric, along with systematic calibration differences.

These results highlight that LLM judges can serve as useful evaluators and optimization signals in complex professional workflows, yet their utility is limited by metric selection and calibration issues.

Review

Original Source: https://arxiv.org/abs/2609.13422

[h] Back to Home