NeFut Logo NeFut
中 Admin Login

[CS.AI] Style, Not Self: Surface Cues Explain Zero-Shot Code Attribution by Large Language Models

Published at: 2026-09-26 22:00 Last updated: 2026-09-28 00:49
#AI #Machine Learning #LLM

We examined whether current commercial large language models (LLMs) can recognize code they have generated under zero‑shot conditions. Five models produced solutions for the MBPP, HumanEval, and DS‑1000 benchmarks, while seven additional models generated only for MBPP. These models were then used as evaluators across four tasks:

  1. Selecting their own solution from a pair;
  2. Judging whether a single solution is theirs;
  3. Identifying which of two solutions was written by a named model;
  4. Blindly rating code quality.

In the single‑solution task, balanced accuracy ranged from 49% to 58%, whereas raw accuracy (38%‑67%) mainly reflected each model’s propensity to claim authorship. In the pairwise task, accuracy across 14 evaluator‑opponent combinations correlated with the evaluator’s solution length at $r=0.93$. Attribution to a named model succeeded on some pairs but was consistently inverted on others.

We introduced a rule‑based normalization that strips docstrings, comments, type hints, and local variable names. This procedure preserved Pass@1 while reducing ten of twelve re‑tested results to chance level; the remaining two followed the length difference left after normalization, although a trained classifier still separated most normalized pairs. Claude Haiku’s self‑preference also vanished after normalization.

Recommendation: report balanced accuracy, include heuristic baselines, and verify label consistency to avoid over‑interpreting a model’s self‑attribution ability.

Review: The work shows that LLMs rely heavily on surface cues such as code length rather than deeper semantic signatures when attributing code, highlighting the need for more robust evaluation designs to prevent inadvertent collusion or self‑bias.

Original Source: https://arxiv.org/abs/2609.30048

[h] Back to Home