NeFut Logo NeFut
中 Admin Login

[CS.AI] COMPASS: Finding Where Reasoning Lives in Language Models

Published at: 2026-10-07 22:00 Last updated: 2026-10-08 01:25
#AI #Machine Learning #LLM

This paper introduces COMPASS, an inference‑time steering technique that discovers and exploits a latent direction inside large language models (LLMs) capable of eliciting reasoning. Prior work shows that explicitly prompting reasoning (e.g., Chain‑of‑Thought) improves performance, but it typically requires a predefined characterization of reasoning via prompt design, contrastive CoT, or SAE‑derived features. For mathematical reasoning with verifiable answers, the authors demonstrate that a far simpler signal—the correctness of the model’s own direct answer attempts—is sufficient to define a useful reasoning direction.

The correctness direction can be decoded from the activations of most attention heads, yet only a small subset of heads can be effectively intervened upon. COMPASS identifies these heads at inference time using a logit‑space attribution score and steers their activations along the correctness direction, requiring only per‑head activation statistics and no additional fine‑tuning.

Experiments span three model families (including GPT‑style and open‑source models) and multiple math benchmarks. Key findings include:

Overall, COMPASS reveals that a minimal correctness signal can uncover the internal reasoning substrate of LLMs and provides an efficient, transferable method for steering model reasoning.

Review: This work highlights a promising avenue for reasoning enhancement that bypasses complex prompt engineering, offering valuable insights for interpretability and model control.

Original Source: https://arxiv.org/abs/2610.07469

[h] Back to Home