Scientific ideation is increasingly powered by large language models (LLMs), yet most existing systems are trained and evaluated on immediate proxies such as novelty, clarity, and feasibility. This leaves open the question of whether delayed scholarly uptake signals can steer models toward research directions with higher expected impact. We adopt citation‑normalized impact as a noisy but scalable proxy for academic adoption and investigate its usefulness. A dataset of over 100K computer‑science papers is built by extracting goal‑conditioned idea descriptions and assigning each paper an ordinal, year‑normalized citation label. A goal‑conditioned reward model is then trained to predict these citation labels from goal‑idea pairs, and the reward is used to align an idea generator via supervised fine‑tuning followed by reinforcement learning (RL). To avoid circularity, evaluation follows a held‑out, reference‑grounded protocol: for a given research goal, model‑generated ideas are compared against historical ideas, and judgments are weighted by the reference idea’s citation label. Experiments show that the RL‑tuned model consistently yields ideas with higher estimated impact than both the base model and supervised‑fine‑tuning baselines. Our findings demonstrate that scientific impact can serve as a practical, outcome‑grounded feedback signal for aligning LLMs in open‑ended scientific discovery.
Review