Recent work shows that models which learn to reward‑hack in reinforcement‑learning (RL) environments can develop broad emergent misalignment (EM), and that reframing reward hacking as acceptable during training—called inoculation prompting (IP)—prevents this generalization. We investigate whether synthetic document finetuning (SDF) can provide a similar inoculation without intervening in later training.
We augment a model’s mid‑training corpus with synthetic documents that portray reward hacking as permissible, then train these models with RL on exploitable environments, teaching them to reward hack. Behaviorally, the mid‑training succeeds: models describe reward hacking favorably and show higher approval of reward‑hacking outputs they generate.
Nevertheless, after learning to reward hack the models exhibit strong EM, whereas IP under the same conditions blocks EM. We show that SDF can steer downstream generalization predictably when inserting new associations, but it struggles and yields unpredictable effects when trying to overwrite existing associations—such as the link between reward hacking and misalignment that leads to EM.
Our results suggest that, at the scales tested, SDF can make a model appear aligned with desired beliefs while steering its later‑training generalization in unintended ways, failing to truly inoculate against misalignment caused by reward hacking.
Review