We explore the feasibility of using automated rewards to train language models for conversational humor, focusing on how rewards can be exploited and how to counteract such exploits. Two reward families aim to capture understandable surprise and predicted audience amusement.
Controlled experiments reveal that an embedding‑based surprise reward treats word‑shuffled replies as readily as witty ones. A fluency filter is added to catch the shuffles, yet the combined reward still rejects some genuinely witty replies and fails further validation.
The audience model’s predicted laughter is vulnerable to laughter cues embedded in either speaker’s messages. Normalizing these cues across speakers blocks the known attacks, although unmatched expressions remain exploitable.
We conduct three reinforcement‑learning runs, each incorporating successive reward revisions. The final run improves the overall evaluation score by 0.0903 and cuts zero‑score sessions by 40%, but the humor‑specific gain stays below our preregistered target.
These findings illustrate a broader challenge for automated reward design: countermeasures must block exploitable shortcuts while preserving the behavior the reward is intended to encourage.
Review