Research paper
Manipulability of Trajectory-Geometric Conversation Rewards and a Role-Restricted, Task-Conditioned Construction
23 July 2026 · 13 pages · PDF
Abstract
A dense reward that scores a conversation on every turn is what reinforcement learning wants where the true outcome is sparse and late — and what Goodhart's law warns will be gamed. We show it need not be. Trajectory geometry is a genuine within-domain failure signal — it predicts failure before a task-oriented call ends — but as a reward it is gameable by selection over real turns alone. Its exploitability, however, is not a property of "content": it is a property of which role the adversary can write. Score the fixed evidence — user turns, tool calls, tool results — admit the model's reasoning and response only as its consistency with that evidence, and condition on the task: the construction recovers 97-99% of the predictor's accuracy while remaining un-gameable by the selection adversary that breaks the naive proxy.
Read it here
Trouble viewing? Open the PDF in a new tab with the button above.
Reproduction. The paper ships as a self-contained kit on GitHub — a portable pipeline spec plus the generator scripts, regenerating every artifact from public data with one command. Details in the reproducibility statement and appendices.