Manipulability of Trajectory-Geometric Conversation Rewards and a Role-Restricted, Task-Conditioned Construction

Reinforcement learning wants a dense reward — something that scores every turn — in exactly the settings where the true outcome is sparse and arrives late. A customer-service call resolves or fails once, at the end; an agent’s task succeeds or does not after fifty tool calls. Trajectory geometry offers a tempting substitute, because a simple head reading the shape of the conversation so far genuinely predicts eventual failure before the conversation ends. This paper asks whether that predictor can serve as the reward, and answers with a measured no and a construction that survives.

The attack, and why its numbers are lower bounds

Optimisation pressure was applied without training a policy. The reward is attacked by selection over real turns: substituting in agent turns drawn from the corpus that historically accompanied success, then asking how many of the head’s failure alerts disappear. No gradient, no fine-tuning, no synthesis — just picking from things a real agent actually said.

That adversary is weak by construction, and deliberately so. It cannot invent the novel turn a generative adversary would find, so every manipulability figure in the paper is a floor rather than a ceiling. The finding is that even this floor is far too high.

The naive reward fails, and the reason is not what it looks like

A head reading all roles reaches 0.788 AUC on ToolBench, and selection alone suppresses half its failure alerts (manipulability 0.50); on the customer-service corpus the same attack suppresses 87% of them. The predictor is a monitor, not a reward, and no transform or combination of the adversary-controlled channel fixes it. The RL optimiser is irrelevant to the vulnerability — the exploit is available before any training begins.

The instructive part is the diagnosis. Exploitability turns out to be a property of which role the adversary can write, rather than a property of scoring “content” at all. A model optimising the reward controls only its own turns. The user’s turns, the tool calls, and the tool results are fixed evidence of what actually happened, and no amount of optimisation pressure lets the policy author them.

ClassRolesWho writes it
Fixed evidenceuser, tool call, tool resultThe world and the record
Adversary-controlledmodel / agent turnsThe policy being optimised

The construction

Every stage of the paper’s construction is a decision about which of those two classes enters the score.

  1. Drop the written role. Scoring only the fixed evidence gives 0.769 AUC against the naive head’s 0.788 — 97% of the accuracy — at manipulability exactly zero, by construction rather than by measurement. The attack rewrites turns the score never reads.
  2. Re-admit the model’s writing exactly once, as conditional typicality. Scoring the model’s turns only as their consistency with the fixed evidence, never as free content, brings accuracy to 0.782 and manipulability to −0.10: mildly anti-gameable, because the substituted turns fit the surrounding evidence worse than the originals did.
  3. Condition on the task. Where the outcome is recomputable from the visible actions, fooling the reward requires actually completing the task. Task-blind, the same detector is fooled by generic success-shaped actions 82% of the time; task-conditioned, that slack falls to zero.

The recovered accuracy is not error-counting in disguise. A running error-rate baseline over the same tool results reaches 0.649, and the fixed-role head beats it by 0.120 AUC with disjoint bootstrap intervals. What the geometry reads is the returned evidence itself, which early prefixes carry long before an error count becomes informative.

The early warning is real, and bounded

At a matched false-alarm rate, the fixed-role head catches 37% of ToolBench failures and 40% of ABCD failures at least two decision points early, where a same-turn-content baseline at the identical operating point catches 29% and 35%. Median lead is two decision points on ToolBench and six on the longer customer-service dialogues. The head also survives transfer to real spoken transcripts, attenuated.

Two boundaries are just as measured. On code-agent trajectories the head collapses to chance, so this is a dialogue result rather than a general agent result. And per-role velocity — the degree-1 slope, the direction the conversation is moving — adds essentially nothing to binary outcome prediction (+0.006 AUC) while carrying an order of magnitude more signal for per-turn satisfaction (+0.08 to +0.12 Spearman). Level and velocity answer different questions, and the deployment is one degree-0 head read at two positions.

What the paper does not claim

The manipulability figures are lower bounds against a substitution adversary; a generative one would find more. “Un-gameable” is scoped to that adversary, and the decisive substance result is scoped further to task-conditioning, where the label is recomputable from the actions — the 82% task-blind slack is the standing reminder of what a task-blind detector concedes. No policy was trained and no RL loop was run. And the corpora are public role-play, predominantly Wizard-of-Oz or crowdworker, rather than production traffic.

The reproduction kit

The repository carries the manuscript, the full machine-checked record it curates from, the pre-registrations, every artifact, the scripts, and a pipeline that regenerates all of it.

./run.sh check     # seconds: validate both documents against their shipped artifacts
./run.sh           # hours: regenerate everything from raw text

The check regenerates 523 printed table cells — 242 in the manuscript, 281 in the record — and compares each against what is printed, verifies every artifact came from a clean checkout, resolves all 46 source tags to real scripts, and rebuilds the committed PDF for a word-level comparison. The pipeline’s DAG is data-gated, so a fresh clone with nothing fetched runs green and you can grow the reproduction one corpus at a time.

Two disciplines beyond the harness are the paper’s own. Researcher degrees of freedom were fixed in a committed pre-registration before each round’s outcome-correlated results were seen, kill conditions in writing. And judging is against verified outcome labels and per-turn human ratings, never an LLM judge — an earlier spike found the judge more biased toward the exploit than the reward was exploitable, which makes it the wrong instrument for measuring exploitation.

One reproducibility detail worth knowing: exact manipulability figures are bit-reproducible only under the pipeline’s pinned single-thread BLAS. Multithreaded float-summation order drifts the last decimals, so ./run.sh check is authoritative and an ad-hoc rerun of a single stage is not.

The takeaway for anyone building one of these

Trajectory geometry earns its place as a monitor. Turning it into a reward requires deciding, explicitly, which channels the thing being optimised is allowed to write — and then scoring only what it cannot. The accuracy cost of that discipline, measured here, is about three points.