Analysis of Conversation Trajectory Representations
A conversation can be read as a trajectory through semantic space, one point per turn, and a growing line of work scores conversation quality from the geometry of how that trajectory moved. The most prominent instance is TRACE (Gooding and Grefenstette, 2025), which predicts conversation-level satisfaction from roughly thirty hand-crafted scalar features — self-similarity of model turns, drift from a stated goal, turn-to-turn volatility, trends in relevance — fed to a random forest. This paper asks which summary of the trajectory actually predicts satisfaction, and measures the answer on three corpora.
The scalarisation problem
Every hand-crafted geometric feature embodies the same unexamined decision. The trajectory is a sequence of points in a few hundred dimensions; a feature such as “volatility” or “drift from the goal” collapses that object to one number, almost always through a distance or a cosine, before any learning happens. Thirty such features means thirty separate decisions about which direction of the geometry mattered, each made in advance by a human, each discarding whatever it did not anticipate.
The alternative is older and simpler than feature engineering: fix an orthogonal basis, project the trajectory onto it, and hand the learner the coefficient vectors. There is no selection step, the residual is quantified, and no direction is lost because nobody thought to measure it. The comparison this paper runs is between those two approaches, with everything downstream held fixed — identical off-the-shelf turn embeddings feed every arm, TRACE included, so no result can be blamed on turn-representation quality.
What the measurements say
The projection wins by a wide margin, and the margin is regime-dependent. Against TRACE’s geometric features the projection is 5-12 points better on all three corpora. On short, first-person-rated chats, though, the trajectory-shape channels collapse and the hand-crafted scalars become a genuine complement rather than a competitor. The dividing line is corpus and label construct, not conversation length — length is ruled out by an explicit deconfound arm that varies it directly.
The projection’s structure matters more than the basis it uses. Read as one interleaved sequence, a conversation alternates between two regions of embedding space at the highest frequency the grid supports, and a low-degree fit spends its structure describing that alternation. Splitting the trajectory by speaker — projecting each party’s turns on their own grid and concatenating — gives the learner each speaker’s content on its own path. That representation, keeping only a per-role mean and a per-role drift, is the paper’s recommendation: training-free, one formula, no feature selection, and never distinguishably worse than any measured alternative in either regime.
What the split buys is the content separation. The speaker-alternation aliasing it also removes is a representation-fidelity fact rather than the source of the accuracy — worth knowing, and worth not overselling.
The advantage is informational, not capacity. A 1,536-dimensional representation beating 26 scalars invites the obvious objection. Compressed per fold to the suite’s own 26 dimensions, the role-split projection still beats it by +4.1 and +4.6 points on the two task-oriented corpora, while a random projection to the same width is significantly worse. The information is in the representation, not in the room to store it.
The caution the paper ends on
This is offline preference prediction, and the winners are plausibly the more gameable ones. A representation built on content — mean pooling, or the projection coefficients — is more directly Goodhart-hackable than content-agnostic scalar geometry, because a model optimising against it can write the content. The ranking measured here need not transfer to a reward signal, and may invert. Anyone tempted to take the top of this leaderboard and optimise a policy against it should treat that as an open question rather than an implementation detail. Measuring it is what the manipulability paper does.
The paper’s Section 8 lists twelve limitations. The material ones: the comparison suite is our reimplementation from the published specification, since four of thirty features need timestamps that public corpora lack; the regime boundary is corpus-attributed rather than fully causal, with label construct, provenance, and goal grounding still bundled; and the comparison class is fixed representations, so “never distinguishably worse” is a claim within that class rather than about the best achievable encoder.
The reproduction kit
The repository carries the manuscript source, every artifact the paper cites, the roughly forty scripts that produced them, and a pipeline that regenerates all of it from public data.
./run.sh check # seconds: validate the paper against its shipped artifacts
./run.sh # many hours: regenerate everything from raw textThe check runs five gates on light dependencies alone — no datasets, no torch, no GPU. All 407 printed table cells are regenerated from the stored artifacts and compared cell by cell against the manuscript; every artifact is verified to have come from a clean checkout; all 83 <!-- src: --> tags are resolved to real scripts; and the committed PDF is rebuilt and compared word for word. CI runs the same gates on every push, so the paper cannot silently drift from its evidence.
Two facts about the setup are worth stating plainly. The comparison suite is our reimplementation of TRACE’s feature specification, with every deviation documented feature by feature in the appendix. And the LLM-judge arms make zero API calls: the ratings are frozen raw data, committed to the repository, and replayed — so a reproduction of those arms costs nothing and produces the same numbers rather than whatever the judge model does this month.
None of the three corpora is redistributed. Each is fetched from its original source at a pinned revision, so a reproduction reads exactly the bytes the paper read, under the original terms.
The practical recommendation
Split the trajectory by speaker, keep each speaker’s mean and drift, and hand that to whatever learner you were going to use. It requires no training, no feature selection, and no tuning; it beats a thirty-feature hand-crafted suite everywhere measured; and on the one regime where curated scalars help, adding them is an optional graft that has never been measured to hurt.