Gram Projections for Conversation Trajectories
Embed each turn of a conversation and the conversation becomes a path through semantic space. Before a simple learner can use that path, it has to be squeezed into a fixed-size vector, and one principled way to do it is to project onto an orthogonal basis and keep the first few coefficients. The Discrete Cosine Transform is the reflexive choice for that job, carried over from signal processing along with a compaction argument about where energy sits in the spectrum. This paper asks whether that argument survives contact with conversation trajectories, which are short, aperiodic, and end somewhere different from where they began.
The alternative it introduces is the Gram projection: the same least-squares projection, onto discrete orthogonal polynomials rather than cosines.
G_k = sum_i phi_k(i) * e_i / sum_i phi_k(i)^2
Here e_1 … e_T are the L2-normalised turn embeddings and phi_k is the degree-k Gram polynomial on the uniform grid. Each coefficient is itself a vector, and the representation is their concatenation.
Coefficients you can read
The polynomial basis buys stable semantics, and the paper establishes four properties exactly rather than asymptotically.
- Degree 0 is the mean. Mean pooling, the most common trick in the field, is the degree-0 special case of this representation.
- Truncation never changes what you kept. Coefficients come from independent inner products, so adding a degree leaves the lower ones untouched. Every prefix is the optimal least-squares summary of its degree.
- Degree 1 is the least-squares drift. Twice the negated degree-1 coefficient is exactly the end-to-end displacement of the fitted line, so “the conversation drifted toward X” becomes a statement about one coefficient.
- The drift coefficient means the same thing at every length, which is what makes a six-turn chat and a forty-turn call comparable on the same axis.
The DCT shares the first two exactly. Readable coefficients and the length-stable drift belong to the polynomial basis alone.
The measured verdict: the choice is free almost everywhere
Interpretability usually costs accuracy, so the interesting question is the price. Across three conversation-satisfaction corpora at matched dimensionality, there is none worth reporting.
Truncated Gram and DCT spans overlap 94-96%. Downstream accuracies tie at the instrument’s measured resolution of about ±2.5 points, with every interleaved-arm confidence interval spanning zero. Structural changes — splitting the projection by speaker so each party’s turns get their own fit — transfer across bases, which says the gain belongs to the representation’s structure rather than to the polynomial family. Structure is worth more design effort than basis choice.
Cost arguments do not rescue the default either. At truncation, Gram and DCT are the same matmul shape with the same asymptotics; the DCT’s celebrated O(T log T) advantage applies to full spectra that sequence representations never compute. Both are dominated end to end by the embedding step, by a measured 780-22,000x across recorded hosts. Every representation compared in the paper costs the same to run.
Where the tie breaks, and how far the claim goes
One exception has a location. On populations mixing very short per-role sequences — one to two turns per speaker — the polynomial basis wins by a measured 2.3 points, CI [+0.3, +4.4], p = 0.022. The proposed mechanism is that the DCT’s leading coefficient confounds drift magnitude with sequence length, which predicts exactly that localisation.
The paper is deliberate about how far this goes. The mechanism stands as a supported hypothesis rather than a demonstrated law: one interventional design was algebraically unable to detect the mechanism’s variance form, the corrected design was under-powered on this corpus family, and the localisation test ran on the corpus that generated the hypothesis, so a content or label-construct confound on short first-person chats is not excluded. What the evidence supports is a one-sided practical asymmetry. Nothing here weakens the tie at conversational lengths, and short-chat populations have a measured, mechanism-consistent reason to prefer the readable basis.
Leaving the “DC-plus-smooth” class entirely carries its own price, and one cheap ablation prices it in advance. The Discrete Sine Transform has no constant column and so cannot represent the mean; it sits at the Gram level on two corpora and loses 3.5 points on the one corpus where trajectory shape carries marginal signal beyond the mean. The cost is zero where shape adds nothing and appears exactly where it does.
The reproduction kit
The repository is the paper’s evidence rather than an appendix to it. It holds the manuscript source, every result artifact the paper cites, the scripts that produced them, and a pipeline that regenerates all of it from public data.
Validation takes seconds and needs no datasets, no torch, and no GPU:
./run.sh checkThat runs five gates, the same ones CI runs on every push. Every table cell is regenerated from the stored artifacts and compared against the cell printed in the manuscript — 138 cells across five tables, strict equality. Every artifact is checked for provenance, so a number computed from a dirty working tree does not count. Every <!-- src: --> tag in the paper resolves to a real script, so no figure or table floats free of the code that made it. And the committed PDF is rebuilt and compared word for word against the shipped one.
A full reproduction is one command more:
./run.shIt re-embeds both corpora from raw text, regenerates every artifact, figure, and table, and rebuilds the PDF, resuming cleanly if interrupted. Regenerated artifacts overwrite the committed ones in place, so git diff afterwards is a direct diff of your numbers against the paper’s. An empty diff is the reproduction.
Neither corpus is redistributed. The USS repository carries no license file and PRISM is gated behind terms each user accepts in their own name, so both are fetched from the original source at a pinned revision — a reproduction reads exactly the bytes the paper read, under the original terms.
The recommendation
Within the DC-plus-smooth class, choose the basis whose coefficients you can read. The accuracy is free, the compute is identical, and the result is a representation where a trend is a single number rather than a linear combination of cosines that nobody can interpret at a glance.