A side note to Part I, developing one sentence: the quality of data embedding sets the ceiling of what the system can do; the quality of data encoding determines how close it operates to that ceiling.
One Model, Four Harnesses
Yu et al. ran a single set of weights - Qwen3-30B-A3B-Thinking, one benchmark, one task set - through four agent harnesses and reported this:1
| Column 1 | Column 2 | Column 3 | Column 4 | Column 5 |
|---|---|---|---|---|
| ClawEval | ZeroClaw | ReACT-style loop | Codex | OpenClaw |
| pass@1 | 32.5 | 26.1 | 12.2 | 11.4 |
| pass@3 | 44.7 | 39.8 | 18.6 | 19.3 |
The only variable is the machinery wrapped around the model, and the spread is nearly threefold.
A model-generation upgrade delivering a threefold gain would be the story of the decade; the same magnitude is sitting in the harness, unclaimed, and almost nobody is reporting it. A benchmark number quoted without its harness is close to meaningless.
If the harness decides that much, then the structure feeding it - what the model can see, what it may do, what counts as done - is doing work that no amount of model improvement substitutes for. In the enterprise, that structure has a name: the ontology. And if the harness decides that much, it belongs inside the training loop rather than outside it, which is what OpenForgeRL: Train Harness-native Agents in Any Environment was built to make possible.
The Ceiling and the Operating Point
Two qualities sit behind LLM performance, and they are worth separating.
The first is what a model is capable of under ideal conditions - given the right context, the right tools, and an unambiguous task, how good is the output? Call this the ceiling. It moves with pretraining scale, architecture, post-training, and reasoning methodology. Benchmarks are the closest thing to a reading on it, and they get there by stipulating the conditions - a fixed task set, a fixed harness, a fixed sampling budget, all written down. That stipulation is what makes the number mean anything, and it is what almost every public conversation about AI progress is actually about.
The second is what a model achieves in a real deployment, where a retrieval system of uncertain quality assembles the context, the tools are partially documented, the task arrives underspecified, and half the relevant state lives in systems the model cannot see. Call this the operating point. Nobody stipulated these conditions. They vary per query, and they go unreported.
The gap between them is large, and nothing guarantees it closes on its own. A spread of that size, sitting entirely in the machinery, is wide enough to swamp the difference between model generations - which would put a frontier model in a poorly assembled deployment below a weaker model in a well-assembled one. Nobody has measured that ordering directly, and the inference assumes the harness effect carries across models. It is still the inference a buyer has to price.
The table above is that gap made concrete. One set of weights, four harnesses: the ceiling did not move, and the operating point spanned a threefold range.
Part I’s taxonomy explains why. Data embedding - what is compressed into the weights - fixes the ceiling. Data encoding - what enters through the context window at inference time - determines where in the space beneath that ceiling the system lands. The two levers are complementary rather than sealed off from each other: a model post-trained on longer horizons and better tool-use traces does not simply sit higher, it gets better at exploiting whatever context it is handed, which raises the return on harness work rather than substituting for it.
A model upgrade buys a higher ceiling and leaves the distance to the operating point largely intact, because that distance is set by facts about the deployment - what the retrieval returns, which tools are documented, what state is visible at all - that no change to the weights reaches. The industry invests overwhelmingly in the ceiling lever anyway, for the same reason Part I identified in the scaling debate: it is legible. You can price a model upgrade. It is considerably harder to price “we fixed how the agent finds things.” That is a leverage-gradient error with the exact shape Part I already drew, one level up: there, the field collapsed scaling models into scaling compute and turned a research question into a procurement one - buy GPUs, move along the known curve. Here, a model upgrade is the procurement answer to a question about harnesses, priced and forecastable in exactly the way GPU acquisition was, and aimed at the lever that was already being pulled hardest.
None of which says the bitter lesson is wrong. It says the lesson isn’t sufficient, and the insufficiency sits in a specific place. Sutton’s argument governs the relationship between the first two categories. It has no account of the third, so it can tell you that handcrafted knowledge loses without telling you which encodings to learn, or where the learning should happen. Yu et al. supply one answer: if the harness decides this much of realized performance, it belongs inside the training loop rather than bolted on outside it. Almost everyone who noticed that the harness mattered responded at inference time, with better prompts and better retrieval. OpenForge trains through it instead. The open question Sutton leaves is not whether human structure belongs in the system, but which kind survives as models improve - a question about what the structure does, not how much of it there is.
Attention Is a Budget
The phrase “context window” invites a false picture: a container you fill, where more capacity is strictly better and irrelevant content is merely wasted space.
The reality is closer to a contested budget. Attention distributes finite weight across everything present. Adding irrelevant material does not simply fail to help - it competes. Tokens that do not bear on the task still absorb attention mass, still participate in retrieval, and still supply plausible-looking material for the model to reason from. A retrieval system that returns fifty documents of which three are relevant has not given the model more to work with. It has given the model a haystack and asked it to do the search the retrieval system was supposed to do.
This is why long context has not dissolved the problem the way many expected. Expanding the window increases the amount of material you can supply without improving your judgment about what to supply. If the selection is poor, a larger window degrades precision while enlarging the cost. The constraint was never storage. The constraint was selection, and selection is a modeling problem, not a capacity problem.
Yu et al. supply a pointed illustration. Their most elaborate harnesses were not their best. Of the four they evaluate, the two simplest - a ReACT-style loop and ZeroClaw, a lightweight shell that makes tool registration easy - reached the highest performance. OpenClaw, with its rich set of built-in tools and control flows, saw only moderate gains from training “while consuming far longer prompts and contexts than the others.”2
More apparatus, longer context, worse result. Part of that penalty is the adapter: neither OpenClaw nor Codex exposes a custom-tool interface, so ClawEval’s tools reach them as skill files instead.3 That wrapping cost is not a defect in the experiment - it is what deploying into a rich, closed harness actually feels like, paid in context on every rollout.
The technical attention mechanism and the ordinary sense of “focus” are not the same thing, and the argument does not depend on conflating them. The claim is narrower and empirical: performance on a task is a function of the signal-to-noise ratio of the supplied context, not merely of whether the necessary information is present somewhere within it. Presence is not access.
The Flat Map Problem
In Attention, Hierarchy, and Traversal, I argued that attention and hierarchy belong to the traversal over a graph under resource constraints. An agent moves through a network under resource constraints, forced at every branch point to choose which edge to follow, and the ordering that choice implies comes from the walk rather than being read off the map. A qualification is owed, because the graph is itself an artifact of its author’s choices. Whoever built the representation decides what gets a node, what gets a typed edge, and what is left out altogether, whether by intention or ignorance. Inclusion is already a coarse act of prioritization - whatever is in the graph outranks everything omitted from it. That is the only ranking a map supplies. Among the things it does include, it is silent on which matters now, for this query, under this budget, and that ordering has to be generated afresh every time. All maps are flat; the walk is hierarchical.
Retrieval systems built on embedding similarity hand the model a flat map and ask it to construct the walk from nothing, on every single query.
Similarity search returns what is semantically near. It has no representation of what is causally, contractually, or structurally connected. Those are different relations, and the difference is not academic. Consider a query like: which of our contracts are exposed if this port closes for three weeks? No document in the corpus contains that answer. The answer exists only in the joins: port → vessel → shipment → purchase order → contract → penalty clause. It is a multi-hop traversal over typed edges, and the number of documents semantically similar to the query is close to zero because the query is about a relationship rather than a topic.
An ontology is a precomputed traversal structure. It is not a store of answers; it is a store of edges, typed and named, so that the walk is cheap. The model still does the reasoning. The ontology removes the requirement that the model first rediscover the shape of the world before it can reason about it, at a per-query cost, starting from no privileged information about which of forty thousand tables are relevant.
Framed economically, the ontology amortizes a discovery cost. The relational structure has to come from somewhere. Either it is declared once and maintained, or it is inferred repeatedly at inference time. The second is not impossible, and it gets cheaper every year as models get better at schema inference. But cheaper is not free, and per-query costs multiplied by enterprise query volumes remain the dominant term.
This gives the argument a falsification condition: if a model with a large context and naive similarity retrieval matches an ontology-grounded system on multi-hop relational queries at comparable cost, the ontology-value claim above is wrong.
What an Ontology Commits To
Palantir’s Ontology is the clearest worked example of such a structure. It defines objects, properties, links, and actions over an organization’s data: a shipment is an object; it links to a supplier, a vessel, a port, a contract; those links are declared, typed, and traversable. And critically, the Ontology exposes actions - typed write-backs carrying authorization rules and submission criteria, so that an edit to an object, a property value, or a link is validated and permissioned before it commits.4
The graph here is a logical structure, not a storage decision. It matters little whether the backing store is a property-graph engine or a set of relational tables with a resolution layer over them; what matters is that relationships between entities are first-class declared objects, not something to be rediscovered by joining on a hunch. Calling it “a graph database” undersells it. The Ontology is a committed model of what exists in this organization and how it connects.
The interesting part is that none of this was built for language models. Foundry dates to around 2016, and the Ontology took its current shape as a first-class product in the two years before ChatGPT; human analysts did the reasoning, and they needed integrated data they could act on under permission. When AIP arrived in April 2023, the layer needed no redesign.5 Models could be pointed at a declared, permissioned map of the enterprise because one was already sitting there.
The bet was about integration, resolution, and permissioning - that the hard part of enterprise software sits below the analytics rather than in them. The LLM era rewarded that bet for a reason nobody had underwritten, and the layer changed jobs rather than becoming redundant: it stopped being the thing a human reasoned over and started being the thing a model reasons through.
That transition generalizes well beyond Palantir.
What the Symbolic Layer Is Actually For
Neurosymbolic AI has a bad reputation, and mostly it has earned it. The classical program - a symbolic reasoner performing inference, with a neural network relegated to perception at the edges - is close to the canonical bitter-lesson failure. The symbolic reasoner held the human knowledge; it was fused to the architecture, and it lost.
The version that is working now inverts the division of labor, and the inversion is the whole point. The neural system does the reasoning. The symbolic system does three other jobs:
- Index. What can be seen. The ontology determines the retrievable surface and the shape of traversal.
- Action schema. What can be done. Typed, permissioned operations with declared preconditions and effects. The model chooses; the schema constrains the choice to well-formed, authorized, reversible operations.
- Verifier. What counts as done. Type checks, constraint satisfaction, referential integrity, test execution, policy evaluation - deterministic judgments about whether an output is admissible.
None of these three requires the symbolic layer to perform inference. Each constrains the interface rather than the inference.
That distinction resolves what would otherwise be a direct contradiction with The Encoding Problem. Part IV argued that encoding helps when it satisfies separability, granularity preservation, and override capability, and was explicitly suspicious of agent frameworks on the third count - rigid workflow rules that force a model through predetermined steps even when a different approach would be better.
The suspicion was correctly aimed but imprecisely targeted. What fails the override test is encoded procedure: a hand-written pipeline that dictates the order of reasoning. What passes it comfortably is encoded structure: a declaration of what entities exist, how they relate, what operations are legal, and what a correct result looks like. A workflow rule that says “first query the database, then validate, then format” constrains processing. A link type that says “a shipment has exactly one vessel” constrains nothing about processing - it enlarges what the model can access and narrows only the space of malformed outputs.
The harness comparison above reads as a small natural experiment on precisely this distinction. ZeroClaw offers minimal procedure and an easy path to declare new tools - a thin interface, cleanly typed. OpenClaw and Codex offer elaborate built-in control flow and resist custom tools - thick procedure, awkward interface. The thin-interface harnesses won, and they also responded best to training. An encoded interface helped; an encoded procedure cost context and returned less.
Encoded structure raises granularity rather than reducing it. This is the opposite of the feature-engineering failure mode, where reducing images to edge maps destroyed information the model might have needed. A well-built ontology destroys nothing; it exposes relationships that were previously latent in the storage layer and unavailable to any consumer without a hand-written join.
So the encoding effectiveness principle survives, sharpened. Constrain the interface, not the inference. Declare structure, not procedure.
Answering the Bitter Lesson Objection
The strongest reply to everything above is the standard one:
All of this is scaffolding. Every generation of AI practitioners has built elaborate structures to compensate for model deficiencies, and every generation has watched the next model absorb those structures. Ontologies are the opening books of the 2020s. Retrieval architectures are this decade’s feature engineering. Longer contexts, better native retrieval, and agentic post-training will make the semantic layer a curiosity. Building your company on it is building on sand.
The objection is partly right and importantly wrong.
It is right about the specific scaffolds. Chunking strategies, reranking heuristics, prompt templates that coax structured output, hand-written schema descriptions - these are compensations for current deficiencies, and they will be absorbed. Anyone whose position depends on them has a depreciating asset.
It is wrong to generalize from those scaffolds to the whole category, for a reason that turns on what kind of knowledge is being encoded. Sutton’s warning concerns knowledge about how to solve the problem - the human’s theory of which features matter, which openings are strong, which phonemes compose which words. That knowledge is lost because the data contains a better theory and the system can find it.
An ontology encodes something else: knowledge about the state of a particular world. That this company has these subsidiaries, that this contract has that indemnity clause, that this part number supersedes that one as of last Tuesday. This information was never in the pretraining distribution, cannot be derived from general capability, and changes weekly. No amount of raw intelligence supplies a fact about your organization that your organization has never written down in a machine-readable form.
The bitter lesson says nothing about the second category, because it is not a competing theory of the domain. It is the domain. A model that reasons superhumanly about supply chains still cannot tell you which of your suppliers is affected, and the gap is not a reasoning gap.
The honest version of the ontology position, then, is narrower than its usual marketing: the value migrates. Schema description is being absorbed and will keep being absorbed. Relational assertion, action typing, and verification are not, because they are not descriptions of how to think - they are the model’s only channel to a specific world, and the only mechanism by which its outputs are permitted to touch that world.
The Harness Decides What Counts
Which brings the argument to its second half, and to one of the defining ideas of agentic AI: the harness is becoming part of the model.
A lot of agent training happens in a clean environment. Then the model is dropped into Codex, Claude Code, OpenClaw, or a complicated internal system full of tools, prompts, state, retries, permissions, and strange failure modes, and is expected to perform. There is an obvious mismatch here, and OpenForgeRL - the source of the harness numbers in the opening table - is one of those papers that feels obvious the second you see it: train the agent through the harness it will actually use.
The reason nobody has is a tooling constraint, and the paper names it precisely: open SFT and RL stacks “cannot natively express stateful, multi-process harness inference.” Reinforcement learning infrastructure was built on the assumption of a clean, single-process environment that the trainer controls. A real harness is a stateful mess of subprocesses, tool servers, and permission boundaries, and existing stacks had no way to represent it. The fix is architectural rather than algorithmic: a lightweight proxy serves the harness’s model calls while recording them as training data for a standard RL codebase such as veRL, and a Kubernetes orchestrator runs each rollout in its own remote container. Training and inference are decoupled, so any harness can be paired with any environment without the RL stack needing to understand either.
This is a bottleneck-migration story like the one Part I describes. The constraint on harness-native training was never the learning algorithm. It was that no one could express the environment. And the payoff for removing it is steep: using only hundreds to a few thousand tasks, their models beat open baselines of similar size on nearly all benchmarks, and in the GUI setting match or surpass models several times larger.6 Training in the real environment is not merely more faithful. It is dramatically more sample-efficient, as you would expect if much of what an agent needs to learn is specific to the machinery it operates rather than the task domain.
This matters because the harness makes a series of consequential decisions on the model’s behalf. It determines what the model can see, what it can do, when it gets another attempt, how its work is judged, and whether a failure becomes useful evidence or disappears. Every one of those shapes behavior at least as much as a post-training run does, and none of them is plumbing.
Orchestration, verification, retries, candidate selection, tracing, failure recovery - these are not incidental to an agent system; they are its behavior, in any sense a user would recognize. The same model, in two different harnesses, is two different products.
A compounding loop is available here, and it is the real prize. Production runs generate traces. Traces expose failure modes. Verifiers turn those failures into signal. The system trains inside its actual operating environment. The product teaches the agent how to use the product. A better harness creates better evidence, better evidence creates a better agent, and a better agent creates better runs.
Only the harness owner can access this loop. That is a structural point about who captures value in agentic AI, and it is closer to the ontology moat than it first appears: both are assets that accumulate from operation rather than from procurement.
The Price of Compiling the Harness
Training through the harness moves it along an encoding gradient, and the encoding framework predicts a cost for that.
The Encoding Problem described knowledge traveling from ephemeral to persistent to compiled, with a consistent direction of travel toward compilation driven by economics - compiled behavior is free at inference time. The cost paid is rigidity. The harness sits at the persistent end of that gradient: separable from the model, revisable in minutes, swappable between deployments. Training through it moves it toward the compiled end, and the obvious worry is specialization. A model reinforcement-trained through one harness becomes a model fitted to that harness’s tool names, error formats, and retry semantics - engineering decisions that were arbitrary when they were made and are now entangled with the weights.
Yu et al. test transfer directly. A model trained only on ZeroClaw still improved on harnesses it had never seen during training - by 3.3 points on OpenClaw and 4.6 on Codex over the untrained base, modest gains, exactly what specialization would predict. Training on all three harnesses together did better everywhere, with the largest gains on the harder harnesses - 20.9 against 14.7 on OpenClaw, and 32.5 against 16.8 on Codex. The result that actually overturns the lock-in framing is narrower: the multi-harness model beat the ZeroClaw-only model on ZeroClaw itself, 48.5 against 46.0.7 Training on other harnesses improved performance in the harness it would actually be deployed in.
Harness diversity behaves like data augmentation. What the model acquires from a spread of harnesses is not idiom but something closer to harness-invariant competence - the general shape of operating machinery under uncertainty - and that generalizes back to any particular machine.
So the strategic conclusion inverts. The risk was never compiling a harness; it is compiling exactly one. Part IV’s advice to compile skeptically becomes something more specific here: compile across a portfolio. Train through several harnesses, including ones you do not ship, and you get deployment realism without buying the rigidity.
The moat survives in weakened form. Most of the gain stays home: zero-shot transfer is worth 3.3 and 4.6 points against the 13.5 points training buys in the harness itself, so whoever owns the harness still captures the larger share of what the compounding loop produces - a genuine structural advantage, but not an exclusive one. A competitor training across a portfolio of harnesses closes most of the gap without owning any of them - the multi-harness model gains 20.3 points on Codex and 9.5 on OpenClaw over training on those harnesses singly.
Why Error Recovery Lags
Reinforcement learning through the harness bought capabilities unevenly. Error recovery improved - Yu et al. define it as the share of rollouts that still solve the task given that at least one command has already failed - and reinforcement learning roughly doubled it off a low base. But every other capability they track ends up between 60 and 86, while error recovery ends up around 50. It improved, and it is still the floor. Their own reading is that these abilities are “difficult to acquire from RL alone, and may require dedicated data or training methods.”8
Self-verification, tool coverage, and step efficiency are in-distribution competencies. They describe executing well along paths the environment routinely produces, and a training loop that samples from the harness’s own behavior will be dense in exactly those situations. Reinforcement learning is very good at improving what it sees often. Note that self-verification, defined as reading back a write to confirm it, is a habit the model can practice on every single successful trajectory.
Error recovery is defined by its precondition: something has already gone wrong. It is the competence required when the environment leaves its own distribution - a tool returning something structurally unlike its schema, a state inconsistent in a way nobody anticipated, a partial write leaving the world in a configuration no design document describes. A loop that samples naturally from production undersamples precisely these cases, because they are rare by construction. The rarer and stranger the failure, the more valuable recovery is and the less training signal exists for it.
If that diagnosis is right, scaling the same loop will not fix it, and the paper’s hypothesis that dedicated methods are needed has a concrete instantiation: perturb the environment deliberately. Fault injection, adversarial harness behavior, synthetic corruption of tool responses and state - manufacture off-distribution situations densely enough to learn from. The harness has to be engineered to break itself on purpose, in training, in the ways production breaks by accident.
This is testable: if agentic error recovery closes the gap over the next generation through scaled natural-distribution rollouts alone, the explanation above is wrong.
What This Adds to the Framework
Part I proposed three ways knowledge enters a general problem solver: structural embedding, data embedding, data encoding. The third category, like the second before it, does too much work: the consequential distinction inside it is between two things that look alike but behave nothing alike.
Encoded procedure is a human theory of how to solve the problem, expressed at inference time. It reproduces the bitter lesson’s failure mode with a delay, and it deserves the suspicion Part IV directs at agent frameworks. Empirically, it also explains why the elaborate harnesses were expensive and hard to learn.
Encoded structure declares what exists, what may be done, and what counts as correct. It constrains the interface without constraining the inference. It does not compete with the model’s capability; it is the channel through which that capability reaches a particular world.
Sutton’s framework has no category for the second, because in 2019 no general capability sat idle for want of access. There is now. The ceiling keeps rising, and whether anything is done with it is a separate engineering discipline that is currently the cheaper lever.
- Per-harness performance becomes a reported quantity. The variance is no longer in question - one base model spanning 11.4 to 32.5 on the same benchmark settles that. The prediction is about disclosure: model releases begin publishing results conditioned on harness, because a number reported without its harness is close to meaningless. Falsified if headline benchmark numbers continue to be published harness-free and nobody minds.
- Ontology value migrates rather than evaporating. Schema description gets absorbed into general capability; relational assertion, action typing, and verification do not. Falsified if a naive-retrieval system with a long context matches a structured one on multi-hop relational tasks at comparable cost.
- Error recovery lags until environments are made adversarial in training and is falsified if it closes the gap through scaled natural-distribution rollouts alone.
OpenForgeRL is an early blueprint for where serious agent engineering goes, and the interesting part is not the proxy or the orchestrator. It is the admission underneath them - that the environment is not a detail surrounding the model, but a component of it. Once you concede that, the question of how good the models are stops being answerable on its own. You can only ask how good a model is at something.
Related
- Beyond the Bitter Lesson - Part 1 - A Deeper Understanding of AI Innovation
- Beyond the Bitter Lesson - The Encoding Problem (Part IV)
- Beyond the Bitter Lesson - Who Wins and Why (Part VI)
- Attention, Hierarchy, and Traversal
- Communication as Graph Traversal
- What is Scaling?
Footnotes
-
Xiao Yu, Baolin Peng, Ruize Xu, Hao Zou, Qianhui Wu, Hao Cheng, Wenlin Yao, Nikhil Singh, Zhou Yu, Jianfeng Gao. OpenForgeRL: Train Harness-native Agents in Any Environment. Columbia University, Dartmouth College, Microsoft Research. arXiv:2607.21557v3 [cs.AI], submitted 23 July 2026, v3 7 August 2026. Harness figures are in Table 4: the untrained Qwen3-30B-A3B-Thinking base model on ClawEval, pass@1 and pass@3. ↩
-
Yu et al., §5. ZeroClaw is their minimal reference harness; ReACT is a standard tool-calling loop. The quoted clause is theirs. ↩
-
Yu et al., §4. Neither OpenClaw nor Codex exposes a custom-tool interface, so ClawEval’s tools reach them as skill files - a wrapping cost paid in context on every rollout. ↩
-
Palantir, https://www.palantir.com/docs/foundry/ontology/overviewOntology overview and https://www.palantir.com/docs/foundry/action-types/overviewAction types overview. Object types, property types, link types, and action types are Palantir’s own four-part decomposition; action types carry authorization rules, submission criteria, rules, and side effects, and commit through an object type’s writeback dataset. ↩
-
Foundry dates to roughly 2016; the Ontology was consolidated as a first-class product around 2021-22; AIP launched April 2023, four months after ChatGPT. Palantir’s own framing of the ontology-as-grounding-layer argument is in https://blog.palantir.com/reducing-hallucinations-with-the-ontology-in-palantir-aip-288552477383Reducing Hallucinations with the Ontology in Palantir AIP. ↩
-
Yu et al., abstract: OpenForgeClaw reaches ClawEval 31.7 (pass@1) and 55.9 (pass@3), QwenClawBench 33.7, MCPAtlas 28.1; OpenForgeGUI, at 8B, reaches OSWorld-Verified 37.7, Online-Mind2Web 63.0, WebVoyager 72.3. ↩
-
Yu et al., Table 5. Trained on ZeroClaw alone: +3.3 on OpenClaw, +4.6 on Codex over the untrained base. Trained on all three: ZeroClaw 48.5 (against 46.0 single-harness), OpenClaw 20.9 (against 14.7), Codex 32.5 (against 16.8). ↩
-
Yu et al., §5.3. Error recovery is the share of rollouts that still solve the task conditional on at least one command having already failed. It roughly doubles under RL and still lands lowest among the capabilities they track. ↩