Perverse Incentives

A perverse incentive is a reward structure that produces the opposite of what it was meant to produce. The incentive is set up to promote some goal; the behaviour it actually rewards undermines that goal. What makes it perverse rather than merely misaligned is the inversion: the mechanism does not just fail to help, it actively pays people to make things worse, and the harm is a direct consequence of people responding rationally to the reward on offer.

The structure is always the same. A principal wants outcome X but cannot observe or reward X directly, so they reward a proxy for X. Agents then optimise the proxy, and wherever the proxy and the goal come apart, the agents follow the proxy - because that is what gets paid. The gap between proxy and goal is where the perversity lives.

Canonical Examples

  • The Hanoi rat bounty (1902) - the French colonial administration paid a bounty per rat tail to reduce the rat population. Rat-catchers began cutting off tails and releasing the rats to breed, some started farming rats for the bounty, and others smuggled rats in from the countryside. The population rose. (The often-cited “cobra effect” - a British bounty on cobras in Delhi producing cobra farms - has the same shape but is poorly sourced; the Hanoi case is documented.)
  • Fee-for-service medicine - paying providers per procedure rewards volume of treatment, not health, so it over-treats.
  • Sales quotas without quality checks - Wells Fargo’s cross-selling targets (exposed 2016) led staff to open some 3.5 million unauthorised accounts, because the reward was for accounts opened, not for customers served.
  • Publish-or-perish - rewarding paper count and citation metrics produces salami-sliced papers, p-hacking, and citation rings. The proxy (publication) diverges from the goal (knowledge).
  • Paying per bug fixed - the canonical thought experiment rather than a documented case (it is a 1995 Dilbert strip: “I’m gonna write me a new minivan”). Rewarding bugs fixed rewards the existence of bugs, so there is no reason to write code that lacks them.

Each is an instance of Goodhart’s Law - in Marilyn Strathern’s phrasing, “when a measure becomes a target, it ceases to be a good measure.” (Goodhart’s own 1975 version was narrower: “any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes.“) The moment the proxy is what gets paid, it is optimised for its own sake and stops tracking the thing it was standing in for.

Relationship to Neighbouring Concepts

  • Externality - a perverse incentive frequently works by externalisation: the agent captures the reward and pushes the cost onto someone the reward scheme does not account for. The rat-catcher gets paid; the city gets the rats.
  • Tragedy of the Commons - the commons problem is a perverse incentive at group scale. Private benefit, socialised cost is precisely the proxy/goal gap: each actor is rewarded for extraction, the group is punished for it. What distinguishes the commons is that no one designed the incentive - it emerges from the structure of shared ownership. Perverse incentives in the narrower sense are usually somebody’s policy.
  • Collective Action Problems - the broader family in which individually rational action produces collectively bad outcomes. Perverse incentives overlap it but are not contained by it: a single agent responding to a single bounty is a perverse incentive with no collective action problem in sight.
  • Moral hazard - a sibling: insulating someone from the downside of their choices (insurance, bailouts, limited liability) rewards risk-taking. Same mechanism, specific to the removal of consequences.
  • Principal-agent problem - the general frame. Perverse incentives are what the principal-agent problem looks like when the contract has been written badly.

Talking Your Book

One form deserves its own heading because it is easy to miss. When the person making a forecast profits from the forecast being believed, regardless of whether it comes true, the reward has been decoupled from accuracy. The barber’s opinion on whether you need a haircut, the fund manager’s view on the asset they hold, the founder’s prediction about the industry they are building - each is paid for persuasion, not correctness. This is not the same as lying; the speaker may believe every word. The incentive simply removes any pressure that would correct them if they were wrong, and adds pressure in the direction of whatever sells.

The test is to ask what the speaker gains if the audience agrees, and whether that gain depends on the claim being true. If the gain arrives either way, the claim is advertising wearing the costume of analysis.

The Way Out

The remedies map onto the diagnosis - close the gap between proxy and goal, or make the agent bear the cost of exploiting it:

  • Reward outcomes, not activity - pay for health, not procedures; for working software, not bugs fixed. Harder to measure, which is why the proxy was chosen in the first place.
  • Skin in the game - make the agent bear the downside. Clawbacks, deferred compensation, personal liability. The stake-creates-alignment argument is the general form: badly designed stake produces perverse incentives, well-designed stake dissolves them.
  • Multiple proxies - a single metric is easy to game; a basket of partially independent metrics is harder, because gaming one usually damages another.
  • Detection and sanction - Ostrom’s route for the commons applies here too. Make exploitation of the proxy visible and costly rather than hoping nobody notices the gap.

The bound on all of these is that no incentive scheme is proxy-free. Every reward is a stand-in for something that cannot be measured directly, so Goodhart eventually bites everywhere. The practical stance is the one from collective action: design for agents who will exploit the gap rather than agents who will not, and put the effort into detection rather than into hoping the proxy holds.