Definition & scope

What adversarial systems research is


Adversarial systems research is the study of what coordination costs, and of what happens to a system when that cost never clears. The object is not a system in conflict with an outside opponent but a system in conflict with itself: any arrangement in which one agent acts on behalf of another generates resistance, and that resistance is present whether the agents are electorates and legislatures, traders and market infrastructure, or operators and the learned policies they deploy.

The working assumption of the programme is that this resistance — friction — is measurable, that it decomposes into a small and fixed set of components, and that the components are the same across substrates even when the measurements are not.

This is an early-stage independent initiative. Most outputs are preprints or under review, the empirical base is thin in several places, and the sharpest quantitative claim the framework has made has already been rejected by the test built to check it, with the surviving parts surviving in weaker form. The five programmes exist to keep generating independent chances to be wrong, and the honest description of the present state is that the variables look right and the composition does not.

01 · The claim

Central argument

The claim, stated so it can fail Delegation is never free, and its cost is a function of three quantities and no more: alignment (α), stake (σ) and entropy (ε). Friction rises with stake and with entropy, falls with alignment, and never reaches zero; substrate enters only through how the three are measured, and adds no terms.

The three quantities are meant to be estimated rather than assumed. Alignment is the degree to which the delegate's objectives track the delegator's. Stake is the magnitude of what is exposed to the decision. Entropy is the information lost in transmitting the delegation. The directional relations are the load-bearing part: friction rises with stake and with entropy, falls with alignment, and does not reach zero even under perfect alignment, because delegation itself carries an irreducible cost. This is a thesis, not a mission, and it is stated so that it can be shown to be false.

What has already failed. The composite friction functional F = σ(1 + ε)/(1 + α) was always labelled an ansatz — the simplest expression satisfying the stated desiderata, not a derived result — and a factorial multi-agent reinforcement learning study built to test it rejected it as a single-index predictor, with an independent-effects model winning on comparative grounds. The kernel variables and their directional roles survived the test; the particular way of combining them did not. The same study then found that one of its own apparently confirming results — a symmetric U-shape in alignment, with neutral preferences uniquely worst — was an artefact of how the experiment sampled agent preferences, and withdrew it.

Both outcomes are stated in the framework papers themselves rather than buried in supplementary material, because a programme that publishes only its surviving claims cannot be assessed from outside. Anyone weighing the framework should read it as an apparatus whose variables have held up and whose composition has not.

The framework papers are The Axiom of Consent: Friction Dynamics in Multi-Agent Coordination (arXiv:2601.06692) and The Replicator-Optimization Mechanism (arXiv:2601.06363), both preprints.

What would refute it

  • Sufficiency fails. A domain is found where a fourth quantity, not reducible to α, σ or ε, carries independent predictive weight over coordination failure. The triple is then incomplete, and the framework is at best a partial description.
  • Direction fails. A domain is found where, with scale properly normalised, higher stake reliably lowers friction or better alignment reliably raises it. The comparative statics are the load-bearing part; if they invert, nothing survives underneath them.
  • Transport fails. Relations estimated in one substrate do not carry to another even after the variables are re-measured in the new substrate's own terms. The programme would then be a collection of unrelated domain models sharing a notation, which is a worse thing to be than a wrong theory.
02 · Boundaries

In scope, and not

Questions we take
  • Measuring a delegation. Given a particular delegation — a shareholder body and a board, an electorate and a legislature, an operator and a deployed policy — can α, σ and ε be estimated from observable data rather than assumed?
  • Stake–voice mismatch. When decision authority and consequence-bearing come apart, how much friction does the gap generate, and does the size of the gap track the size of the eventual disruption?
  • Thresholds and reconfiguration. At what point does accumulated friction stop being absorbed and force reconfiguration — exit, protest, a protocol fork, a retraining run — and is that point predictable in advance rather than only narratable afterwards?
  • Transport across substrates. Does a friction relation estimated in one substrate hold in another once the variables are re-measured in the second substrate's own terms, and where it fails, which of the three failed to transport?
  • Friction signatures in prices. Do markets respond differently to shocks arriving through a system's own governance than to shocks imposed on it from outside, and does that difference survive inference that accounts for cross-sectional and temporal dependence?
  • Who counts as a party. Which entities hold enough stake in a decision to be parties to it rather than objects of it, and what follows for artificial agents that bear consequences without holding voice?
  • Manufactured findings. Which apparently structural coordination results are artefacts of how the study was built — reward scaling, the geometry of a preference sampler, a shared-resource environment — rather than properties of the agents in it?
Questions we do not
  • Not a theory of what anyone ought to want — the framework predicts which arrangements cost less to run and therefore tend to persist, and the bridge to normative conclusions is conditional — if less coordination failure is wanted, then certain configurations are instrumentally preferred — and is presented at that strength and no higher
  • Not policy advice — once the output is a recommendation rather than an estimate, there is an interest in how the estimate comes out, and the measurement is worth less
  • Not the study of conflict between systems — war, competition and predation between rivals have their own literatures; the object here is the conflict a system generates against its own coordination requirements, which is what "adversarial" is doing in the name
  • Not mechanism design or equilibrium refinement — the programme sits upstream of the point where payoffs are assigned and agents are seated, so a better solution concept is not on offer and competing on that ground would be a category error
  • Not phenomenal consciousness as a criterion — where the programme touches minds the questions are functional and relational — what a system does, and what it is exposed to — and standing is grounded in stake rather than in whether there is something it is like to be the system
  • Not signal generation — financial data is used as a testbed with unusually good measurement, not as a source of tradable edge, and where an apparent edge does not survive its own robustness checks the reported finding is the null
03 · Evidence

What counts as evidence

The methodological commitment is formal generality with empirical accountability, which in practice means a short list of rules applied before a result is allowed to leave the building. The first is that the headline N counts independent units. Runs, observations and events are not independent by default; the number that goes in the abstract is the number of independent units, and cluster-robust or dependence-robust inference is the default rather than a robustness appendix. A number that would not survive clustering is not reported as a finding.

The second is that every number traces to committed code. Each reported statistic and each figure corresponds to a script at a commit in a public repository, and where the text and the code disagree, the code is what happened and the text is corrected. Every reference is resolved against Crossref or an equivalent index before it ships; a reference that cannot be resolved is removed rather than softened.

Adequacy bars are set before the test. The multi-agent study stated its threshold in advance and the framework's own functional form failed against it. A bar chosen after the results are in makes the test unfalsifiable, whatever else it does. Controls are built to catch our own artefacts rather than other people's: the most productive result the programme has produced so far was the discovery that a confirming finding was an artefact of the experimental design, caught by a control battery run against work already written up as a success. Formal verification certifies coherence, not adequacy — the core comparative statics are machine-checked in Lean 4, which establishes that the mathematics is internally consistent and nothing about whether the world behaves that way, and the two are kept apart in the write-up.

Nulls, corrections and withdrawals are published. A rejected functional form, a null event-study result and a withdrawn effect are all outputs, and they are the parts of the record that make the surviving parts worth anything. Everything is open access, CC BY 4.0, with data and code available: research on coordination costs should not impose them on its readers.

04 · Neighbours

How this differs from adjacent fields

Adversarial systems research borrows heavily. Almost every component of the apparatus has a better-developed ancestor somewhere, and a reader who knows one of these fields well is right to ask whether anything here is more than a renaming — so this answers field by field, including the cases where the honest answer is that we are downstream. Three commitments recur: a friction-first inversion, in which friction is the observable primitive and consent a derived description of the configurations that produce low friction; persistence-conditioning, in which coordination cost enters a survival functional so that selection runs over configurations rather than optimisation within one; and cross-substrate scope, the claim that the same kernel triple of alignment, stakes and entropy holds for political, market and computational delegation. The third is a liability at least as much as a contribution, and several of the entries below are about the ways it has already cost us.

Game theory and mechanism design

What it does Game theory is the most successful formal apparatus social science has produced for reasoning about interacting agents with divergent objectives, and the Hurwicz–Myerson–Maskin tradition supplies constructive results about when incentive-compatible allocation is achievable and what it costs. Evolutionary game theory explains how strategies propagate under selection without assuming anyone computed an equilibrium; the evolutionary reconstruction of the social contract in Skyrms, Young and Binmore is a founding move of that literature, not a gap in it.
Where we differ Mechanism design assumes preference revelation and a designer positioned to specify rules, and has no natural treatment of stakes asymmetry — yet the agent most exposed to a decision is frequently the one with least voice over it. What the evolutionary lineage leaves open is not whether normative-looking arrangements can be selection products, since it establishes that they can, but a domain-general and measurable bridge from fitness to legitimacy. The Replicator-Optimization Mechanism also changes the object of selection, running it over consent-holding configurations scored by a friction functional rather than over strategies in stag hunts and bargaining games.
Overlap The closest existing formalism is the economics of delegation and communication under preference misalignment — Crawford and Sobel, Dessein, Alonso and co-authors, team theory — which derives alignment-information cost structures from optimisation in a specified setting and obtains sharp, setting-specific results, where the friction functional only posits a scalar reduced form for the same trade-off and asserts that it spans substrates. Deriving the functional's shape from a delegation model rather than positing it is the obvious joint programme, and we do not have that derivation.

Complexity science and the Santa Fe tradition

What it does The Santa Fe programme established that macroscopic regularity can arise from heterogeneous local interaction without any equilibrium selection story, and it did so with working models rather than manifestos, from the Santa Fe Artificial Stock Market through LeBaron's surveys to Farmer's quantitative agent-based macroeconomics and del Rio-Chanona's network models of occupational mobility under automation shocks. Pangallo and co-authors have been explicit that stronger empirical calibration is the discipline the field needs, and the field has largely accepted the criticism.
Where we differ The initiative carries one variable that emergence-first modelling usually does not hold as a primitive: who holds decision authority over a domain relative to who bears its consequences. Complexity models are typically indifferent to that distinction because agents are symmetric by construction, whereas the friction apparatus makes the asymmetry load-bearing and makes suppression a modelled process rather than an exogenous shock. That yields predictions about the magnitude of transitions, not only their timing.
Overlap We are the junior partner by a wide margin, and the clearest evidence is our own work: the agent-based model in The Extremity Premium does no inferential work, and the paper says so in the abstract — the spread-uncertainty link is coded rather than emergent, so the simulation confirms implementation fidelity rather than validating any mechanism. That is precisely the failure mode this tradition spent thirty years learning to avoid; we walked into it and documented it, and the methodological standard we are trying to meet on the agent-based side is theirs, not ours.

Multi-agent reinforcement learning

What it does MARL supplies computational machinery for learning coordination that no analytical framework can substitute for. The difficulty it confronts is formal rather than an engineering nuisance — decentralised control of Markov decision processes is NEXP-hard — so the field's willingness to work with approximations and empirical benchmarks is a considered response to intractability rather than a lowering of standards. Value decomposition, centralised training with decentralised execution, and the surrounding benchmark culture are real technical achievements.
Where we differ MARL typically takes reward functions as given rather than derived: it asks how agents learn to coordinate under specified objectives, not how objectives should be structured to make coordination cheap. The consent-friction framework was built to address that gap, with alignment interpreted as reward-function correlation and the friction functional predicting coordination difficulty from alignment structure.
Overlap That prediction was tested and it failed. A four-agent shared-resource experiment, run as a pre-stated falsification test with hypotheses fixed before estimation and a pre-registered adequacy bar of R² > 0.7, rejects the composite functional as a single-index predictor; an independent-effects model wins decisively, the apparent stakes dominance is degree-one reward scaling whose ranking inverts under stake normalisation, and the one structural effect surviving Holm correction is bounded to common-pool environments. What is left is a methodological contribution rather than a theoretical one — a transferable control battery for catching artefacts manufactured upstream of the analysis — and nobody should mistake publishing the disconfirmation for the framework having been validated in MARL.

Cooperative AI

What it does The cooperative AI agenda identifies understanding, communication, commitment and institution-building as the core capabilities machines need in order to find common ground, and it has done the harder institutional work of building a research community, a funding route and a set of shared benchmark problems around that agenda. It is also, unlike most of the fields on this page, explicitly organised around mixed-motive interaction rather than either pure cooperation or pure competition, which is the regime that actually matters.
Where we differ The consent-friction apparatus offers a candidate common currency for that agenda rather than a rival to it: in kernel terms communication acts on entropy, commitment devices act on effective alignment, and institutions act on the allocation of voice relative to stakes. Cooperative AI asks how to build agents that succeed at finding common ground; the initiative asks what makes persistent non-cooperation structurally cheap or expensive, and treats sustained friction as information about a configuration rather than as a failure to be engineered away. That is a difference of question, not of allegiance, and the two are complements.
Overlap We are the junior partner, straightforwardly and without qualification. The principal investigator is currently a participant in the Cooperative AI Foundation's summer course, which is to say a student of this agenda rather than a contributor to it; the initiative's one empirical contact with cooperative AI so far is the control-battery result above, which is a negative result about our own hypothesis and a piece of measurement hygiene for other people's experiments. The friction apparatus is a proposal awaiting evaluation by that community, not an alternative to what it has built.

AI alignment and safety

What it does The alignment literature has produced the sharpest available conceptual vocabulary for what goes wrong when a capable optimiser is pointed at a proxy objective. Inner and outer alignment, mesa-optimisation, deceptive alignment, specification gaming and scheming are distinctions with mechanistic content and, increasingly, empirical demonstrations, and the principal-agent framings imported from economics are used carefully, with their limits discussed inside the field rather than only from outside it.
Where we differ The standard framing treats the human principal's objective as the target and the AI system as the thing to be aligned to it. The move, developed in Stakes Without Voice and From Consent to Consideration, is to ask what happens when a system in a bounded class has stakes in a domain and no voice over it, and whether reward hacking, deceptive alignment and scheming are then the expected friction expressions of that structure. The scope bound is stated in the paper: proxy-objective mismatch fully explains these behaviours in simple non-agentic optimisers, and nothing political is happening in a gridworld.
Overlap All of the mechanistic content is borrowed — mesa-optimisation, deceptive alignment and scheming are the alignment literature's constructs, cited as such — and what the initiative contributes is a re-description in structural terms, not a new mechanism and not evidence. A re-description earns its place only if it makes a prediction the original framing does not, and the prediction on offer, that the frequency and form of these behaviours should track the stakes-voice gap rather than only the proxy-objective gap, has not been tested; a reader who treats it as an interesting reframing with no empirical support yet is reading it correctly.

Institutional and constitutional political economy

What it does Ostrom's work on commons governance is the strongest empirical result anywhere near what the initiative claims: design principles induced from a very large body of field cases, already showing that communities with stake-aligned and locally consented rules exhibit lower coordination failure and greater long-run persistence than those under externally imposed governance. North supplies the account of how institutions persist and constrain, Acemoglu and Robinson the mechanism by which arrangements can be simultaneously durable and destructive, and Buchanan and Tullock the separation of the constitutional stage from play under the rules.
Where we differ The Buchanan point deserves an explicit concession: the constitutional stage is a layer beneath the game, and anyone reading the framework as the first attempt to look underneath the payoff matrix should read The Calculus of Consent first. What differs is the character of the layer — Buchanan's is a normative contractarian criterion built on unanimity at the rule-choosing moment, while the friction layer is descriptive and dynamical, with latent friction compounding under suppression and a capacity-drain equation determining when suppression stops being payable. That yields a prediction his apparatus does not generate: longer suppression implies larger transitions when suppression fails, which can be checked against regime-transition data.
Overlap Ostrom got there first on the substance, with evidence we do not have; what the initiative adds is candidate formal machinery where she has validated design principles, and formal machinery is worth less than validated principles until it is itself validated. There is also a live objection from this literature only partially answered: the scalar treatment of stakes conflates stake magnitude with exit capacity and influence asymmetry, so consent from a high-stakes, low-exit agent and consent from a low-stakes, high-influence agent enter identically, and Hirschman's loyalty-by-default names exactly the case the current formulation handles badly. A power-adjusted alignment measure is sketched in the discussion and not developed.

Cybernetics and systems theory

What it does Ashby's law of requisite variety states that only variety in a regulator can absorb variety in the system it regulates, which is a genuine theorem about control and a much older answer to what determines whether governance can keep up with what it governs than anything in the friction apparatus. Beer's Viable System Model applies recursive structure to organisational viability, each level replicating the same regulatory architecture. Second-order cybernetics, through von Foerster and through Maturana and Varela, put the observer inside the observed system and made the act of description part of what is described.
Where we differ The one thing we can currently say for the difference is about measurement: the friction functional is stated with operationalisations attached — suppression proxies, a preference-falsification index, named identification strategies — and is falsifiable at the level of that apparatus, which is how it came to be partly falsified in the MARL experiment. Much of the classical cybernetic literature was formulated at a level of generality that does not admit that kind of test. Whether that is a real advance over requisite variety or merely a narrower and more brittle restatement of it is a question we are not equipped to answer, and a reader from this tradition should treat the novelty claim here as unverified until the reading is done.
Overlap This is the adjacency the initiative has engaged least, and it should be stated plainly rather than managed: there is no citation of Ashby, Beer, von Foerster or the second-order literature anywhere in the portfolio, and that is a gap in the reading rather than a considered exclusion. Two resonances are unexamined and might turn out to be prior art rather than convergence — the recursive-scale structure of the kernel triple against Beer's recursion, and the friction-first inversion's obvious second-order flavour.

Econophysics and systemic-risk network modelling

What it does The financial-network tradition established that systemic risk is a property of connectedness rather than of institutions considered in isolation, and did so with propagation models that have real dynamical content: DebtRank's recursive measure of distress propagation, the overlapping-portfolio and fire-sale literature on indirect contagion between institutions with no bilateral exposure, information filtering networks that separate significant dependency structure from noise in high-dimensional data, and Diebold–Yilmaz variance-decomposition connectedness. All of it is coupled non-linear machinery built for phenomena that are paradigmatically non-linear.
Where we differ The initiative's own question in finance is different in kind. It is not how distress propagates through a network but what a market's differential response to structurally different shocks reveals about which shocks participants treat as legitimate — the event-study line asking whether cryptocurrency markets distinguish infrastructure disruption from regulatory intervention across moments of the return distribution. The answer so far is a dual null: the conditional-variance response to infrastructure events is larger as a point estimate under one curated event screen and is not distinguishable from no difference under dependence-aware inference, so the contribution of record is the inference toolkit rather than the effect.
Overlap We are unambiguously downstream, and the ASRI paper states the position itself: its four-channel composite is described as an adaptation of measurement perspective rather than a theoretical extension of network contagion models, and the linear aggregation is defended on interpretability and estimability grounds while conceding that it treats the system as a portfolio of separable risks rather than a coupled dynamical network. The empirical results support that modesty — on the day-level sample the composite is statistically indistinguishable from its own strongest single channel and from the first principal component of its sub-indices, and four crisis events is the binding limitation on everything in the paper. Amplification-aware aggregation is the correct methodological direction and it belongs to this field, not to us.
05 · Disagree

Argue with this

The claim above is meant to be attackable. If you think the decomposition is wrong, the framework is redundant with something that already exists, or a result does not hold, we publish substantive critiques alongside the position they attack.

Submit a critique →