What adversarial systems research is
Adversarial systems research is the study of what coordination costs, and of what happens to a system when that cost never clears. The object is not a system in conflict with an outside opponent but a system in conflict with itself: any arrangement in which one agent acts on behalf of another generates resistance, and that resistance is present whether the agents are electorates and legislatures, traders and market infrastructure, or operators and the learned policies they deploy.
The working assumption of the programme is that this resistance — friction — is measurable, that it decomposes into a small and fixed set of components, and that the components are the same across substrates even when the measurements are not.
This is an early-stage independent initiative. Most outputs are preprints or under review, the empirical base is thin in several places, and the sharpest quantitative claim the framework has made has already been rejected by the test built to check it, with the surviving parts surviving in weaker form. The five programmes exist to keep generating independent chances to be wrong, and the honest description of the present state is that the variables look right and the composition does not.
Central argument
The three quantities are meant to be estimated rather than assumed. Alignment is the degree to which the delegate's objectives track the delegator's. Stake is the magnitude of what is exposed to the decision. Entropy is the information lost in transmitting the delegation. The directional relations are the load-bearing part: friction rises with stake and with entropy, falls with alignment, and does not reach zero even under perfect alignment, because delegation itself carries an irreducible cost. This is a thesis, not a mission, and it is stated so that it can be shown to be false.
What has already failed. The composite friction functional F = σ(1 + ε)/(1 + α) was always labelled an ansatz — the simplest expression satisfying the stated desiderata, not a derived result — and a factorial multi-agent reinforcement learning study built to test it rejected it as a single-index predictor, with an independent-effects model winning on comparative grounds. The kernel variables and their directional roles survived the test; the particular way of combining them did not. The same study then found that one of its own apparently confirming results — a symmetric U-shape in alignment, with neutral preferences uniquely worst — was an artefact of how the experiment sampled agent preferences, and withdrew it.
Both outcomes are stated in the framework papers themselves rather than buried in supplementary material, because a programme that publishes only its surviving claims cannot be assessed from outside. Anyone weighing the framework should read it as an apparatus whose variables have held up and whose composition has not.
The framework papers are The Axiom of Consent: Friction Dynamics in Multi-Agent Coordination (arXiv:2601.06692) and The Replicator-Optimization Mechanism (arXiv:2601.06363), both preprints.
What would refute it
- Sufficiency fails. A domain is found where a fourth quantity, not reducible to α, σ or ε, carries independent predictive weight over coordination failure. The triple is then incomplete, and the framework is at best a partial description.
- Direction fails. A domain is found where, with scale properly normalised, higher stake reliably lowers friction or better alignment reliably raises it. The comparative statics are the load-bearing part; if they invert, nothing survives underneath them.
- Transport fails. Relations estimated in one substrate do not carry to another even after the variables are re-measured in the new substrate's own terms. The programme would then be a collection of unrelated domain models sharing a notation, which is a worse thing to be than a wrong theory.
In scope, and not
- Measuring a delegation. Given a particular delegation — a shareholder body and a board, an electorate and a legislature, an operator and a deployed policy — can α, σ and ε be estimated from observable data rather than assumed?
- Stake–voice mismatch. When decision authority and consequence-bearing come apart, how much friction does the gap generate, and does the size of the gap track the size of the eventual disruption?
- Thresholds and reconfiguration. At what point does accumulated friction stop being absorbed and force reconfiguration — exit, protest, a protocol fork, a retraining run — and is that point predictable in advance rather than only narratable afterwards?
- Transport across substrates. Does a friction relation estimated in one substrate hold in another once the variables are re-measured in the second substrate's own terms, and where it fails, which of the three failed to transport?
- Friction signatures in prices. Do markets respond differently to shocks arriving through a system's own governance than to shocks imposed on it from outside, and does that difference survive inference that accounts for cross-sectional and temporal dependence?
- Who counts as a party. Which entities hold enough stake in a decision to be parties to it rather than objects of it, and what follows for artificial agents that bear consequences without holding voice?
- Manufactured findings. Which apparently structural coordination results are artefacts of how the study was built — reward scaling, the geometry of a preference sampler, a shared-resource environment — rather than properties of the agents in it?
- Not a theory of what anyone ought to want — the framework predicts which arrangements cost less to run and therefore tend to persist, and the bridge to normative conclusions is conditional — if less coordination failure is wanted, then certain configurations are instrumentally preferred — and is presented at that strength and no higher
- Not policy advice — once the output is a recommendation rather than an estimate, there is an interest in how the estimate comes out, and the measurement is worth less
- Not the study of conflict between systems — war, competition and predation between rivals have their own literatures; the object here is the conflict a system generates against its own coordination requirements, which is what "adversarial" is doing in the name
- Not mechanism design or equilibrium refinement — the programme sits upstream of the point where payoffs are assigned and agents are seated, so a better solution concept is not on offer and competing on that ground would be a category error
- Not phenomenal consciousness as a criterion — where the programme touches minds the questions are functional and relational — what a system does, and what it is exposed to — and standing is grounded in stake rather than in whether there is something it is like to be the system
- Not signal generation — financial data is used as a testbed with unusually good measurement, not as a source of tradable edge, and where an apparent edge does not survive its own robustness checks the reported finding is the null
What counts as evidence
The methodological commitment is formal generality with empirical accountability, which in practice means a short list of rules applied before a result is allowed to leave the building. The first is that the headline N counts independent units. Runs, observations and events are not independent by default; the number that goes in the abstract is the number of independent units, and cluster-robust or dependence-robust inference is the default rather than a robustness appendix. A number that would not survive clustering is not reported as a finding.
The second is that every number traces to committed code. Each reported statistic and each figure corresponds to a script at a commit in a public repository, and where the text and the code disagree, the code is what happened and the text is corrected. Every reference is resolved against Crossref or an equivalent index before it ships; a reference that cannot be resolved is removed rather than softened.
Adequacy bars are set before the test. The multi-agent study stated its threshold in advance and the framework's own functional form failed against it. A bar chosen after the results are in makes the test unfalsifiable, whatever else it does. Controls are built to catch our own artefacts rather than other people's: the most productive result the programme has produced so far was the discovery that a confirming finding was an artefact of the experimental design, caught by a control battery run against work already written up as a success. Formal verification certifies coherence, not adequacy — the core comparative statics are machine-checked in Lean 4, which establishes that the mathematics is internally consistent and nothing about whether the world behaves that way, and the two are kept apart in the write-up.
Nulls, corrections and withdrawals are published. A rejected functional form, a null event-study result and a withdrawn effect are all outputs, and they are the parts of the record that make the surviving parts worth anything. Everything is open access, CC BY 4.0, with data and code available: research on coordination costs should not impose them on its readers.
How this differs from adjacent fields
Adversarial systems research borrows heavily. Almost every component of the apparatus has a better-developed ancestor somewhere, and a reader who knows one of these fields well is right to ask whether anything here is more than a renaming — so this answers field by field, including the cases where the honest answer is that we are downstream. Three commitments recur: a friction-first inversion, in which friction is the observable primitive and consent a derived description of the configurations that produce low friction; persistence-conditioning, in which coordination cost enters a survival functional so that selection runs over configurations rather than optimisation within one; and cross-substrate scope, the claim that the same kernel triple of alignment, stakes and entropy holds for political, market and computational delegation. The third is a liability at least as much as a contribution, and several of the entries below are about the ways it has already cost us.
Complexity science and the Santa Fe tradition
Multi-agent reinforcement learning
Cooperative AI
AI alignment and safety
Institutional and constitutional political economy
Cybernetics and systems theory
Econophysics and systemic-risk network modelling
Argue with this
The claim above is meant to be attackable. If you think the decomposition is wrong, the framework is redundant with something that already exists, or a result does not hold, we publish substantive critiques alongside the position they attack.