Four Ways to Manufacture a Coordination Finding
A Control Battery for Mixed-Motive MARL
Abstract
Mixed-motive multi-agent reinforcement-learning sweeps can manufacture findings before analysis. In a four-agent shared-state task, four choices make reward scale look structurally dominant, manufacture a symmetric treatment profile, turn state contention into a general correlation effect, and misattribute attainable-objective costs to learning. The treatment is signed correlation among target coordinates (ideal points), not a manipulation of objective- or reward-function alignment. The battery finds that the original sampler never produced negative target correlation and that stake dominance was degree-one reward scaling. Separable IQL is correlation-invariant and the VDN slope was not detected; on legacy-Gaussian support, slope estimates steepened across four fixed first-k shared-coordinate configurations. Frozen held-out evaluation preserved the target-correlation gradient. Translating the distribution inside the feasible box made it three to four times steeper. An exact finite-horizon oracle attributed 54–63% (95% bootstrap intervals 51–72%) of that gradient to feasibility and 37–46% (28–49%) to a support-dependent frozen-policy residual. The full proposed index F = σ(1+ε)/(1+α) has no valid confirmatory verdict: on a post-diagnostic ε=0 signed grid, its reduced restriction loses to a degree-one scale-homogeneous alternative by AICc, and the observation-noise-variance proxy interacts with signed target correlation in the opposite direction. These post-diagnostic controls are exploratory demonstrations. The battery separates treatment validity, algebraic scale, environmental scope and benchmark validity before judging learned policies.