Cooperative Multi-Agent Reinforcement Learning via Epigraph-Form Guided Exploration

Sungil Son*  Hoseong Jung*  Dahyun Oh*  H. Jin Kim
Seoul National University
*Equal contribution
NeurIPS 2026

Exploration That Does Not Break Coordination

VMAS-balance: four agents carry a payload on a rod to the target

Both policies receive the same successor-distance intrinsic reward. The only difference is how exploration is combined with the task objective.

Linear reward mixing  r = rext + λ rint

Coordination collapse. The intrinsic bonus distorts the objective. Agents chase novelty, the rod tilts, and the payload is dropped.

EFXPLORER  min{ Aext, Vint − z }

Emergent coordination. Exploration is allowed only while the budget lasts and task progress is preserved, so the team explores briefly, then carries the payload to the target together.

TL;DR: EFXPLORER lets agents explore only as much as the team can afford. Instead of adding an intrinsic bonus to the task reward, it maximizes exploration subject to a constraint that task performance does not degrade, solves this in epigraph form with adaptive per-agent exploration budgets, and drives exploration with long-horizon successor-distance novelty. The result is new cooperative strategies on SMAX, VMAS, and MPE with no reward-mixing coefficient to tune.

More Demonstrations

SMAX 5m_vs_6m: five allied marines (blue) versus six enemy marines (green). Winning a numerically inferior fight requires a maneuver that plain reward maximization rarely discovers.

MAPPO MAPPO on SMAX 5m_vs_6m: the blue team engages head-on and is eliminated.

No exploration. The team engages head-on and is eliminated within a few steps.

EFXPLORER-MAPPO EFXPLORER on SMAX 5m_vs_6m: the blue team pulls back, stretches the pursuing enemy into a line, and defeats them.

Explored maneuver. The team pulls back to a corner, stretches the pursuing enemy into a line, and defeats them.

Why Not Just Add an Intrinsic Bonus?

Successful coordination occupies only a tiny subspace of the joint state-action space, so cooperative MARL leans on task-agnostic exploration signals. But a linearly mixed bonus can overshadow the task objective, and one agent's greedy exploration makes the environment non-stationary for its teammates, who mistake the temporary drop for failure and abandon the joint behavior.

Cooperative balancing task: (a) linear reward combination distorts objectives and collapses coordination; (b) EFXPLORER gates exploration with a min operator so coordination emerges.

Cooperative balancing task with four agents transporting a payload on a rod to a target. (a) Naive linear reward combinations induce objective misalignment, destabilizing coordination. (b) EFXPLORER leverages a min-function to conservatively gate exploration, enabling coordination to emerge progressively without sacrificing task performance.

Method

EFXPLORER maximizes each agent's intrinsic return subject to a non-degradation constraint on the joint extrinsic return, rewrites the problem in epigraph form, and solves it by alternating between a policy update against a stepwise min-bottleneck and an outer update of the exploration budgets.

Overview of EFXPLORER: (a) conservative exploration framework with exploration budgets and an epigraph min objective; (b) successor distance-based intrinsic reward from factorized per-agent episodic novelty.

Overview of EFXPLORER. (a) Conservative exploration: the epigraph objective uses a min operator to balance task progress and budget-adjusted exploration, while an outer loop adapts the budget \(z\). (b) Factorized episodic novelty: per-agent successor distance (SD) networks estimate temporal proximity to derive individual intrinsic rewards \(r^{\mathrm{int}}_{i,t}\).

Constrained exploration, in epigraph form

\[ \begin{aligned} &\max_{z}\; z := \sum_{i\in\mathcal{N}} z^i \\[0.35em] \text{s.t.}\quad &\max_{\boldsymbol{\pi}}\;\min\Big\{\, \color{#b91c1c}{J_{\mathrm{ext}}(\boldsymbol{\pi}) - J_{\mathrm{ext}}(\hat{\boldsymbol{\pi}})}\,,\; \color{#1d4ed8}{J_{i,\mathrm{int}}(\pi^i) - z^i} \Big\} \ge 0,\quad \forall i \in \mathcal{N} \end{aligned} \]
\[ \begin{aligned} &\max_{z}\; z := \sum_{i\in\mathcal{N}} z^i \\[0.35em] \text{s.t.}\;\; &\max_{\boldsymbol{\pi}}\;\min\Big\{\, \color{#b91c1c}{J_{\mathrm{ext}}(\boldsymbol{\pi}) - J_{\mathrm{ext}}(\hat{\boldsymbol{\pi}})}\,, \\[0.2em] &\qquad\qquad\;\; \color{#1d4ed8}{J_{i,\mathrm{int}}(\pi^i) - z^i} \Big\} \ge 0,\;\; \forall i \in \mathcal{N} \end{aligned} \]

The task-progress term and the intrinsic surplus over the budget \(z^i\) are coupled by a single min operator: exploration counts only when the team is not losing task performance relative to the reference policy \(\hat{\boldsymbol{\pi}}\). No trade-off coefficient is introduced.

Stepwise epigraph bottleneck

The episode-level constraint is turned into a per-step signal by augmenting the state with the remaining budget. Each agent's update is driven by the smaller of the reference task advantage and its intrinsic surplus, which gives a conservative sufficient condition for feasibility and a standard likelihood-ratio policy gradient.

\(\mathcal{B}^{\mathrm{epi}}_{i} = \min\{A^{\mathrm{ext}}_{\hat{\boldsymbol{\pi}}}(s_t,\mathbf{a}_t),\; S^{\mathrm{int}}_{i}(\tilde{s}^i_t,\mathbf{a}_t)\}\)

Adaptive exploration budgets

The outer loop grows the budgets \(z^i\) only from rollouts whose Monte Carlo task return beats a running reference. Among those feasible rollouts it picks the one with the largest collective intrinsic return and moves the budgets toward it with rate \(\alpha_z\). If nothing is feasible, the budgets stay put.

\(z^{i,k+1}_0 \leftarrow \alpha_z\, \hat{z}^{\,i,k,m^\ast}_0 + (1-\alpha_z)\, z^{i,k}_0\)

Long-horizon novelty via successor distance

Each agent learns a quasimetric successor distance on its own state space with a symmetric InfoNCE objective. The intrinsic reward is the minimum learned temporal distance to the states already visited in the episode, so agents are pushed toward temporally distant states rather than locally surprising ones.

\(r^{\mathrm{int}}_{i,t}(s^i_t) = \min_{k \in [0,t)} d_{\phi_i}(s^i_k, s^i_t)\)

Results

Comparison to Cooperative MARL Baselines

We evaluate on cooperative tasks from SMAX, VMAS, and MPE, including MPE-corridor, a new scenario that forces decentralized coordination in a narrow passage. All methods share the same backbone architecture and evaluation protocol. EFXPLORER improves normalized returns on both centralized-training (MAPPO) and independent-learner (IPPO) backbones. Baselines that inject intrinsic signals through social influence, entropy regularization, or enforced diversity (COIN, HASAC, DiCO) let those signals dominate learning, which destabilizes teammates and can collapse coordination.

Normalized average return curves on VMAS-balance, VMAS-give way, VMAS-transport, VMAS-wheel, MPE-corridor, MPE-tag, SMAX-3s5z_vs_3s6z, and SMAX-5m_vs_6m for EFXPLORER-MAPPO, EFXPLORER-IPPO, MAPPO, IPPO, HASAC, COIN, and DiCO.

Performance comparison of EFXPLORER with cooperative MARL baselines. Solid lines show the mean over 5 random seeds, while shaded regions correspond to the standard deviation.

3
benchmark suites (SMAX, VMAS, MPE)
8
cooperative tasks in the main comparison
2
backbones (MAPPO, IPPO)
5
seeds per curve

Epigraph Form vs. Linear Scalarization and Lagrangian Relaxation

The same constrained objective can be attacked with a scalarized reward \(J_{\mathrm{ext}} + \beta \sum_i J_{i,\mathrm{int}}\) (MAPPO-Linear) or with a primal-dual multiplier (MAPPO-Lagrangian). Both are brittle: on MPE-corridor, \(\beta = 0.1\) explores well while \(0.02\) or \(1\) degrade performance, and the Lagrangian variant needs a heavily weighted initialization (\(\lambda_0 = 1\)) to avoid coordination collapse. EFXPLORER stays competitive across all tasks with no such coefficient to tune.

Normalized average return on MPE-corridor, VMAS-balance, VMAS-give way, and SMAX-3s5z_vs_3s6z for EFXPLORER versus MAPPO-Linear with beta 0.02, 0.1, 1 and MAPPO-Lagrangian with lambda0 0.2 and 1.

Performance comparison with alternative optimization strategies. Solid lines show the mean across 5 random seeds; shaded regions indicate the standard deviation. Linear(β) and Lag(λ0) denote the scalarization weight and the multiplier initialization, respectively.

Exploration Budget Ablation

Fixed budgets fail at both extremes: EF-z0 suppresses exploration entirely, while EF-zmax over-drives it and destabilizes coordination. The adaptive budget resolves this by matching exploration pressure to task feasibility. The update rate \(\alpha_z\) sets how reactive the budget is: a small value (0.1) suits simpler tasks like VMAS-balance, while a large value (0.7) accelerates discovery in the tight MPE-corridor bottleneck. Further ablations on state factorization, the intrinsic reward design, and the feasibility set are in the paper.

Exploration budget ablation on MPE-corridor and VMAS-balance: adaptive budgets with different alpha_z versus fixed EF-z0 and EF-zmax.

Exploration budget \(z\). Ablation of fixed-budget variants and update rates \(\alpha_z\), evaluated with EFXPLORER-MAPPO. Solid lines show the mean over 5 seeds; shaded regions indicate the standard deviation.

Case Study: MPE-corridor

Eight agents, four per side, must pass through a corridor that fits at most two at a time to reach mirrored goals on the opposite side. The distance-to-goal reward becomes locally misleading at the corridor entrance, so pure goal pursuit produces collisions and stalls the team.

Qualitative analysis on MPE-corridor: (a) episode snapshots t1 to t5; (b) per-agent intrinsic reward landscapes, chosen actions, and reward traces at t2, t3, t4.

Qualitative analysis on MPE-corridor. (a) Episode snapshots (\(t_1\)–\(t_5\)) illustrating multi-agent interactions in the narrow passage. (b) Stepwise diagnostics for agents 1–3 (\(t_2\)–\(t_4\)): intrinsic reward landscapes over candidate next states, action directions maximizing extrinsic or intrinsic reward, executed actions, and corresponding instantaneous reward traces. When novelty points toward unproductive motion (toward a wall or away from the goal), EFXPLORER suppresses it and executes a task-aligned action (\(t_2\), \(t_4\)). At the corridor bottleneck (\(t_3\)) it permits a conservative detour to avoid collision: the extrinsic reward dips while the intrinsic reward spikes, so exploration is triggered exactly to resolve the interaction bottleneck.

Benchmarks

Tasks span continuous-control team sports from VMAS, particle-world coordination from MPE, and StarCraft micromanagement from SMAX, run through the JaxMARL and BenchMARL interfaces.

VMAS balance task
VMAS-balance
VMAS give way task
VMAS-give way
VMAS transport task
VMAS-transport
VMAS wheel task
VMAS-wheel
MPE corridor (new) MPE tag SMAX 3s5z_vs_3s6z SMAX 5m_vs_6m

Abstract

Motivation

Discovering cooperative behavior in multi-agent reinforcement learning (MARL) is challenging due to the combinatorial complexity of joint state-action spaces, which hinders the emergence of coordinated behaviors from trial-and-error alone.

Problem

Intrinsic rewards are often used to aid discovery, but naively combining them with team objectives can distort the learning signal, compromising task performance.

Approach

We propose EFXPLORER, a constrained exploration framework that maximizes exploration objectives subject to a constraint that preserves established task performance. We solve the resulting problem using an epigraph reformulation that introduces adaptive exploration budgets. This approach separates intrinsic rewards from task objectives and regulates exploration through task feasibility.

Method

To further encourage diverse and temporally extended exploration, we incorporate a successor distance-based intrinsic reward that captures long-horizon dependencies.

Results

Empirically, our method outperforms strong baselines and induces novel cooperative strategies across SMAX, VMAS, and MPE benchmark suites.

Future Work

Our current evaluation focuses on state-based setups; scaling EFXPLORER to high-dimensional observations such as raw images is an important next step. A further direction is broadening the epigraph formulation to jointly manage exploration, safety constraints, and human preferences within a unified framework.