World-Model Policy Arbiter for
Goal-Conditioned Reinforcement Learning

1University of Toronto2Vector Institute
WMPA overview: frozen policies are rolled out in a learned world model, the imagined futures are scored by a learned value function, and the best policy acts before the next reassessment

WMPA starts from a bank of frozen goal-conditioned policies trained by different offline GCRL algorithms. At each arbitration step it rolls every policy forward in a learned state-space world model, scores the imagined futures with one shared goal-conditioned value function, and executes the highest-scoring policy for a short commitment interval before choosing again. No policy is retrained.

WMPA in Action

No single offline GCRL algorithm is best across tasks, or even across the phases of one task. On cube-double-play, where a robot arm must stack two cubes, GCIQL succeeds in 36% of episodes and HIQL in only 5%. WMPA, choosing among the six frozen policies within each episode, succeeds in 69%.

A cube-double-play episode in which WMPA alternates between HIQL and GCIQL until both cubes are stacked

One cube-double-play episode. Top: the real state at four arbitration steps and at success, with the policy WMPA selects there. Bottom: the policy executed over the whole episode, one block per arbitration. Control alternates between HIQL, which moves the free gripper to the next cube, and GCIQL, which grasps and places it.

Abstract

Offline goal-conditioned reinforcement learning (GCRL) has produced a diverse set of goal-reaching algorithms, yet no single algorithm performs best across environments, goals, and even different phases of the same task. Rather than deploying only the best-performing policy, we ask whether a set of frozen goal-conditioned policies can be used collectively as a portfolio, deciding at every state which policy should act. Choosing a policy at each state is not straightforward. The policies’ own value functions cannot be compared directly: they may use different scales, and some policies have no value function. We need to judge each policy by the states it is likely to reach, even though we can execute only one policy at a time. We also need to avoid switching so often that control becomes unstable. To address these challenges, we introduce World-Model Policy Arbiter (WMPA), a test-time framework that, given a set of frozen policies as input, rolls out each frozen policy in a learned state-space world model, evaluates the imagined futures with a shared goal-conditioned value function, and executes the highest-scoring policy for a short commitment interval before the next round of arbitration (policy selection). WMPA assumes access to a bank of frozen goal-conditioned policies and requires neither policy retraining nor privileged task-specific knowledge. Under the official OGBench evaluation protocol on 18 state-based datasets spanning maze navigation as well as cube, scene, and puzzle manipulation, WMPA improves the macro-average success rate from the 44% achieved by the best policy selected per dataset to 58%, with statistically significant gains on 12 datasets. These gains include +33 percentage points on cube-double-play and +36 percentage points on scene-play.

Results

We evaluate WMPA under the official OGBench protocol on 18 state-based datasets spanning maze navigation and cube, scene, and puzzle manipulation. The bank holds the six OGBench reference learners (GCBC, GCIVL, GCIQL, QRL, CRL, HIQL), and every method runs on the same episodes (5 goals × 50 episodes × 3 bank seeds). WMPA raises the macro-average success rate from 44% for the best single policy per dataset to 58%, with statistically significant gains on 12 datasets. The largest gains are on multi-object manipulation: +41 points on scene-noisy, +36 on scene-play, +35 on cube-double-noisy, and +33 on cube-double-play. Switching to a random policy at the same interval reaches only 39%, and on 7 datasets WMPA significantly exceeds a per-episode hindsight oracle that never switches, so part of the gain comes from composing policies within an episode.

Success rate per dataset: best single policy versus WMPAHorizontal bars for 18 OGBench datasets. WMPA averages 58% against 44% for the best single policy.0255075100Success rate (%)Best single policy (per dataset)WMPAΔ (points)MazeCubeScenePuzzlepointmaze-medium-navigate | best single policy (QRL): 76 ± 17% | WMPA: 79 ± 3% | Δ +3 [-17, +17]pointmaze-medium-navigate+3antmaze-large-navigate | best single policy (HIQL): 90 ± 3% | WMPA: 91 ± 3% | Δ +0 [-3, +4]antmaze-large-navigate+0cube-single-play | best single policy (GCIQL): 71 ± 3% | WMPA: 85 ± 3% | Δ +14 [+10, +18]cube-single-play+14cube-single-noisy | best single policy (GCIQL): 99 ± 1% | WMPA: 100 ± 0% | Δ +1 [+0, +2]cube-single-noisy+1cube-double-play | best single policy (GCIQL): 36 ± 7% | WMPA: 69 ± 3% | Δ +33 [+25, +40]cube-double-play+33cube-double-noisy | best single policy (GCIQL): 26 ± 9% | WMPA: 60 ± 7% | Δ +35 [+30, +40]cube-double-noisy+35cube-triple-play | best single policy (CRL): 5 ± 3% | WMPA: 1 ± 1% | Δ -4 [-8, -1]cube-triple-play−4cube-triple-noisy | best single policy (GCIVL): 9 ± 3% | WMPA: 18 ± 3% | Δ +9 [+6, +12]cube-triple-noisy+9scene-play | best single policy (GCIQL): 52 ± 5% | WMPA: 87 ± 4% | Δ +36 [+30, +41]scene-play+36scene-noisy | best single policy (GCIVL): 30 ± 3% | WMPA: 70 ± 5% | Δ +41 [+36, +45]scene-noisy+41puzzle-3x3-play | best single policy (GCIQL): 91 ± 4% | WMPA: 100 ± 0% | Δ +9 [+5, +13]puzzle-3x3-play+9puzzle-3x3-noisy | best single policy (GCIQL): 91 ± 3% | WMPA: 93 ± 2% | Δ +1 [-1, +4]puzzle-3x3-noisy+1puzzle-4x4-play | best single policy (GCIQL): 24 ± 5% | WMPA: 51 ± 8% | Δ +27 [+18, +36]puzzle-4x4-play+27puzzle-4x4-noisy | best single policy (GCIQL): 34 ± 5% | WMPA: 58 ± 11% | Δ +24 [+16, +32]puzzle-4x4-noisy+24puzzle-4x5-play | best single policy (GCIQL): 14 ± 3% | WMPA: 18 ± 3% | Δ +4 [+3, +6]puzzle-4x5-play+4puzzle-4x5-noisy | best single policy (GCIVL): 20 ± 3% | WMPA: 20 ± 3% | Δ +0 [+0, +1]puzzle-4x5-noisy+0puzzle-4x6-play | best single policy (GCIQL): 11 ± 2% | WMPA: 17 ± 3% | Δ +6 [+3, +9]puzzle-4x6-play+6puzzle-4x6-noisy | best single policy (GCIQL): 19 ± 3% | WMPA: 20 ± 3% | Δ +1 [-0, +2]puzzle-4x6-noisy+1Average over 18 datasets | best single policy: 44% | WMPA: 58% | Δ +13Average (18 datasets)+13
Success rate (%) on each dataset: the best single policy in the bank versus WMPA over the same frozen policies, averaged over 3 bank seeds × 250 episodes. Δ is WMPA minus the best policy. Hover over a row for the 95% intervals.
Show the numbers as a table
FamilyDatasetBest single policyBest (%)WMPA (%)Δ [95% CI]
Mazepointmaze-medium-navigateQRL76 ± 1779 ± 3+3 [−17, +17]
Mazeantmaze-large-navigateHIQL90 ± 391 ± 3+0 [−3, +4]
Cubecube-single-playGCIQL71 ± 385 ± 3+14 [+10, +18]*
Cubecube-single-noisyGCIQL99 ± 1100 ± 0+1 [+0, +2]*
Cubecube-double-playGCIQL36 ± 769 ± 3+33 [+25, +40]*
Cubecube-double-noisyGCIQL26 ± 960 ± 7+35 [+30, +40]*
Cubecube-triple-playCRL5 ± 31 ± 1−4 [−8, −1]*
Cubecube-triple-noisyGCIVL9 ± 318 ± 3+9 [+6, +12]*
Scenescene-playGCIQL52 ± 587 ± 4+36 [+30, +41]*
Scenescene-noisyGCIVL30 ± 370 ± 5+41 [+36, +45]*
Puzzlepuzzle-3x3-playGCIQL91 ± 4100 ± 0+9 [+5, +13]*
Puzzlepuzzle-3x3-noisyGCIQL91 ± 393 ± 2+1 [−1, +4]
Puzzlepuzzle-4x4-playGCIQL24 ± 551 ± 8+27 [+18, +36]*
Puzzlepuzzle-4x4-noisyGCIQL34 ± 558 ± 11+24 [+16, +32]*
Puzzlepuzzle-4x5-playGCIQL14 ± 318 ± 3+4 [+3, +6]*
Puzzlepuzzle-4x5-noisyGCIVL20 ± 320 ± 3+0 [+0, +1]
Puzzlepuzzle-4x6-playGCIQL11 ± 217 ± 3+6 [+3, +9]*
Puzzlepuzzle-4x6-noisyGCIQL19 ± 320 ± 3+1 [−0, +2]
Average (18 datasets)4458+13

Mean success rate ± half-width of the 95% hierarchical-bootstrap interval. Δ is the paired per-episode difference with its 95% interval; * marks an interval that excludes zero.

BibTeX

@article{quan2026wmpa,
  title   = {World-Model Policy Arbiter for Goal-Conditioned Reinforcement Learning},
  author  = {Quan, Junwei and Opryshko, Evgenii and Rhinehart, Nicholas and Gilitschenski, Igor},
  journal = {arXiv preprint},
  year    = {2026}
}