Agent Priors-guided Policy Learning

One structural prior shapes how a skill is learned and tells the runtime agent when to use it.

Puming Jiang* Tianrun Hu* Haozhe Du Yibo Li Zhiwei Xue Xinhu Li Harold Soh

National University of Singapore  ·  *Equal contribution

APPL composes prior-specific skill policies at runtime, with objects shifted beyond the demonstrations.

Abstract.

Robots that learn from a few demonstrations often require two forms of generalization. Compositional generalization recombines skills to solve new tasks, and skill generalization lets the learned policy behind each skill work in new situations. The two depend on each other, yet information is lost between composition and the skills it calls. Where a skill works is determined by the structure its policy is trained with, while composition sees the skill only through a separate description, such as a name, an instruction, or a symbolic operator, that omits this structure. Our key idea is to use each policy's structural prior as part of the interface between composition and the skill. A structural prior states what a behavior depends on, for example that a grasp depends only on the gripper's pose relative to the object. Built into training, it shapes where the policy generalizes; stated in language, it tells composition where the policy applies. We instantiate this idea in Agent Priors-guided Policy Learning (APPL). A construction agent segments complete demonstrations into reusable skills, proposes several structural priors for each skill, and trains and verifies one policy per prior. A runtime agent then selects among these prior-specific policies and composes them toward new task goals using their interfaces. Across MetaWorld and long-horizon ManiSkill tasks, APPL improves out-of-distribution skill generalization and enables previously unseen skill compositions; ablating the interface information substantially reduces performance. These results support the use of training-time structural assumptions as a bridge between skill learning and skill composition.

What is a prior?

What a skill depends on, built into how it learns.

Train5 demos
Testdrag the block

How it works.

APPL pipeline: construction agent, frozen skill library, runtime agent
Build prior-specific policies offline, compose them online. Orange marks the prior's two roles.

Long-horizon rollouts.

Twelve demonstrations per task. Objects shifted beyond them. Logged episodes, replayed exactly.

  • Reference frame and gripper frame
  • Δ gripper moves, drawn ×12
  • Object path first, then gripper
  • Frame change: scene dims, axes span the view

The interface: each policy's prior, as the agent reads it

1Open drawer

Prior
Gripper targets in world coordinates; a phase head separates approach, pull and release.
Use when
Drawer at its calibrated place; the entry stage is uncertain.

2Red block to pad

Prior
Gripper position as an offset from the red goal pad.
Use when
Red block grasped; the pad may have moved.
Prior
Anchor switches with the phase: world, red block, pad, gripper, blue block.
Use when
The phase is clear; red, the pad or blue may have moved.

3Blue block into drawer

Prior
Progress along the straight line from the blue block to its slot, plus a small offset.
Use when
Drawer open; the blue block may have moved.
Prior
Gripper position as an offset from the blue slot inside the drawer.
Use when
Blue block carried; alignment with the drawer slot is the risk.

Results.

89.6%OOD skill success, 2 demosvs 37.9% fixed prior
50.0%shifted objectsvs 10.0% Diffusion Policy
92.5%task-level variantsresume or stop early
8/16new compositionsvs 2/16 without interface

Exp. 1: agent-designed priors for skill generalization

Six MetaWorld tasks, success (%) with N demonstrations.

MethodN = 2N = 5N = 10N = 20Mean
IIDOODIIDOODIIDOODIIDOODOOD
Diffusion Policy (B0)73.3328.9682.5037.2997.5038.54100.0046.6737.86
Relational prior (B1)85.8337.9293.3342.71100.0052.29100.0056.2547.29
APPL, first proposal (q1)84.1756.0495.8363.54100.0073.33100.0078.7567.92
APPL, best of three (q3)90.8380.00100.0081.46100.0090.42100.0088.9685.21
APPL (q4)96.6789.58100.0093.54100.0093.33100.0093.1392.40

Exp. 2: long-horizon composition

Five ManiSkill tasks. M: shifted objects · T: task-level variants · Comp.: new compositions.

MethodDrawerBufferRetrievePegPourTotal (%)Comp.
MTMTMTMTMTMT
DP0/8–0/8–1/8–3/8–0/8–10.0––
SinglePrior1/8–0/8–2/8–1/8–0/8–10.0––
Agent+VLA1/88/81/84/85/88/85/84/81/84/832.570.06/16
APPL w/o interface information0/84/84/85/82/86/80/87/82/84/820.065.02/16
APPL w/o prior information4/87/80/84/84/86/83/86/87/86/845.072.53/16
APPL w/o HL agent0/84/82/84/80/84/83/84/84/85/822.552.51/16
APPL w/o verification2/85/83/86/82/85/88/86/86/85/852.567.57/16
APPL3/87/82/88/87/88/81/87/87/87/850.092.58/16

Selection needs the interface

Same library, interface hidden: 50.0 → 20.0% · 92.5 → 65.0% · 8/16 → 2/16.

Priors enable new handoffs

The return policy's prior ignores particles, so the agent used it to return a full cup.

BibTeX

@article{jiang2026appl,
  title   = {Agent Priors-guided Policy Learning},
  author  = {Jiang, Puming and Hu, Tianrun and Du, Haozhe and Li, Yibo and
             Xue, Zhiwei and Li, Xinhu and Soh, Harold},
  journal = {arXiv preprint arXiv:2609.35690},
  year    = {2026}
}