SPACeR: Self-Play Anchoring with Centralized Reference Models

Published in International Conference on Learning Representations (ICLR), 2026

Check out our website: SPACeR! Arxiv here

Abstract:

SPACeR leverages a pretrained tokenized autoregressive motion model as a centralized reference policy to guide decentralized self-play. Through log-likelihood rewards and KL divergence constraints, it aligns self-play RL policies with human driving distributions, outperforming pure self-play methods on WOSAC while achieving 10x faster inference and 50x fewer parameters than imitation learning models.