SPACeR: Self-Play Anchoring with Centralized Reference Models
Published in International Conference on Learning Representations (ICLR), 2026
Check out our website: SPACeR! Arxiv here
Abstract:
SPACeR leverages a pretrained tokenized autoregressive motion model as a centralized reference policy to guide decentralized self-play. Through log-likelihood rewards and KL divergence constraints, it aligns self-play RL policies with human driving distributions, outperforming pure self-play methods on WOSAC while achieving 10x faster inference and 50x fewer parameters than imitation learning models.
