Post-Training Generative Recommenders with Advantage-Weighted Supervised Finetuning

246 · Netflix, Inc. · Oct. 25, 2025, 10:10 p.m.
Summary
This blog discusses the challenges and innovations in post-training generative recommender systems, specifically through the lens of a novel algorithm called Advantage-Weighted Supervised Fine-tuning (A-SFT). The authors explore how generative recommenders can improve user experience by integrating user feedback, address issues with traditional reinforcement learning techniques, and demonstrate the effectiveness of A-SFT compared to existing methods in optimizing recommendations.