SheepNav
新上线2个月前0 投票

Weight-Space Geometry of Offline Reasoning Training

arXiv:2606.23740v1 Announce Type: new Abstract: Offline reinforcement-learning losses (RFT, RIFT, DFT, Offline GRPO, DPO) are widely used to distill reasoning from large teachers into smaller students, and are typically compared on downstream accuracy alone. We ask whether they are mechanistically distinct or converge to a similar weight update. Training six methods (SFT, RFT, DFT, RIFT, Offline GRPO, DPO) on identical math rollouts from a single base model (Qwen3-4B) with attention-only LoRA, w

延伸阅读

  1. AI面试的终极形态:两个机器人互相面试
  2. OpenAI被指“协助教唆”加拿大Tumbler Ridge枪击案,面临30起新诉讼
  3. 纽约市新政策:高中前全面禁止学生使用AI
查看原文