Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks
arXiv:2607.22724v1 Announce Type: new Abstract: Group-based policy optimization has been increasingly used to train large language model (LLM) agents from sparse outcome rewards by comparing trajectories or steps within a group. However, on difficult long-horizon tasks, this comparison can suffer from a sampling imbalance: repeated or low-effect actions dominate the high-probability region of the policy while useful state-changing actions remain under-sampled. This imbalance produces many all-fa
延伸阅读
相关资讯
An Integrated Deep Learning and Statistical Framework for Whole-Network Gene--Environment Association with Leaf Vascular Architecture
今天Hierarchical Grading in Large Language Models
今天Learning to Access Computation: Accessibility Plasticity as a Principle of Adaptive Intelligence
今天QFedPolyp: A Communication- and Inference-Efficient Federated Learning Framework for Polyp Segmentation
今天