SheepNav
精选昨天0 投票

SAAG: Structured Agent Assessment and Grounding

arXiv:2607.18245v1 Announce Type: new Abstract: Exact-match evaluation of agent-calling obscures qualitatively different failure modes: a model may select the right function yet hallucinate argument values, or satisfy a schema while choosing a agent for the wrong reason. Existing benchmarks collapse these distinctions into a single binary score, leaving practitioners unable to diagnose where agent calls fail. We propose SAAG a cascaded diagnostic framework that decomposes agent-calling evaluatio

延伸阅读

  1. 随机原始-对偶解码:为多目标生成式推荐系统而生
  2. NEXUS:为工具调用型LLM智能体构建结构化运行时安全监控
  3. 大语言模型的信息辨别能力:研究发现模型在来源可信度和事实判断上均存在显著缺陷
查看原文