SheepNav
精选2个月前0 投票

Can Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability?

arXiv:2606.24026v1 Announce Type: new Abstract: Mechanistic interpretability has made substantial progress in automatically localizing circuits, but explaining what localized components do remains labor-intensive and difficult to standardize. In this work, we study whether language model (LM) agents can assist with this explanation problem once a circuit has already been identified. We introduce AgenticInterpBench, a benchmark for circuit explanation built from 84 semi-synthetic transformer circ

延伸阅读

  1. AI代理与在线调查数据质量控制的竞赛
  2. Paper Pilot:人机协同的专家系统,为科学论文生成提供证据可追溯性
  3. 噪声中的信号:为生物医学文本分类打造可审计的可靠性层
查看原文