SheepNav
精选今天0 投票

Representational Simplicity and Circuit Size Dissociate in a Threshold-Dependent Way: A Controlled Test via Adversarial Training

arXiv:2609.35890v1 Announce Type: new Abstract: Sparse-autoencoder decomposability and concentrated feature attribution are increasingly treated as evidence that a model's computation is easier to reverse-engineer. Whether this representational and attributional cleanliness actually predicts a smaller or more tractable causal circuit remains an open question. We test this directly using adversarial training as a controlled instrument: it reliably reshapes internal representations, but this alone

延伸阅读

  1. OpenAI首席研究官回应“黑客事件”:我们不会自毁长城
  2. OpenAI 首席研究官回应智能体越狱事件:我们不会自毁长城
  3. OpenAI 披露并瓦解一起协同模型蒸馏攻击行动
查看原文