CausalGate: Causal Importance Distillation for Transformer Module Pruning
arXiv:2607.22720v1 Announce Type: new Abstract: Existing adaptive inference methods for Large Language Models rely on observational heuristics, such as hidden-state similarity or activation magnitudes, to drop redundant modules. However, these correlation-based metrics often fail to capture subtle, non-linear structural computations vital for semantic accuracy. We introduce CausalGate, an intervention-guided framework for compute-efficient transformer inference. During a calibration phase, Causa
延伸阅读
相关资讯
An Integrated Deep Learning and Statistical Framework for Whole-Network Gene--Environment Association with Leaf Vascular Architecture
今天Hierarchical Grading in Large Language Models
今天Learning to Access Computation: Accessibility Plasticity as a Principle of Adaptive Intelligence
今天QFedPolyp: A Communication- and Inference-Efficient Federated Learning Framework for Polyp Segmentation
今天