AI 编程 › AI 资讯 › 正文

Mind the Spike: Mechanisms and Brittleness of Visual Massive Activations in Large Vision-Language Models

arxiv · arXiv · 2026-09-27 01:30 · 评分 76

Large vision-language models (LVLMs) inherit massive activations from their text-only bases: spikes where a few fixed hidden channels receive values thousands of times above the typical magnitude. The text spike systematically appears in early layers at a fixed initial position, independently of input content. Visual spikes vary across images, but whether their formation follows a consistent pattern across LVLMs and how they respond to image perturbations remain open questions. We find that some

原文:arXiv | 返回 AI 资讯列表