AI 编程 › AI 资讯 › 正文

Copy the Same, Distill the Difference: Initializing Linear Vision Transformers

arxiv · arXiv · 2026-09-29 01:55 · 评分 80

研究提出利用成熟 Softmax 注意力 ViT 初始化线性 ViT 的策略,解决线性注意力模型需从零预训练且性能落后的问题,有望降低训练成本并提升推理效率。

原文:arXiv | 返回 AI 资讯列表