AI 编程 › AI 资讯 › 正文

Diagnosing On-Policy Self-Distillation for Reasoning Language Models

arxiv · arXiv · 2026-09-30 14:54 · 评分 80

诊断在线策略自蒸馏(OPSD)在语言模型数学推理中的行为:无需外部奖励或更强教师,自教师利用特权信息提供稠密信号;研究解释了为何结果从适度提升到行为崩溃不一,并给出原因分析。

原文:arXiv | 返回 AI 资讯列表