Distillation Scaling Laws

cs.AI updates on arXiv.org 07月28日 12:43

Distillation Scaling Laws

提出一种基于计算预算和分配的蒸馏模型性能评估法则，优化教师和学生模型性能，提供针对不同场景的蒸馏方案，提升大规模蒸馏效果。

arXiv:2502.08606v2 Announce Type: replace-cross Abstract: We propose a distillation scaling law that estimates distilled model performance based on a compute budget and its allocation between the student and teacher. Our findings mitigate the risks associated with large-scale distillation by enabling compute-optimal allocation for both the teacher and student to maximize student performance. We provide compute-optimal distillation recipes for two key scenarios: when a teacher already exists, and when a teacher needs training. In settings involving many students or an existing teacher, distillation outperforms supervised learning up to a compute level that scales predictably with student size. Conversely, if only one student is to be distilled and a teacher also requires training, supervised learning is generally preferable. Additionally, our large-scale study of distillation increases our understanding of the process and helps inform experimental design.

Fish AI Reader

AI辅助创作，多种专业模板，深度分析，高质量内容生成。从观点提取到深度思考，FishAI为您提供全方位的创作支持。新版本引入自定义参数，让您的创作更加个性化和精准。

FishAI

鱼阅，AI 时代的下一个智能信息助手，助你摆脱信息焦虑

联系邮箱 441953276@qq.com

相关标签

蒸馏模型计算预算性能优化

相关文章

Sparse Maximal Update Parameterization (SμPar): Optimizing Sparse Neural Networks for Superior Training Dynamics and Efficiency

Node.js 最佳实践：开发人员指南

Webassembly：网络应用程序的近原生性能

This AI Paper from Databricks and MIT Propose Perplexity-Based Data Pruning: Improving 3B Parameter Model Performance and Enhancing Language Models

是时候向谷歌字体说再见了：缓存性能 (2020)

用于连接处理的简单、高效和稳健的哈希表

利用 Zig 的分配器

This AI Research Discusses Achieving Efficient Large Language Models (LLMs) by Eliminating Matrix Multiplication for Scalable Performance

Rails 上的异步 Ruby

使用 SIMD 指令更快地扫描 HTMLChrome 浏览器版