Browse topics
1 article #efficiency
-
Model Compression via Knowledge Distillation
How knowledge distillation trains a smaller model from a teacher, including the core loss, practical variants, code, and common failure modes.
Model Compression via Knowledge Distillation
How knowledge distillation trains a smaller model from a teacher, including the core loss, practical variants, code, and common failure modes.