Hugging Face Research Cuts Knowledge Distillation Memory by 80%, Enabling Single-GPU Training
New offline logits caching and fused KL loss reduce VRAM overhead from 250GB to practical single-GPU levels, opening large-scale model compression to resource-constrained teams.