Hugging Face Research Cuts Knowledge Distillation Memory by 80%, Enabling Single-GPU Training
New offline logits caching and fused KL loss reduce VRAM overhead from 250GB to practical single-GPU levels, opening large-scale model compression to resource-constrained teams.
Synthesia Pivots to Performance Assessment With AI Roleplay Coach Platform
Synthesia launches interactive conversational training, positioning itself as a talent-analytics play rather than a content-generation vendor.
Hugging Face Launches PyTorch Profiler Tutorial Series for Performance Optimization
A new multi-part guide demystifies torch.profiler traces, starting with matrix operations and scaling to large language model optimization.
IBM's Granite 4.1 Shows Data Discipline Can Beat Bigger Models
IBM's new trio of fully-dense LLMs reaches 512K-token context and outperforms a larger mixture-of-experts predecessor through rigorous data curation alone.