Hugging Face Research Cuts Knowledge Distillation Memory by 80%, Enabling Single-GPU Training
New offline logits caching and fused KL loss reduce VRAM overhead from 250GB to practical single-GPU levels, opening large-scale model compression to resource-constrained teams.
LinkedIn Freezes Data Center Expansion as GPU Efficiency Gains Offset AI Demand
LinkedIn will keep compute spending flat through June 2027 after doubling GPU efficiency, bucking industry trend of aggressive infrastructure buildouts.
OpenAI's GPT-5.6 family targets efficiency across the cost-capability spectrum
OpenAI released GPT-5.6 Sol, Terra, and Luna with stacked infrastructure optimizations aimed at balancing frontier intelligence with operational cost.
Liquid AI releases LFM2.5-Encoders: sub-billion-parameter models for CPU-based long-context NLP
Two new encoder models from Liquid AI match larger baselines on GLUE and SuperGLUE while maintaining 8,192-token context and 3.7× CPU speed advantage over ModernBERT-base.
Google's 'Frozen v2' chip targets 6-10x efficiency gains for Gemini inference
Alphabet is developing a custom AI accelerator designed to reduce per-token power consumption and reduce dependence on Nvidia's dominance.
Why AI Models Fail Where Babies Excel: The EgoBabyVLM Challenge
A new benchmark reveals that cutting-edge vision-language models struggle with the messy, multimodal learning that infants master effortlessly.
Subquadratic Claims Breakthrough in LLM Efficiency With Independent Validation
Miami-based startup Subquadratic released third-party benchmarks for SubQ, its new model claiming 12x context scaling and lower energy consumption than existing LLMs.
Remote's AI-Driven Payroll Efficiency: 50% Revenue Growth Per Employee Without Hiring
Amsterdam-based payroll platform Remote reached $300M ARR and achieved 50% revenue-per-employee growth by embedding AI across all departments, not just engineering.
OlmoEarth v1.1 cuts satellite-imagery inference costs by 3x through token optimization
Allen Institute releases OlmoEarth v1.1, a more efficient earth-observation model family that maintains v1 performance while reducing compute through shorter token sequences.