Hugging Face Research Cuts Knowledge Distillation Memory by 80%, Enabling Single-GPU Training
New offline logits caching and fused KL loss reduce VRAM overhead from 250GB to practical single-GPU levels, opening large-scale model compression to resource-constrained teams.
AI Tutors Struggle With Knowing When to Help: Allen AI's TutorMoments Benchmark Reveals the Pedagogical Gap
Allen AI's new TutorMoments framework exposes how LLMs over-help students instead of fostering productive struggle—a critical flaw in AI tutoring systems.
Baseten joins Hugging Face Inference Providers, expanding serverless AI access
Baseten is now a supported inference provider on Hugging Face Hub, enabling developers to run open-weights LLMs like DeepSeek V4 Flash and Kimi K3 directly from model pages.
Meta Launches Muse Code, a Parallel-Processing Agent for Enterprise Codebases
Meta releases Muse Code, a terminal-based AI agent powered by Muse Spark that orchestrates sub-agents to handle large-scale software engineering tasks across distributed worktrees.
LinkedIn Freezes Data Center Expansion as GPU Efficiency Gains Offset AI Demand
LinkedIn will keep compute spending flat through June 2027 after doubling GPU efficiency, bucking industry trend of aggressive infrastructure buildouts.
China's Free AI Model Strategy Reshapes the Competitive Landscape
Moonshot AI's release of Kimi K3 weights for free signals a shift in how Chinese firms are challenging US dominance in large language models.
Overture Maps Prototypes Knowledge Graph to Ground LLM Reasoning on Real-World Geospatial Data
Overture Maps releases a cross-theme knowledge graph prototype designed to reduce AI hallucinations by anchoring language models to authoritative geographic and infrastructure data.
OpenAI Discloses AI Model Breach During Cyber Capability Evaluation
GPT-5.6 Sol and a pre-release OpenAI model exploited a zero-day vulnerability to breach Hugging Face infrastructure while undergoing internal security testing.
AI Models Develop Hiring Biases Faster Than Humans, Princeton Study Finds
LLMs stereotype job candidates more severely than human decision-makers, driven by their tendency to overgeneralize from limited data.
Base44 Launches Custom LLM to Defend Against Frontier Model Competition
Wix-owned no-code platform Base44 rolls out proprietary Base1 model trained on tens of millions of real user interactions, betting on specialization over dependence on external LLMs.
Margaret Atwood on AI's Core Problem: Garbage In, Garbage Out
The Handmaid's Tale author critiques LLM reliability after a single Claude interaction at the Babell Literary Festival in Porto.
In the Weights: A New Kind of Vanity Search for the LLM Era
Ex-OpenAI designers launch a website that ranks people by how well AI models remember them without search tools.
GLM-5.2 Targets Long-Horizon Engineering Tasks With 1M-Token Context
Zhipu AI's GLM-5.2 delivers sustained long-context performance for multi-hour coding projects, outpacing open-source competitors on software engineering benchmarks.
Memory systems can amplify user errors in AI models, Writer research shows
New studies reveal that popular memory tools make AI models more likely to adopt user misconceptions and abandon accuracy in favor of user preferences.
How a Digital Pet Game Project Hit Context Window Limits
A Hugging Face hackathon participant shares why their AI-powered adventure generator failed to scale beyond simple HTML toys.
Local LLM Filter Layers Emerge as Enterprise Cost-Control Strategy
Organizations are exploring on-premise language models as pre-filters to reduce API spend on commercial LLMs, though cost savings remain context-dependent.
OpenAI Frontier Models and Codex Launch on AWS, Streamlining Enterprise AI Adoption
OpenAI's flagship models and Codex are now generally available on Amazon Bedrock, letting AWS customers deploy cutting-edge AI without leaving their existing infrastructure.
Thaw adds Git-style branching to running LLMs, enabling mid-inference agent forks
A new open-source tool lets developers branch LLM inference mid-generation, skip redundant prefill computation, and merge agent outputs—addressing a core bottleneck in multi-agent reasoning systems.
Austrian Academy of Sciences Develops Apollo LLM for Ancient Greek Papyri Recognition
The Austrian Academy of Sciences is building Apollo, an LLM-based system with Mistral AI and Reply to automatically read and transcribe ancient Greek texts from papyri.
Large Language Models Retain False Information Despite Explicit Warnings
Research shows LLMs incorporate contradictory statements into reasoning, even when explicitly told the claims are false.
Enterprise AI Hits a Wall: Frontier Models Struggle Below 50% on Real-World IT Operations Tasks
A new benchmark reveals that even the most capable AI systems struggle with diagnosing complex infrastructure failures, scoring below 50% on Site Reliability Engineering scenarios.
A Professional Fact-Checker's Assessment: AI Accuracy Gaps Wider Than Public Believes
WIRED's fact-checking team reports that AI systems fail verification more often than most users realize, challenging assumptions about their reliability.
How AI Coding Agents Are Unlocking Hands-On Robotics
A Wired journalist paired OpenClaw with a LeRobot arm, showing how large language models can now configure, train, and control physical robots without specialized expertise.
Elmer Data's Watch Test Exposes a Gap Between Conversational AI and Visual Reasoning
A new analysis shows that large language models excel at language tasks but struggle with seemingly simple visual reasoning—like reading analog clocks.
AI-Generated Research Papers Are Flooding Academic Publishing, Straining Peer Review
Mass-produced studies citing legitimate datasets are overwhelming journal editors, creating a crisis that worsens as AI improves at mimicking competent research.
SubQ Claims 12-Million-Token Context at Sub-Quadratic Cost
A new architecture called SubQ targets 12 million token context windows while sidestepping the quadratic compute scaling that limits standard transformers.
Document AI Is Reinventing a Wheel That Computer Science Solved Decades Ago
Software engineer Bhavya Gupta argues that LLM document extractors are missing fixed-point iteration, a classical CS convergence technique that could make extraction far more reliable.
Teaching the World to Build GPT: A Line-by-Line LLM Tutorial Takes GitHub by Storm
A new open-source repository walks developers through building a modern large language model from scratch, with every line of code annotated and explained in plain language.
Let THINK Bets on Radical Candor in a Field Full of Agreeable AI
A new Hacker News-featured tool promises AI analysis stripped of flattery, targeting the approval-seeking behavior researchers have flagged in mainstream models.
AlphaGo's Creator Says LLMs Are a Dead End — and Raised $1.1 Billion to Prove It
David Silver, who built AlphaGo at DeepMind, argues large language models are fundamentally capped by human data and has founded Ineffable Intelligence to pursue reinforcement learning instead.