Baseten joins Hugging Face Inference Providers, expanding serverless AI access
Baseten is now a supported inference provider on Hugging Face Hub, enabling developers to run open-weights LLMs like DeepSeek V4 Flash and Kimi K3 directly from model pages.
MacPaw and Liquid AI Partner on On-Device Inference for SetApp Developer Ecosystem
Ukraine-based MacPaw partners with Liquid AI to build locally hosted AI inference, positioning SetApp as a platform for AI-native applications with credit-based pricing.
Runware launches modular Sonic Inference Pod to compete with fixed hyperscaler data centers
The AI infrastructure company deploys portable compute units across three continents, betting distributed inference closer to users will outpace centralized facilities.
OpenAI's Intelligence-Per-Dollar Play: GPT-5.6 Pricing Cuts Redefine Model Economics
OpenAI cut GPT-5.6 Luna prices 80% and introduced speed-tiered variants, signaling a shift from capacity competition to cost-per-outcome optimization.
OpenAI cuts GPT-5.6 Luna pricing by 80%, introduces Fast mode for Sol
OpenAI reduces GPT-5.6 Luna costs by 80% and Terra by 20%, while launching Fast mode for GPT-5.6 Sol to deliver 2.5× faster inference at doubled pricing.
Coinbase Reportedly Cuts AI Spending 50% by Switching to Chinese Models GLM and Kimi
Coinbase is reported to have migrated inference workloads to Zhipu AI's GLM and Moonshot's Kimi, reportedly achieving a 50% cost reduction, though the claim lacks independent verification.
Etched AI chip startup reaches $10.3B valuation on transformer-optimized inference gains
The Harvard-founded chip maker closes a $300M Series C led by Sequoia, doubling its valuation in seven months on custom prefill and decode silicon.
Google's 'Frozen v2' chip targets 6-10x efficiency gains for Gemini inference
Alphabet is developing a custom AI accelerator designed to reduce per-token power consumption and reduce dependence on Nvidia's dominance.
Inference Chip Collateral Becomes Real: Upper90's $400M Bet on the Post-GPU Era
A tech investment firm finances AI infrastructure using inference-specific silicon as collateral, signaling a shift from training chips to efficient model-serving hardware.
Why Model Routing in Agents Fails: Cost, Complexity, and Latency Are Deceptive
IBM Research and Hugging Face show that intelligent model selection in agentic systems requires optimizing infrastructure and workload patterns, not just picking cheaper models for easy tasks.
Hugging Face and Cerebras Demonstrate Real-Time Speech-to-Speech with Gemma 4
A modular voice AI pipeline achieves sub-second latency by pairing Google DeepMind's Gemma 4 model with Cerebras inference acceleration, powering conversational robots and assistants.
Etched AI Chip Startup Reaches $5B Valuation on $1B in Booked Orders
The inference-focused chipmaker has raised $800M total and secured $1B in system orders as competition in custom AI silicon intensifies.
Base44 Launches Custom LLM to Defend Against Frontier Model Competition
Wix-owned no-code platform Base44 rolls out proprietary Base1 model trained on tens of millions of real user interactions, betting on specialization over dependence on external LLMs.
Hugging Face Jobs Now Supports vLLM Servers via Single-Command Deployment
Run a private, OpenAI-compatible LLM endpoint on HF infrastructure with one command—no Kubernetes, billed per-minute.
OpenAI Launches Jalapeño, Its First Custom AI Processor for Inference
OpenAI unveiled Jalapeño, an ASIC chip co-developed with Broadcom to handle AI inference workloads and reduce dependence on Nvidia GPUs.
OpenAI's Jalapeño chip signals full-stack AI infrastructure play against Nvidia dependency
OpenAI unveiled Jalapeño, a custom inference processor built with Broadcom, marking the company's entry into purpose-built silicon and a bid to reduce reliance on Nvidia GPUs.
OpenAI and Broadcom Unveil Jalapeño, a Custom LLM Inference Accelerator
OpenAI and Broadcom announced Jalapeño, a purpose-built AI chip for language model inference, designed to improve efficiency and reduce costs at scale.
Local Models Triage OpenClaw PRs at Scale—No API Costs
Hugging Face demonstrates real-time GitHub issue classification using local open-weights models on NVIDIA hardware, eliminating API dependency.
Baseten's $1.5B Round Exemplifies AI Inference Funding Frenzy—and Valuation Tactics
The inference infrastructure startup is raising at a $13B valuation just five months after a $300M Series E, but split pricing reveals how startups are gaming funding rounds.
The Great Model Downgrade: Why Tech Companies Are Ditching Expensive AI
As inference costs soar, enterprises are discovering that smaller models handle 80% of workloads just fine—and the economics could reshape OpenAI and Anthropic's path to IPO.
AutoMegaKernel: RightNow AI's LLM-to-CUDA Compiler Aims for Provably Correct Inference Kernels
A GitHub research project claims to compile LLM computation graphs into single CUDA kernels with formal correctness guarantees, but lacks published benchmarks or third-party validation.
Thousand Token Wood v2: Multi-Model Finance Sim Shows How Heterogeneous Small Models Enable Complex Emergent Behavior
A Hugging Face hackathon project demonstrates that serving four different small models in a single agent economy is tractable when infrastructure abstracts tokenizer variance.
Local LLM Filter Layers Emerge as Enterprise Cost-Control Strategy
Organizations are exploring on-premise language models as pre-filters to reduce API spend on commercial LLMs, though cost savings remain context-dependent.
Groq raises $650M to scale inference cloud after Nvidia licensing deal
The AI chip startup is shifting focus to its inference-as-a-service platform following its $20B partial exit with Nvidia.
Groq raises $650M to scale inference cloud after $20B Nvidia technology licensing deal
The AI chip startup is pivoting toward inference-as-a-service, backed by existing investors including Disruptive and Infinitium.
XCENA's $135M bet: Memory, not compute, is AI's real scaling wall
The Korean chip startup raises Series B at $570M valuation, targeting the data-movement bottleneck that GPUs can't solve alone.
Hugging Face Launches PyTorch Profiler Tutorial Series for Performance Optimization
A new multi-part guide demystifies torch.profiler traces, starting with matrix operations and scaling to large language model optimization.
General Compute bets on SambaNova chips to crack the inference neocloud market
A new inference cloud startup backed by FUSE VC is deploying specialized chips to undercut GPU-heavy competitors in the race for AI inference capacity.
Hugging Face Cuts RL Training Sync Overhead by 98% With Sparse Delta Weights
A new TRL protocol reduces per-step model synchronization from terabytes to tens of megabytes by shipping only changed parameters across distributed training pipelines.
OpenRouter's $1.3B valuation signals shift toward multi-model inference infrastructure
AI gateway OpenRouter raises $113M Series B from CapitalG, doubling its valuation in 12 months as enterprises increasingly avoid vendor lock-in.
NVIDIA's Nemotron-Labs Diffusion Models Generate Multiple Tokens in Parallel, Bypassing Autoregressive Bottleneck
NVIDIA releases diffusion language models at 3B, 8B, and 14B scales that generate and refine tokens in parallel, offering latency improvements for GPU-constrained inference workloads.
Cerebras Systems IPO Surges 108% on First Day, Reaching $66B Valuation
Cerebras Systems priced its IPO at $185/share and opened at $385, closing day one at $311 with a $66B market cap.
Hugging Face Explains Async Continuous Batching: Up to 25% Inference Throughput Gains
Hugging Face's engineering blog details how asynchronous continuous batching eliminates CPU-GPU idle gaps that waste nearly a quarter of LLM inference runtime.
Oracle's $300 Billion Inference Bet: The Riskiest Play in AI Infrastructure
Oracle has staked its enterprise future on a $300B compute deal with OpenAI, betting that AI's real profits live in inference — not model training.