Hugging Face Integrates Nunchaku 4-bit Diffusion into Diffusers Library
SVDQuant-based quantization now runs natively in Diffusers, cutting VRAM requirements from 24GB to 12GB while accelerating inference.
SVDQuant-based quantization now runs natively in Diffusers, cutting VRAM requirements from 24GB to 12GB while accelerating inference.
A new plugin directs lightweight AI tasks to specialized smaller language models, targeting inference cost reduction through intelligent task routing.
A new open-source tool lets developers branch LLM inference mid-generation, skip redundant prefill computation, and merge agent outputs—addressing a core bottleneck in multi-agent reasoning systems.