Hugging Face Integrates Nunchaku 4-bit Diffusion into Diffusers Library
SVDQuant-based quantization now runs natively in Diffusers, cutting VRAM requirements from 24GB to 12GB while accelerating inference.
SVDQuant-based quantization now runs natively in Diffusers, cutting VRAM requirements from 24GB to 12GB while accelerating inference.
Hugging Face releases Holo3.1 with quantized checkpoints for on-device inference, mobile automation support, and cross-framework compatibility.
Tether's QVAC SDK now includes TurboQuant quantization, reportedly enabling 5x context expansion on-device with reduced memory overhead.
The bitsandbytes library applies 4-bit and 8-bit quantization to PyTorch models, making 70B+ parameter LLMs runnable on consumer GPUs and underpinning the QLoRA fine-tuning wave.