QooryBeta
← News

We’re releasing Gemma 4 NVFP4 quants that run 1.5× faster on your GPU. Gemma-4-12B NVFP4 works on 11GB VRAM. 26B-A4B hits 13K tok/s (B200). Unsloth NVFP4 enables faster, more accurate 4-bit Blackwell inference. Blog: https://t.co/EPAHgqe2B2 Gemma NVFP4: https://t.co/RWflncpLPJ

@0xkeenz·Jul 14, 2026·2 sources·positive
Read article
AI Summary

Unsloth releases Gemma 4 NVFP4 quantized models that run 1.5x faster on GPUs, with the 12B model fitting in 11GB VRAM and the 26B-A4B achieving 13K tokens per second on B200.

Related Projects
All Sources