Skip to content
Topic

#Quantization

12 articles on Quantization — news, releases, guides and analysis from the SourceFeed engine.

Convert and Quantize Hugging Face Models to GGUF for llama.cpp
Tutorial 1d ago 0

Convert and Quantize Hugging Face Models to GGUF for llama.cpp

Turn any Hugging Face checkpoint into a 4-bit GGUF that runs fast and small on your own hardware.

Mariana Souza
Latent Reasoning Escapes the Lab, Bolted Onto DeepSeek-V4

Latent Reasoning Escapes the Lab, Bolted Onto DeepSeek-V4

Article · 4d ago2
Run Mixtral 8x7B Locally with llama.cpp and Benchmark MoE vs. Dense

Run Mixtral 8x7B Locally with llama.cpp and Benchmark MoE vs. Dense

Tutorial · 1w ago0
An 8B Fine-Tune Now Fits in 4 GB of VRAM

An 8B Fine-Tune Now Fits in 4 GB of VRAM

Article · 1w ago1
Your SSD Is the New VRAM for Local LLMs

Your SSD Is the New VRAM for Local LLMs

Article · 1w ago1
Cloudflare's Quantization Math Is Right. The Disclosure Isn't

Cloudflare's Quantization Math Is Right. The Disclosure Isn't

Article · 1w ago0
Quantize the Decode, Not the Prefill

Quantize the Decode, Not the Prefill

Article · 1w ago5
An LLM on an $8 Chip Is Mostly a Memory Trick

An LLM on an $8 Chip Is Mostly a Memory Trick

Article · 2w ago0
Bonsai 27B Puts Real Agents on Phones

Bonsai 27B Puts Real Agents on Phones

Article · 4w ago1
Quantize and Run Llama 3.2 on Apple Silicon with llama.cpp

Quantize and Run Llama 3.2 on Apple Silicon with llama.cpp

Tutorial · 1mo ago0
Demystifying Integer Quantization for Neural Network Inference

Demystifying Integer Quantization for Neural Network Inference

Article · 1mo ago0
Xiaomi's MiMo-V2.5-Pro-UltraSpeed Pushes a 1T Model Past 1000 Tokens/Sec on Commodity GPUs

Xiaomi's MiMo-V2.5-Pro-UltraSpeed Pushes a 1T Model Past 1000 Tokens/Sec on Commodity GPUs

News · 2mos ago5