Skip to content
Topic

#Mixture Of Experts

11 articles on Mixture Of Experts — news, releases, guides and analysis from the SourceFeed engine.

Google's 4.7x Qwen 3.5 Speedup Is a Sharding Story
Article 3d ago 3

Google's 4.7x Qwen 3.5 Speedup Is a Sharding Story

Two KV heads and 512 experts pushed Ironwood onto DP+EP, the same topology vLLM now recommends for GPUs.

Priya Nair
Models Aren't Getting Dumber — They're Getting Unbundled

Models Aren't Getting Dumber — They're Getting Unbundled

Article · 1w ago0
Run Mixtral 8x7B Locally with llama.cpp and Benchmark MoE vs. Dense

Run Mixtral 8x7B Locally with llama.cpp and Benchmark MoE vs. Dense

Tutorial · 3w ago0
How a 20B Model Hits 120 tok/s on an iPhone

How a 20B Model Hits 120 tok/s on an iPhone

Article · 3w ago1
Your SSD Is the New VRAM for Local LLMs

Your SSD Is the New VRAM for Local LLMs

Article · 3w ago1
Your SSD Is the New VRAM

Your SSD Is the New VRAM

Article · 3w ago1
AirLLM's 4GB 70B Trick Is Real, and Beside the Point

AirLLM's 4GB 70B Trick Is Real, and Beside the Point

Article · 3w ago1
A 26B Model in 2 GB of RAM, Courtesy of Your SSD

A 26B Model in 2 GB of RAM, Courtesy of Your SSD

Article · 1mo ago2
Kimi K3 Turns a Year of Efficiency Tricks Into 2.8T Parameters

Kimi K3 Turns a Year of Efficiency Tricks Into 2.8T Parameters

Article · 1mo ago1
Inkling Bets Fine-Tuning Beats Frontier Chatbots

Inkling Bets Fine-Tuning Beats Frontier Chatbots

News · 1mo ago5
Inkling Puts 975B Open Weights in Reach

Inkling Puts 975B Open Weights in Reach

Article · 1mo ago2