FreeToken MoE Inference Framework Enables Personal Computers to Run Large Models
Researchers from UC Berkeley, MIT, and UT Austin have released FreeToken, a framework that allows personal computers to run large MoE models at interactive speeds. The framework has achieved impressive results with models like Qwen3.6 35B on an 8GB RTX 4060 laptop at 39 tok/s, DeepSeek-V4-Flash 284B on an RTX 5090 desk
