Researchers from UC Berkeley, MIT, and UT Austin have released FreeToken, a framework that allows personal computers to run large MoE models at interactive speeds. The framework has achieved impressive results with models like Qwen3.6 35B on an 8GB RTX 4060 laptop at 39 tok/s, DeepSeek-V4-Flash 284B on an RTX 5090 desk
FreeToken MoE Inference Framework Enables Personal Computers to Run Large ModelsAI News
Researchers from UC Berkeley, MIT, and UT Austin have released FreeToken, a framework that allows personal computers to run large MoE models at interactive speeds. The framework has achieved impressive results with models like Qwen3.6 35B on an 8GB RTX 4060 laptop at 39 tok/s, DeepSeek-V4-Flash 284B on an RTX 5090 desk
