FreeToken Unlocks Frontier MoE Inference on Consumer Hardware via Dynamic Co-Execution
InfoQ (AI, ML & Data)
Read full postUC Berkeley and MIT researchers developed FreeToken, an open-source inference engine enabling efficient Mixture-of-Experts (MoE) model execution on consumer hardware by dynamically co-scheduling CPU and GPU workloads to overcome PCIe bottlenecks.




