Machine LearningDev9 min reading time

Meet FreeToken: An Edge-Native MoE Serving Engine that Runs 753B GLM-5.2 on a Single Workstation GPU

MarkTechPost
Read full post
Researchers from UC Berkeley and UT Austin developed FreeToken, an edge-native MoE serving engine enabling large models like 753B GLM-5.2 to run on a single workstation GPU by elastically mapping computation across available hardware. FreeToken supports interactive speeds for models ranging from 35B on laptops to 753B on workstations, targeting solo developers and small teams needing cost-effective local inference.

More on this story


More in Machine Learning

Machine Learning3 min read

Nvidia and Palantir fine-tune a 30B Nemotron model for Nvidia’s supply chain. It beats a model 18 times its size.

Covered by 3 sources
Machine Learning6 min read

CoreWeave Puts Field Engineers Inside Customer Teams for Physical AI

Covered by 2 sources
Machine Learning2 min read

Weatherwatch: AI model beats standard methods at predicting cyclones

The Guardian