Machine Learning28 min reading time
Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
Hacker News
Read full postThe WASTE inference engine runs the 2.78 trillion parameter Kimi K3 model on a 64 GB consumer laptop by streaming model experts from disk, using only about 29 GB RAM at 0.5 tokens per second. This approach enables handling extremely large models without requiring server-grade hardware, though inference speed remains slow.


