Machine Learning6 min reading time
Smaller, faster, safer: running Kimi and GLM at scale
Hacker News
Read full postCloudflare's Workers AI enhances serving of large models like Moonshot's Kimi K-series and Z.ai's GLM by quantizing KV cache to 8-bit floats, compressing model weights, and protecting shared caches, doubling context capacity and improving efficiency without accuracy loss.



