Machine Learning6 min reading time

Smaller, faster, safer: running Kimi and GLM at scale

Hacker News
Read full post
Cloudflare's Workers AI enhances serving of large models like Moonshot's Kimi K-series and Z.ai's GLM by quantizing KV cache to 8-bit floats, compressing model weights, and protecting shared caches, doubling context capacity and improving efficiency without accuracy loss.

More in Machine Learning

Machine Learning3 min read

Nvidia and Palantir fine-tune a 30B Nemotron model for Nvidia’s supply chain. It beats a model 18 times its size.

Covered by 3 sources
Machine Learning6 min read

CoreWeave Puts Field Engineers Inside Customer Teams for Physical AI

Covered by 2 sources
Machine Learning2 min read

Weatherwatch: AI model beats standard methods at predicting cyclones

The Guardian