We built a complete production-grade inference service from scratch on a cluster of more than 100,000 Chinese-made AI accelerators. All production inference for GLM-5.3-Flash runs on this system. This was not easy. No one had previously deployed a cluster of Chinese-made accelerators at this scale. We faced relatively limited chip memory capacity and bandwidth, while also needing to support a new model architecture, a 1M-token context window, and multimodal requests. The ecosystem was immature,
CryptoAlerta — análise de criptomoedas e mercado em tempo real