Zhipu (Z.ai) Unmasks The Mystery Ox Alpha Model as GLM-5.3-Flash, Revealing That It Was Run Entirely On Chinese GPUs While Serving 100 Trillion Tokens/Day
Last week, the mystery around the Ox Alpha AI model was so thick that you could cut it with a knife. Yet, a few days later, the enigma has been de-mystified, and that too by none other than the China-based AI lab Zhipu (Z.ai), which has identified the Ox Alpha as its very own GLM-5.3-Flash. Zhipu AI's GLM-5.3-Flash reduces the size of its KV cache by a factor of 4.4x relative to the GLM-5.3 model As we noted last week, someone anonymously dropped Ox Alpha on OpenRouter and OpenCode on August 20, offering a 1-million-token multi-modal (text, audio, video) context [โฆ]
Read full article at https://wccftech.com/zhipu-z-ai-unmasks-the-mystery-ox-alpha-model-as-glm-5-3-flash-revealing-that-it-was-run-entirely-on-chinese-gpus-while-serving-100-trillion-tokens-day/

