GLM-5.3-Flash is 320B-A18B, multimodal and roughly 9x cheaper than GLM-5.2's 744B-A40B. A verified head-to-head on specs, benchmarks, cost and self-hosting -- plus the cases where GLM-5.2 still wins.
Z.ai's GLM-5.3-Flash is the model that ran anonymously as Ox Alpha: a 320B-total, 18B-active natively multimodal MoE with a 1M-token context and MIT open weights. Here are the verified specs, pricing and benchmarks.
Together AI, which serves open-source artificial intelligence models, announced a deal to use compute from Humain in Saudi Arabia, where it can bypass U.S. backlash over data centers.