The config.json diff between Qwen 3.8-27B and Qwen 3.6-27B is empty -- same layers, same hidden size, same vocab. Every gain came from post-training. Here's what actually changed, what the upgrade costs you, and who should stay on 3.6.
The 17GB Apache-2.0 model scoring 52 on Artificial Analysis: which benchmark numbers to trust, the overthinking fix, VRAM needs, and how it compares to Claude and the 2.4T flagship.
Run Qwen3.8-27B on your own GPU: Ollama one-liner, LM Studio, llama.cpp with MTP speculative decoding, VRAM tables, tested 16GB/24GB/Mac configs, and the overthinking fix.
GLM-5.3 claims 2,436 real vulnerabilities found -- and FreeBSD and Red Hat CVEs actually credit the model line. What verifies, what doesn't, and the licence question that matters most.
GLM-5.3 launched August 14, 2026 with big coding claims and cyber capabilities -- but no open weights and no API for two weeks. What shipped, what's verified, and whether to wait.
DeepSeek V4-Pro hit general availability on August 13, 2026. It ranks #2 on SWE-bench Verified at 1/36th the cost -- and last of seven on agentic coding. The split is the story.
From 16:00 UTC on August 16, 2026 DeepSeek moves to peak/off-peak billing. V4-Pro output rises up to 4.55x and cache-hit input up to 12x. Full rate tables and what to do before it lands.
Alibaba's Qwen3.8-Max is a 2.4T MoE with a near-MIT licence and elite algorithmic coding. But its agentic benchmarks reverse under a neutral harness, and you can't run it locally.
xAI's Grok 4.6 numbers verify cleanly -- but the Terminal-Bench version trap, a stale official leaderboard and a 2.3x cost-per-task rise make them easy to misread.
Grok 4.6 improves coding and reasoning but regressed on agentic coding, tripled time-to-first-token and costs 2.3x more per task despite an unchanged price list.