
Ox Alpha Was Z.ai: Inside the GLM-5.3-Flash Stealth Test
Z.ai revealed that Ox Alpha was an early GLM-5.3-Flash preview. The anonymous experiment drew 343.5 million OpenRouter requests before the company attached its name.

Z.ai revealed that Ox Alpha was an early GLM-5.3-Flash preview. The anonymous experiment drew 343.5 million OpenRouter requests before the company attached its name.
Fast reads on model releases, compute strategy, policy pressure, and the companies fighting over the AI stack.

SpecPTC launches safe tool calls while agent code is still streaming. Alex Zhang reports 1–1.2x RLM gains, but the bigger story is runtime scheduling.

Grok 4.6 reaches a 61 Intelligence Index score at $2/$6 API pricing. Grok Bot adds persistent computer use, shared state, routines, and a new governance problem.

Prime Intellect's recursive harness framing explains why AI agents are becoming an orchestration, verification, and training-data business, not just a model race.

Alibaba's 2.4T-parameter Qwen3.8-Max pairs 1M context, $2/$6 API pricing, and vendor-reported agent gains with a promised checkpoint few teams could deploy themselves.

Reasoning effort is not a universal quality slider. It is becoming the control plane that allocates tokens, latency, tools, and verification across AI workloads.

TileRT is not just a faster inference engine. Separate Xiaomi MiMo and Z.ai deployments show why ultra-low-latency runtimes are becoming a new battleground for frontier AI products.

Kimi K3 pairs a 57.1 Artificial Analysis score with 2.8T parameters, 1M-token context, $0.94 task cost, and open weights promised for July 27.

Meta Muse Spark 1.1 API pricing, 1M-token context, benchmarks, coding agents, and what Meta's paid agent platform means for developers.

Kolmogorov-Arnold Networks replace scalar weights with learned functions. Two years of evidence show where KANs work, where they fail, and why the idea survived.

Loop Engineering turns the hidden management work around coding agents into software: triggers, scoped execution, independent verification, durable state, budgets, and explicit stop conditions.

Grok 4.5 combines a 54 Intelligence Index score, 90-token-per-second measured speed, $0.31 benchmark task cost, Cursor-trained agent behavior, and live search in xAI's strongest model release yet.

DeepSpec turns speculative decoding from a hidden serving trick into an open training stack, with DSpark claiming 60% to 85% faster V4-Flash generation.