Anthropic finds collaboration breaks down in agent swarm experiments
Testing 45 agents on vulnerability hunting and game-building revealed a pattern where teamwork falls apart
A Chinese AI developer that makes the chatbot and model "Kimi." It is a regular fixture at the top of open-source model rankings. Headquartered in Beijing, this startup is, despite its name "Moonshot," a software company with no connection to space exploration or aerospace business. Its flagship model Kimi has been released with an emphasis on its ability to handle long context in a single pass, and it is often grouped together with DeepSeek, Alibaba's Qwen, and others as part of the lineage of Chinese-made open-source models. As of August 2026, OpenAI and Anthropic have China AI
Current rank (1M)
31
Rank over the last 30 days
2026.08.17 – 2026.09.15 · High 21 · Low 31 · Now 31
Testing 45 agents on vulnerability hunting and game-building revealed a pattern where teamwork falls apart
Moonshot AI's new benchmark ranks GPT-5.6 Sol first at 59.7%, revealing that many "reasoning errors" are actually perception failures
GPT-5.6 Luna cut 80%, Claude Opus 5 launches at half the price of top-tier model
Released via Expert Mode, V4-Pro edges out Opus-4.8 on Terminal Bench
NVIDIA's open-weight Nemotron 4 is said to reach 1 trillion parameters, but Kimi K3 and DeepSeek V4 are already bigger
Scoring 61 on the intelligence index, it ties GPT-5.6 Sol and closes to within two points of Claude Opus 5 — yet costs less than half as much. A leap achieved not by scaling up but by changing the training recipe alone. The price war sparked by xAI has only just begun.
Matches Kimi-K3 and GLM-5.2 in agentic evaluations like Terminal Bench and HLE
Score jumps 5 points a month after launch, takes the runner-up spot to Claude Opus 5 in agentic tasks
Artificial Analysis tests AI on real-world spreadsheets and documents. Claude Opus 5 scored 54%; 57% of failures anchored on a wrong early hypothesis.
Anthropic is pushing for a September-October listing, but Chinese rival models and data center backlash have emerged as key variables
Amazon Bedrock has adjusted on-demand inference pricing for its OpenAI GPT-5.6 model lineup
Moonshot AI tested Kimi K3 across major inference providers, with Together AI ranking first or tied for first in 3 of 4 benchmarks