One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

DeepSeek V4 Pro GA Benchmarks Leak, Nears Top Open-Source Tier

Matches Kimi-K3 and GLM-5.2 in agentic evaluations like Terminal Bench and HLE

DeepSeek V4 모델과 다른 AI 모델들의 벤치마크 성능 비교표

이미지: X — 뉴스 앰프 화면 갈무리

Summary

  • DeepSeek-V4-Pro-0813 GA benchmark figures leaked on X, showing strong results in agent-related evaluations
  • It scored 87.9 on Terminal Bench 2.1, beating Opus-4.8 (85.0) but falling slightly short of Kimi-K3 (88.3)
  • The GA release, which leans on price competitiveness, follows the low-cost pricing table unveiled on August 12 and reflects the open-source camp's ongoing catch-up
모델명
DeepSeek-V4-Pro-0813 (GA)
Terminal Bench 2.1
87.9점 (Kimi-K3 88.3 / Opus-4.8 85.0)
HLE (with tools)
60.0점 (Opus-4.8 57.9)
비교 대상
GLM-5.2, Kimi-K3, Opus-4.8
유출 경로
X 계정 Chubby(@kimmonismus)

The benchmarks leaked before the announcement

Benchmark figures for the general availability (GA) release of DeepSeek-V4-Pro-0813 circulated on X (formerly Twitter) ahead of the official launch. A table shared by the news account Chubby listed scores across 10 agent-related evaluations, showing the model scoring 87.9 on Terminal Bench 2.1 — ahead of Opus-4.8's 85.0. It fell slightly short, however, of open-source rival Kimi-K3's 88.3. Its HLE (Humanity's Last Exam) score with tool use came in at 60.0, higher than Opus-4.8's 57.9. Chubby noted, "Not a bad upgrade, but it's a shame the comparison is against 4.8 instead of Opus 5."

What's going on here

DeepSeek first unveiled the V4 series back in April, and after a preview phase, has now moved to a GA release. Earlier, on August 12, the model's API pricing was disclosed: $0.435 per million input tokens (on cache miss) and $0.87 per million output tokens. A comparison published the same day showed DeepSeek undercutting Grok 4.6 ($2/$6 per million input/output tokens), Claude Opus 5 ($5/$25), and GPT-5.6 Sol ($5/$30). In other words, the leaked benchmarks emerge against a backdrop where the model already rivals frontier-tier competitors on price, and now appears competitive on performance as well.

Agentic benchmarks measure more than simple Q&A — they assess the ability to write code, operate a terminal, and carry out multi-step tasks autonomously. Categories like Toolathlon-Verified and DSBench-FullStack are representative examples. DeepSeek reaching parity with open-source frontier models like Kimi-K3 and GLM-5.2 in this domain signals not just a narrowing gap with closed models, but intensifying competition within the open-source camp itself.

Why it matters

A limitation of this benchmark leak is that the comparison is against Opus 4.8, a generation behind the latest Opus 5, rather than the current flagship. Even so, the trend of models delivering solid agentic performance at low prices is clear. Together AI's launch of fine-tuning support for DeepSeek V4 Flash 0731 on August 11 fits into this same pattern. For developers, this adds one more option: building agentic workflows on a high-value, cost-effective open-source model instead of a top-tier closed one.