One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

Together AI Publishes DeepSeek·GPT-5.6 Luna Cascade Test Results

DeepSeek-first cascade solves more tasks than Luna alone on DeepSWE benchmark

이미지: X — 인프라·칩

Summary

  • Together AI compared DeepSeek V4 Flash and GPT-5.6 Luna on the DeepSWE benchmark
  • A DeepSeek-first cascade combined with test-suite verification solved more tasks than running Luna alone
  • The cascade approach reportedly cut cost per task by 37%
Video from the source
발표 주체
Together AI
비교 대상
DeepSeek V4 Flash vs GPT-5.6 Luna
벤치마크
DeepSWE
비용 절감
과제당 37% 낮음
발행일
2026-08-08

Together AI has published a comparative analysis of DeepSeek V4 Flash and GPT-5.6 Luna on the DeepSWE benchmark. According to the company, a "DeepSeek-first cascade" approach combined with test-suite verification solved more tasks than using GPT-5.6 Luna alone.

What was found

Together AI tested both models on the same software engineering task benchmark, DeepSWE. Rather than relying on a single model, the company applied a cascade structure in which DeepSeek V4 Flash generates an initial answer, the result is verified against a test suite, and the task is escalated to GPT-5.6 Luna when necessary.

What changes

The cascade approach was confirmed to solve more tasks than running Luna alone while cutting the cost per task by 37%. Together AI stated, "A DeepSeek-first cascade... solved MORE tasks than Luna alone at 37% lower cost." Specific resolution rates or absolute cost figures were not disclosed, so further verification is needed.