
이미지: X — 인프라·칩
Summary
- Together AI compared DeepSeek V4 Flash and GPT-5.6 Luna on the DeepSWE benchmark
- A DeepSeek-first cascade combined with test-suite verification solved more tasks than running Luna alone
- The cascade approach reportedly cut cost per task by 37%
- 발표 주체
- Together AI
- 비교 대상
- DeepSeek V4 Flash vs GPT-5.6 Luna
- 벤치마크
- DeepSWE
- 비용 절감
- 과제당 37% 낮음
- 발행일
- 2026-08-08
Together AI has published a comparative analysis of DeepSeek V4 Flash and GPT-5.6 Luna on the DeepSWE benchmark. According to the company, a "DeepSeek-first cascade" approach combined with test-suite verification solved more tasks than using GPT-5.6 Luna alone.
What was found
Together AI tested both models on the same software engineering task benchmark, DeepSWE. Rather than relying on a single model, the company applied a cascade structure in which DeepSeek V4 Flash generates an initial answer, the result is verified against a test suite, and the task is escalated to GPT-5.6 Luna when necessary.
What changes
The cascade approach was confirmed to solve more tasks than running Luna alone while cutting the cost per task by 37%. Together AI stated, "A DeepSeek-first cascade... solved MORE tasks than Luna alone at 37% lower cost." Specific resolution rates or absolute cost figures were not disclosed, so further verification is needed.



