One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

Qwen3.8 Max Narrowly Trails Open-Source Kimi K3 on Benchmark

Alibaba's new model scores 56 on Intelligence Index at $1.14 per task, but Kimi K3 wins on cost efficiency

이미지: METAL LAB 생성

Summary

  • Alibaba has released Qwen3.8 Max, a MoE model with 2.4T total parameters
  • It scored 56 on the Artificial Analysis Intelligence Index at a cost of $1.14 per task
  • Open-source model Kimi K3 scored 1 point higher while costing 25% less ($0.86)
모델명
Qwen3.8 Max (Alibaba Qwen)
구조
2.4T 총 파라미터 MoE(Mixture of Experts), 알리바바 발표 기준
Intelligence Index 점수
56점
작업당 비용
1.14달러
비교 대상
Kimi K3 (오픈 가중치 리더)
Kimi K3 비용 우위
Qwen3.8 Max 대비 25% 낮은 0.86달러
평가 기관
Artificial Analysis

Alibaba Unveils Qwen3.8 Max

Alibaba's Qwen team has released a new large language model, Qwen3.8 Max. Alibaba stated that the model uses a Mixture of Experts (MoE) architecture with 2.4T (2.4 trillion) total parameters. Evaluation firm Artificial Analysis measured the model using its proprietary Intelligence Index and published a comparison against leading open-weight models.

Benchmark Results: 56 Points, $1.14 Per Task

According to Artificial Analysis, Qwen3.8 Max scored 56 on the Intelligence Index. The cost of achieving this score was measured at $1.14 per task. On the same metric, Kimi K3 — considered the leading open-weight model — scored one point higher while costing less.

ModelIntelligence IndexCost per Task ($)
Qwen3.8 Max56 bar:561.14
Kimi K357 (estimated, 1 point ahead) bar:570.86

Cost Efficiency Gap

Artificial Analysis noted that Kimi K3's cost per task is 25% lower than Qwen3.8 Max's. While the score gap is just one point, the cost gap is substantial, giving Kimi K3 an edge in terms of the price required to achieve comparable performance. The post stated that the "open weights leader Kimi K3 remains 1 point ahead at 25% lower cost per task."

MoE Architecture and Parameter Scale

Qwen3.8 Max is built on a MoE architecture with 2.4T total parameters, according to Alibaba. The MoE approach activates only a subset of total parameters during inference, allowing large-scale parameter counts while managing computational cost. However, this benchmark result suggests that a large parameter count alone does not guarantee cost efficiency.

The latest trends in the open-source LLM ecosystem have also been covered in metallab.ai's article on Kimi K3.

Implications

This comparison is seen as evidence that the gap between closed large models and open-weight models is narrowing. In particular, the open-weight model's advantage on cost-to-performance metrics could influence competitive dynamics in commercial API pricing. However, since these figures are based on Artificial Analysis's own benchmark methodology, results may vary under different evaluation frameworks, warranting caution in interpretation.