
Image: METAL
Summary
- Microsoft has released MAI Code 1.1 Flash, a code model for GitHub Copilot, touting a 25% improvement in token efficiency and roughly a quarter of the previous cost
- It scored 62.9% on Terminal Bench 2.1, well below DeepSeek V4 Flash-0731's 82.7%, and was also more expensive
- The move appears to be part of Microsoft's ongoing strategy of prioritizing cost savings, after recently replacing OpenAI and Anthropic models in Copilot with its own MAI models
On paper, it looks like a win
On the 12th, Microsoft released MAI Code 1.1 Flash, a code-generation model for GitHub Copilot. The company said it improves token efficiency by 25% over its predecessor, MAI-Code-1-Flash (released last June), while cutting costs to about a quarter of the previous level. Microsoft also said the rate at which developers actually adopt code generated by the model rose by 4 percentage points.
The training method stands out as well. Microsoft said it trained the model using "hundreds of thousands of reinforcement learning environments within GitHub Copilot." Reinforcement learning is a training method in which a model writes code on its own, receives scores based on outcomes such as whether tests pass, and improves from there. In other words, feedback drawn from real development environments was used as training material.
The story changes once you look at the full benchmark table
The problem is what it's being compared against. The benchmark table Microsoft published lines up only the previous model and lightweight models from Anthropic and OpenAI. Within that group, MAI Code 1.1 Flash edges ahead by a narrow margin. But once DeepSeek's open-source model, DeepSeek-V4-Flash-0731, is placed alongside the same table, the gap widens considerably.
| Benchmark | MAI Code 1.1 Flash | MAI Code 1-Flash (previous) | Claude Haiku 4.5 | GPT-5.4 mini | DeepSeek V4 Flash-0731 |
|---|---|---|---|---|---|
| SWE-bench Verified | 72.6% | 71.6% | 69.8% | 69.2% | Not disclosed |
| Terminal Bench 2.1 | 62.9% | 51.7% | 49.4% | 60.7% | 82.7% |
Terminal Bench 2.1 measures a model's ability to complete tasks by actually executing commands in a real terminal environment. In this category, DeepSeek came out nearly 20 percentage points ahead.
It falls behind on price too
It's not just performance. Microsoft's model is also more expensive per token. Some point out that when efficiency relative to usage is factored in, the real gap could be even larger than the numbers in the table suggest.
| Model | Input (per million tokens) | Cached input | Output |
|---|---|---|---|
| DeepSeek V4 Flash | $0.14 | $0.0028 | $0.28 |
| MAI Code 1.1 Flash | $0.20 | $0.02 | $1.20 |
| Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 |
The outlet The Decoder noted that Microsoft avoided this kind of direct comparison in its official announcement, instead emphasizing vague improvement metrics such as "a 4% increase in code survival rate" and "a 9% increase in return usage rate." It said the raw benchmark data was buried inside the model card.
Praising open source while avoiding it
DeepSeek V4 Flash is an open-weight model released on April 22. Anyone can download its weights to run on their own servers, or fine-tune it directly through supervised learning or preference optimization on platforms like Together AI.
Together AI begins supporting fine-tuning for DeepSeek V4 Flash 0731
Microsoft has recently sent several messages that seem to signal support for the OpenAI ecosystem. But its actions point in the opposite direction. Earlier, it reduced the share of OpenAI and Anthropic models in Copilot and made its own MAI models the default — a decision that chose lower cost over higher performance. The Decoder's view is that MAI Code 1.1 Flash follows the same pattern: pouring resources into an in-house model that lags on both performance and price, ultimately for the sake of margin management.
So what actually changes
For Copilot users, the app still lets them choose among multiple models for now. But most users simply stick with the default. If Microsoft locks in its own model as the default, it could create a structure where a model that lags in both performance and price still captures market share by default. With the cost-effectiveness of open-weight models already proven through benchmarks, how far Big Tech's calculus for pushing closed models can hold up remains to be seen — and will depend on how Copilot's pricing policy and default model settings evolve going forward.





Comments