
Summary
- Ant Group has released Ling 3.0 Tiny, an open-weight small model that scored 25 on the Artificial Analysis Intelligence Index.
- With 7.9B total parameters and 1.3B active parameters, it sits on the Pareto frontier for "active parameters versus intelligence."
- Its correct-answer rate was 9%, the lowest among the compared models, but its attempt rate was only 37%, giving it the lowest hallucination rate at 30% — while it produced the most output tokens per task, at 51k.
When it doesn't know, it doesn't answer
There's usually one way for a small model to beat a big one: don't pretend to know. Ling 3.0 Tiny, an open-weight small model from Ant Group, only attempted answers on 37% of knowledge questions. The rest it simply skipped. As a result, its hallucination rate — the rate at which it gives wrong answers — came in at 30%, the lowest among the 12 models compared.
The model's Intelligence Index score, as measured by independent evaluator Artificial Analysis, is 25. Total parameters stand at 7.9B, but the active parameters actually engaged when processing a single query are just 1.3B. Artificial Analysis assessed that the model sits on the Pareto frontier for "active parameters versus intelligence" — meaning no smarter model of this size currently exists.

How small is 1.3B active parameters, really
MoE (mixture-of-experts) architectures don't activate the entire model — only the portions needed for a given task. That's why total parameters can be large while active parameters stay small. For comparison, NVIDIA's Nemotron 3.5 Lightning has 30B total parameters with 3B active, while Ant Group's higher-tier Ling 3.0 Flash, released August 7-8, has 124B total. Ling 3.0 Tiny's 1.3B stands out as notably small even within this family. Smaller models consume less memory, and that opens up the possibility of running them on a laptop or a single GPU.
How can it rank last on accuracy but high on the index
The AA-Omniscience Index measures the reliability of a model's knowledge. Correct answers earn points, wrong answers lose points, and declining to answer with "I don't know" incurs no penalty. The scale runs from -100 to 100, where 0 means correct and incorrect answers are equal in number, and negative scores mean wrong answers outnumber right ones. Small models, which simply hold less knowledge to begin with, tend to stay in negative territory.
| Model | Omniscience Index | Correct-answer rate | Hallucination rate |
|---|---|---|---|
| Nemotron 3.5 Lightning | -18 | 14% | 38% |
| Ling 3.0 Tiny | -19 | 9% | 30% |
| Qwen3.6 35B A3B | -22 | 19% | 51% |
| Muse Glimmer (high) | -33 | 27% | 82% |
| Ministral 3 8B | -69 | 13% | 94% |
Reading the table across reveals the pattern. Muse Glimmer has the highest correct-answer rate at 27%, but also an 82% hallucination rate — because it answers even what it doesn't know. Ling 3.0 Tiny has the lowest correct-answer rate on the list at 9%, yet its index score of -19 ranks second. Artificial Analysis attributes this to "the low attempt rate reducing hallucinations."
This distinction matters in practice. For tasks like document summarization or internal search — where a single wrong fact can ruin the entire output — a model that stays silent when unsure is easier to work with than one that gets half its answers right by guessing at the rest. On the other hand, it's poorly suited to chatbots that need to field broad, general-knowledge questions.

The cost comes in tokens
None of this is free. Ling 3.0 Tiny used an average of 51k output tokens to solve a single Intelligence Index task. Of that, 32k were reasoning tokens spent working through the problem before producing an answer.
| Model | Output tokens per task | |
|---|---|---|
| Ling 3.0 Tiny | 51k | 100 |
| Ling 3.0 Flash | 36k | 71 |
| Nemotron 3.5 Lightning | 31k | 61 |
| Qwen3.6 35B A3B | 31k | 61 |
| Gemma 4 31B | 5k | 10 |
According to Artificial Analysis's figures, that's roughly 65% more than Nemotron 3.5 Lightning or Qwen3.6 35B A3B. In effect, the model compensates for its reduced parameter count by "thinking" at greater length. Because the model is small, it generates each individual token quickly — but when the total number of tokens needed is high, both perceived latency and cost rise accordingly. This is why parameter count alone shouldn't be the deciding factor when choosing a small reasoning model.

So what does this actually change
Small open-weight models have been arriving in a rush over the past few weeks. NVIDIA's Nemotron 3.5 Lightning scored 24 on the Intelligence Index, Ling 3.0 Flash scored 38, and now Ling 3.0 Tiny adds a score of 25. Looking at the scores alone, they're all in the same ballpark. What sets them apart is character. Some models compete on speed, others on breadth of knowledge — Ling 3.0 Tiny has carved out its niche as the one that "says fewer wrong things."
Being open-weight, its weights can be downloaded and run independently. For use cases like internal document Q&A — where giving a wrong answer is riskier than giving no answer at all — this adds one more option worth testing on your own servers. Still, whether the low attempt rate seen in benchmarks translates into an overuse of "I don't know" in real-world deployment is something you'll only find out by running it yourself. It's not a call to make from a single benchmark figure.





Comments