METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Ant Group's Ling 3.0 Tiny scores 25 on intelligence index with 1.3B active parameters

Its accuracy was the lowest in its class, but its hallucination rate was also the lowest at 30%. The tradeoff: 50,000 tokens per answer.

Ant Group's Ling 3.0 Tiny scores 25 on intelligence index with 1.3B active parameters

Summary

  • Ant Group has released Ling 3.0 Tiny, an open-weight small model that scored 25 on the Artificial Analysis Intelligence Index.
  • With 7.9B total parameters and 1.3B active parameters, it sits on the Pareto frontier for "active parameters versus intelligence."
  • Its correct-answer rate was 9%, the lowest among the compared models, but its attempt rate was only 37%, giving it the lowest hallucination rate at 30% — while it produced the most output tokens per task, at 51k.

When it doesn't know, it doesn't answer

There's usually one way for a small model to beat a big one: don't pretend to know. Ling 3.0 Tiny, an open-weight small model from Ant Group, only attempted answers on 37% of knowledge questions. The rest it simply skipped. As a result, its hallucination rate — the rate at which it gives wrong answers — came in at 30%, the lowest among the 12 models compared.

The model's Intelligence Index score, as measured by independent evaluator Artificial Analysis, is 25. Total parameters stand at 7.9B, but the active parameters actually engaged when processing a single query are just 1.3B. Artificial Analysis assessed that the model sits on the Pareto frontier for "active parameters versus intelligence" — meaning no smarter model of this size currently exists.

여러 AI 모델의 AA-Omniscience 지수, 정확도, 환각률을 막대그래프로 비교한 차트
이미지: @ArtificialAnlys (X)

How small is 1.3B active parameters, really

MoE (mixture-of-experts) architectures don't activate the entire model — only the portions needed for a given task. That's why total parameters can be large while active parameters stay small. For comparison, NVIDIA's Nemotron 3.5 Lightning has 30B total parameters with 3B active, while Ant Group's higher-tier Ling 3.0 Flash, released August 7-8, has 124B total. Ling 3.0 Tiny's 1.3B stands out as notably small even within this family. Smaller models consume less memory, and that opens up the possibility of running them on a laptop or a single GPU.

How can it rank last on accuracy but high on the index

The AA-Omniscience Index measures the reliability of a model's knowledge. Correct answers earn points, wrong answers lose points, and declining to answer with "I don't know" incurs no penalty. The scale runs from -100 to 100, where 0 means correct and incorrect answers are equal in number, and negative scores mean wrong answers outnumber right ones. Small models, which simply hold less knowledge to begin with, tend to stay in negative territory.

ModelOmniscience IndexCorrect-answer rateHallucination rate
Nemotron 3.5 Lightning-1814%38%
Ling 3.0 Tiny-199%30%
Qwen3.6 35B A3B-2219%51%
Muse Glimmer (high)-3327%82%
Ministral 3 8B-6913%94%

Reading the table across reveals the pattern. Muse Glimmer has the highest correct-answer rate at 27%, but also an 82% hallucination rate — because it answers even what it doesn't know. Ling 3.0 Tiny has the lowest correct-answer rate on the list at 9%, yet its index score of -19 ranks second. Artificial Analysis attributes this to "the low attempt rate reducing hallucinations."

This distinction matters in practice. For tasks like document summarization or internal search — where a single wrong fact can ruin the entire output — a model that stays silent when unsure is easier to work with than one that gets half its answers right by guessing at the rest. On the other hand, it's poorly suited to chatbots that need to field broad, general-knowledge questions.

AI 모델별 지능지수 작업당 출력 토큰 수를 답변과 추론으로 나누어 보여주는 막대그래프
이미지: @ArtificialAnlys (X)

The cost comes in tokens

None of this is free. Ling 3.0 Tiny used an average of 51k output tokens to solve a single Intelligence Index task. Of that, 32k were reasoning tokens spent working through the problem before producing an answer.

ModelOutput tokens per task
Ling 3.0 Tiny51k100
Ling 3.0 Flash36k71
Nemotron 3.5 Lightning31k61
Qwen3.6 35B A3B31k61
Gemma 4 31B5k10

According to Artificial Analysis's figures, that's roughly 65% more than Nemotron 3.5 Lightning or Qwen3.6 35B A3B. In effect, the model compensates for its reduced parameter count by "thinking" at greater length. Because the model is small, it generates each individual token quickly — but when the total number of tokens needed is high, both perceived latency and cost rise accordingly. This is why parameter count alone shouldn't be the deciding factor when choosing a small reasoning model.

다양한 AI 모델의 여러 지능 평가 항목별 성능을 막대그래프로 비교한 종합 차트
이미지: @ArtificialAnlys (X)

So what does this actually change

Small open-weight models have been arriving in a rush over the past few weeks. NVIDIA's Nemotron 3.5 Lightning scored 24 on the Intelligence Index, Ling 3.0 Flash scored 38, and now Ling 3.0 Tiny adds a score of 25. Looking at the scores alone, they're all in the same ballpark. What sets them apart is character. Some models compete on speed, others on breadth of knowledge — Ling 3.0 Tiny has carved out its niche as the one that "says fewer wrong things."

Being open-weight, its weights can be downloaded and run independently. For use cases like internal document Q&A — where giving a wrong answer is riskier than giving no answer at all — this adds one more option worth testing on your own servers. Still, whether the low attempt rate seen in benchmarks translates into an overuse of "I don't know" in real-world deployment is something you'll only find out by running it yourself. It's not a call to make from a single benchmark figure.

Comments