METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Artificial Analysis swaps grading model in Intelligence Index v4.1.1

Patch release upgrades the grader model and incorporates the latest τ³-Banking version

Artificial Analysis swaps grading model in Intelligence Index v4.1.1

Summary

  • Artificial Analysis announced on August 6, 2026 that it had updated its Intelligence Index to v4.1.1.
  • The core of this patch is an upgrade to the grader model used for scoring and the incorporation of the latest τ³-Banking version.
  • The announcement did not include detailed score changes or shifts in model rankings, so further confirmation is needed.

The index moves up to v4.1.1

Artificial Analysis, an AI model evaluation organization, has updated its comprehensive evaluation metric, the Artificial Analysis Intelligence Index, to v4.1.1. The announcement was made via an X post on August 6, 2026, and was explicitly described as a patch release rather than a major overhaul.

The announcement cited two changes. One is an upgrade to the grader model used for scoring, and the other is the incorporation of the latest τ³-Banking version into the Artificial Analysis evaluation stack. The organization explained that this work is intended to keep the index "the most useful comprehensive metric."

Changes confirmed in this patch

ItemDetails
Versionv4.1.1
Release typePatch
Change 1Grader model upgrade
Change 2Incorporation of latest τ³-Banking version
Announcement channelX post
Announcement date2026-08-06

What changes when the grader model changes

The grader model is the entity that automatically judges whether a response is correct or assesses its quality. Therefore, when the grader is replaced, scoring outcomes can shift even on the same evaluation questions. However, this announcement did not include figures on which model is being used as the grader or how scores moved before and after the switch. In other words, how the index values or rankings of individual models actually changed cannot be confirmed from this disclosure alone.

For τ³-Banking as well, only the fact that the latest version was brought in was mentioned; the differences from the previous version and how the scoring reflects it were not explained in this post.

What remains unconfirmed

Version management of evaluation metrics is directly tied to benchmark reliability, but this announcement is closer to a summary-level notice. Detailed changelogs, the scope of items subject to re-grading, and comparison tables against previous v4.1 results need to be confirmed in separate documentation.

Other analyses on AI model evaluation methods and benchmark interpretation can be found in METAL LAB's benchmark-related articles.

Given that scores for the same model shift whenever the metric changes, it is safer not to directly compare leaderboard figures released after this patch with those based on the previous version.

Comments