METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㄱWords you meet while using AI

GIFT-Eval

A public testbed that compares time-series forecasting AI models on accuracy under the same conditions

In plain words

GIFT-Eval is a public testbed where AI models are given the same set of problems and graded on how well they predict data that changes over time, like sales, visitor numbers, or weather. Just like students from different academies all take the same mock exam and compare scores, forecasting models built by different companies are tested on identical data to see whose predictions come out on top.

There are two ways of scoring. One checks how precisely a model hits a single exact number. The other checks how well a model estimates the probability that the real value falls within a certain range. This reveals not just which model nails the average value, but which one gives reliable ranges even under uncertainty.

How it shows up in the news

In the article, Google Research tested its new forecasting model TimesFM-3 on three benchmarks—GIFT-Eval, FEV-Bench, and Time—and it ranked highest on all three. What's easy to misunderstand here is that GIFT-Eval was not built by Google. Google is simply a participant that used this testbed to get a report card for its own model; the testbed itself was created by someone else.

Try it yourself

If you open the publicly available GIFT-Eval benchmark page on Hugging Face, you can see a leaderboard where time-series forecasting models from various companies are scored on the same datasets. Clicking on a model's name in the leaderboard shows detailed breakdowns of what conditions it was tested under and what scores it received.

See also

Stories using this term

Browse every entry