Live·112 models · 112 evaluated·Auto-refreshed from the official leaderboard every 12 hours

GIFT-Eval Leaderboard

GIFT-Eval expands the question beyond raw point accuracy. It looks at how models perform across many datasets, frequencies, and forecasting settings, making it a strong benchmark for teams that care about robustness rather than a single flattering leaderboard slice. This page mirrors the official Salesforce GIFT-Eval leaderboard with search, filters, and per-slice rankings — covering every submitted model, not just the top few — and links each result back to a hosted endpoint on TSFM.ai where available.

CastStar ranks first with an average MASE rank of 16.44 across the GIFT-Eval dataset surface.

What this benchmark answers

Which models stay strong across heterogeneous datasets and probabilistic settings?

Methodology

Models are scored on grouped benchmark slices and ranked by average rank, with Weighted Quantile Loss providing a secondary read on probabilistic accuracy.

GIFT-Eval leaderboard

Showing 112 of 112 models · 38 hosted on TSFM.ai · Sorted by overall MASE rank (lower is better)

# Model MASE rank
1
CastStarHF

USTC-AGI

16.44
2
Falcon-Agent

ant-intl

16.67
3
RacineCast-1-NCHFCode

Racine.ai

18.68
4
RacineCast-1HFCode

Racine.ai

18.72
5
Cobra-AgentHF

Dalpha AI

20.80
6
Toto-2.0-FnFHFCode

Datadog

22.25
7
PrismHF

Birla AI Labs

22.79
8
RAES-Conductance-EnsembleHFCode

California State University Northridge

22.97
9
HistRoute-CV

Hitsz iLearn-ML

23.79
10
Taichu-TimeSeries-AgentHFCode

zidongtaichu

24.03
11
metis-autocastHF

Metis

24.51
12
Toto-2.0-2.5B-FTHFCode

Datadog

25.26
13
TSOrchestraHFCode

Melady Lab @ USC

25.33
14
DeOSAlphaTimeGPTPredictor-2025HF

vencortex®

26.19
15
TimeRouterHFCode

UConn & Morgan Stanley

27.48
16
RAES-Conductance-Ensemble-VHFCode

SETI

27.54
17
MoiraiAgent-leakingHFCode

Salesforce AI Research

28.13
18
MoiraiAgentHFCode

Salesforce AI Research

30.32
19
CredenceHFCode

ContinualIST

30.43
20
Falcon-2.0HFCode

ant-intl

31.60
21
STRIDE (+Chronos-2)HF

Google Cloud AI Research

33.10
22
Samay

Kairosity

33.29
23
TiRex-2-PretrainedHFCode

NXAI

33.66
2434.31
25
ZooCast-Top1Code
34.42
26
Migas-1.0HF

Synthefy

34.46
27
STRIDE (+Timer-S1)HF

Google Cloud AI Research

34.88
28
TimeCopilotHFCode
35.29
29
SynapseHF

Google Cloud AI Research

35.58
3036.19
3137.36
3237.59
33
TurkForecast-FM-Chronos2-LoRA-v1HFCode

TurkForecast (Mert Karatay)

38.68
34
EFG-base

LAIR

39.49
35
TSOrchestra-testHF

Melady Lab @ USC

39.74
36
TiRex-2-ZeroshotHFCode

NXAI

40.40
37
ZeusHFCode

GestaltCog Lab @ ICT

40.81
38
IBM logo

IBM TSFM & Rensselaer Polytechnic Institute

41.26
39
Granite-FlowState-r1.1HFCode

IBM TSFM

41.55
40
Tsinghua University logo

Tsinghua & ByteDance

41.61
41
ValBestSingle-cmttHFCode

N/A

41.78
42
Google logo

Google Research

42.11
4342.43
44
LongSeer-v1.0

LongShine AI Research

44.45
45
Falcon-XHFCode

ant-intl

45.40
4646.18
47
IBM logo

IBM TSFM & Rensselaer Polytechnic Institute

47.00
48
Reverso

MIT

47.16
4947.76
50
Xihe-ultraHF

Ant

48.53
5148.61
5249.44
53
VISIT-2.0
49.91
54
t0-alphaHFCode

The Forecasting Company

50.95
55
TEMPO_ENSEMBLEHF

Melady Lab @ USC

51.66
56
Xihe-maxHF

Ant

51.71
5753.35
58
FlowState-9.1MHFCode

IBM Research

53.63
59
Reverso-SmallHFCode

MIT

53.86
60
A

ShanghaiTech University

55.56
61
Salesforce logo

Salesforce AI Research

57.31
6259.58
63
A

ShanghaiTech University

59.99
6460.54
65
A

ShanghaiTech University

60.71
6660.76
67
CHARMHFCode

C3 AI

61.21
6861.40
69
Google logo

Google Research

62.08
70
Tsinghua University logo

Tsinghua University

64.46
71
xLSTM-MixerHF

AIML Lab @ TU Darmstadt

66.41
72
CleanTS-65MHFCode

Shandong University

67.31
73
Reverso-Nano

MIT

67.40
7467.66
75
PatchFMHFCode

LITIS

69.54
7670.00
7770.74
78
TabPFN-TSHFCode

PriorLabs

70.86
79
TempoPFNHFCode

University of Freiburg

71.00
8074.00
81
Lingjiang

Alibaba Cloud

74.70
82
Salesforce logo

Salesforce AI Research

75.30
83
Salesforce logo

Salesforce AI Research

76.16
84
Metamorph1.0

SRI International

76.25
85
LiteSpecFormerHFCode

FlowVortex (Xidian University)

76.65
8678.65
87
Chronos_baseHFCode

AWS AI Labs

79.31
88
IBM logo

Princeton University

81.90
89
FLAIRHFCode

Mellon Inc.

82.05
90
Chronos_smallHFCode

AWS AI Labs

82.44
91
i_transformer

Tsinghua University

83.87
9284.69
93
Super-LinearHFCode

Ben-Gurion University of the Negev

84.96
94
Google logo

Google Research

86.64
95
TFT

Google Research

86.92
96
VisionTSHF

Zhejiang University

88.77
97
Salesforce logo

Salesforce AI Research

88.84
98
N-BEATS

ServiceNow

89.67
99

IBM Research

90.09
100
Auto_Arima
91.51
10192.85
102
Auto_ETS
94.80
10394.88
104
Auto_Theta
95.09
105
Seasonal_NaiveHFCode
95.61
106
Crossformer

Shanghai Jiao Tong University

95.81
107
DLinear

The Chinese University of Hong Kong

96.70
108
DeepAR

Amazon Research

97.92
109
TIDE

Google Research

97.94
110
ServiceNow logo

Morgan Stanley & Service Now

101.5
111
NaiveHFCode
101.7
112
VISIT-1.0
101.8
112 models · Aggregated live from the official GIFT-Eval leaderboard. Refreshes every 12 hours.Last refreshed Jul 29, 2026, 6:05 AM

Model landscape

GIFT-Eval is not only a TSFM leaderboard — it includes classical statistical baselines (ARIMA, ETS, Theta), deep-learning architectures (PatchTST, iTransformer, TFT), and agentic systems, so foundation-model results can be compared against the full prior art.

Pretrained
31(28%)
Zero-shot
37(33%)
Fine-tuned
6(5%)
Agentic
22(20%)
Deep learning
10(9%)
Statistical
6(5%)

Why robustness across datasets matters

A model that tops one leaderboard can collapse on a different frequency or domain. GIFT-Eval forces models to prove themselves across 23 grouped dataset slices covering different frequencies, horizons, and series shapes. If a model ranks well here, you can be more confident it will not surprise you when your data does not look like the training distribution.

Understanding Weighted Quantile Loss

WQL penalizes both overconfident and underconfident prediction intervals. A model with a low WQL produces forecast distributions that are well-calibrated — the 90th percentile prediction actually lands above the true value about 90% of the time. This matters for capacity planning, inventory, and any decision that depends on reliable uncertainty estimates rather than just the median forecast.

How to interpret it

  • Lower Average Rank is better because the leaderboard aggregates placement across many slices.
  • WQL matters when forecast calibration and uncertainty quality matter to the business.
  • Use GIFT-Eval to sanity-check whether a model is robust beyond a single narrow domain.

Frequently asked questions

What is GIFT-Eval?
GIFT-Eval (General Time Series Forecasting Model Evaluation) is a probabilistic forecasting benchmark from Salesforce that evaluates time series foundation models across 23 diverse dataset groups spanning 7 domains and 10 frequencies. Models are ranked by Average Rank and Average Weighted Quantile Loss.
What does Average Rank mean in GIFT-Eval?
Average Rank is the mean placement of a model across every benchmark slice. A lower value means the model consistently finishes near the top across heterogeneous datasets rather than dominating only one slice.
When should I use GIFT-Eval instead of FEV Bench?
Use GIFT-Eval when you care about probabilistic forecast quality (calibrated uncertainty) and robustness across many different data domains, rather than just point-forecast accuracy on a single leaderboard.
Does GIFT-Eval test multivariate forecasting?
GIFT-Eval includes both univariate and multivariate slices. The leaderboard on TSFM.ai aggregates both variate types from the upstream grouped-by-univariate file, and you can filter per-slice rankings to inspect multivariate-only performance.
What is the 'test leak' column?
Some GIFT-Eval submissions are from models whose training data overlaps with the evaluation datasets. A 'Yes' in the test-leak column means the authors disclosed partial or full pretraining-data overlap; use the filter to compare only models with no known test-data leakage.
What model types appear on the leaderboard?
GIFT-Eval ranks pretrained foundation models, zero-shot models, fine-tuned models, agentic systems, classical deep-learning architectures (PatchTST, iTransformer, TFT, TiDE, N-BEATS), and statistical baselines (ARIMA, ETS, Theta). The page shows every type so you can benchmark a TSFM against classical prior art.
How often is the leaderboard refreshed?
TSFM.ai refetches the upstream Salesforce GIFT-Eval results every 12 hours via Next.js ISR. The 'last refreshed' timestamp at the bottom of the leaderboard reflects the most recent successful refresh.

Related reading

Compare with other TSFM benchmarks

FEV Bench

How well does a model generalize to unseen real-world forecasting tasks?

BOOM

How do models behave on observability telemetry instead of academic datasets?

Impermanent

Does model performance hold up as real time passes and the data distribution shifts?

ARFBench

Which multimodal models can reason about anomalies, timing, magnitude, and cross-series structure in production telemetry?

Sources