Live·130 models · 130 evaluated·Auto-refreshed from the official leaderboard every 12 hours

GIFT-Eval Leaderboard

GIFT-Eval expands the question beyond raw point accuracy. It looks at how models perform across many datasets, frequencies, and forecasting settings, making it a strong benchmark for teams that care about robustness rather than a single flattering leaderboard slice. This page mirrors the official Salesforce GIFT-Eval leaderboard with search, filters, and per-slice rankings — covering every submitted model, not just the top few — and links each result back to a hosted endpoint on TSFM.ai where available.

STRIDE w/ Synapse ranks first with an average MASE rank of 14.62 across the GIFT-Eval dataset surface.

What this benchmark answers

Which models stay strong across heterogeneous datasets and probabilistic settings?

Methodology

Models are scored on grouped benchmark slices and ranked by average rank, with Weighted Quantile Loss providing a secondary read on probabilistic accuracy.

GIFT-Eval leaderboard

Showing 130 of 130 models · 38 hosted on TSFM.ai · Sorted by overall MASE rank (lower is better)

# Model MASE rank
1
STRIDE w/ SynapseHF

Google Cloud AI Research

14.62
2
EXAONE-Forecast-Agent

LG AI Research

18.87
3
LS-MoE

LongShine AI Research

19.31
4
LS-Agent

LongShine AI Research

21.25
5
Falcon-Agent

ant-intl

22.15
6
CastStarHF

USTC-AGI

22.23
7
limix_moe
22.24
8
TIMEHEDGE

Neurogica Inc.

23.66
9
RegLab-AH-1

RegLab

24.31
10
RacineCast-1-NCHFCode

Racine.ai

25.64
11
RacineCast-1HFCode

Racine.ai

25.75
12
FastCastrage

sme

27.27
13
TimesFM-3HFCode

Google Research

27.70
14
Cobra-AgentHF

Dalpha AI

28.60
15
Toto-2.0-FnFHFCode

Datadog

29.41
16
RAES-Conductance-EnsembleHFCode

California State University Northridge

30.47
17
PrismHF

Birla AI Labs

31.19
18
EXAONE-Forecast

LG AI Research

31.61
19
HistRoute-CV

Hitsz iLearn-ML

32.00
20
Taichu-TimeSeries-AgentHFCode

zidongtaichu

32.52
21
metis-autocastHF

Metis

32.78
22
Toto-2.0-2.5B-FTHFCode

Datadog

33.26
23
TSOrchestraHFCode

Melady Lab @ USC

33.55
24
DeOSAlphaTimeGPTPredictor-2025HF

vencortex®

34.54
25
ForecastMate

Zhejiang University

34.73
26
RAES-Conductance-Ensemble-VHFCode

SETI

35.95
27
TimeRouterHFCode

UConn & Morgan Stanley

36.25
28
MoiraiAgent-leakingHFCode

Salesforce AI Research

37.40
29
Falcon-2.0HFCode

ant-intl

39.61
30
Granite-PatchTST-FM-r2HFCode

IBM TSFM & Rensselaer Polytechnic Institute

39.77
31
MoiraiAgentHFCode

Salesforce AI Research

40.08
32
CredenceHFCode

ContinualIST

40.14
33
STRIDE (+Chronos-2)HF

Google Cloud AI Research

41.22
34
Samay

Kairosity

43.30
35
TiRex-2-PretrainedHFCode

NXAI

43.72
3643.91
37
STRIDE (+Timer-S1)HF

Google Cloud AI Research

44.20
38
Migas-1.0HF

Synthefy

44.68
39
ZooCast-Top1Code
44.69
40
TimeCopilotHFCode
45.89
41
SynapseHF

Google Cloud AI Research

45.99
4246.46
4347.63
44
TurkForecast-FM-Chronos2-LoRA-v1HFCode

TurkForecast (Mert Karatay)

47.84
4548.03
46
EFG-base

LAIR

49.96
47
TSOrchestra-testHF

Melady Lab @ USC

50.39
48
ZeusHFCode

GestaltCog Lab @ ICT

51.15
49
tafsutHF

Tafsut-FM (Huawei GTS x EURECOM)

51.16
50
TiRex-2-ZeroshotHFCode

NXAI

51.25
51
Granite-FlowState-r1.1HFCode

IBM TSFM

52.09
52
IBM logo

IBM TSFM & Rensselaer Polytechnic Institute

52.26
53
Tsinghua University logo

Tsinghua & ByteDance

52.59
54
ValBestSingle-cmttHFCode

N/A

52.94
5553.05
56
Google logo

Google Research

53.12
57
Falcon-XHFCode

ant-intl

55.86
58
LongSeer-v1.0

LongShine AI Research

56.20
5957.84
60
IBM logo

IBM TSFM & Rensselaer Polytechnic Institute

58.58
61
Reverso

MIT

59.12
6259.53
63
Xihe-ultraHF

Ant

60.11
6460.47
65
TEMPO_ENSEMBLEHF

Melady Lab @ USC

61.21
6661.60
67
VISIT-2.0
62.02
68
t0-alphaHFCode

The Forecasting Company

63.36
69
Xihe-maxHF

Ant

63.84
7065.88
71
FlowState-9.1MHFCode

IBM Research

66.02
72
Reverso-SmallHFCode

MIT

66.54
73
A

ShanghaiTech University

68.02
74
Salesforce logo

Salesforce AI Research

70.04
7572.82
76
A

ShanghaiTech University

73.02
7773.65
78
A

ShanghaiTech University

74.03
7974.24
80
CHARMHFCode

C3 AI

74.78
81
Google logo

Google Research

74.84
8274.95
83
Tsinghua University logo

Tsinghua University

77.97
84
xLSTM-MixerHF

AIML Lab @ TU Darmstadt

80.15
85
recursive-moirai-2HFCode

Emilio Cantu

80.66
86
CleanTS-65MHFCode

Shandong University

81.36
8781.38
88
Reverso-Nano

MIT

81.71
89
PatchFMHFCode

LITIS

83.74
90
goia-forecast-nano-v0HFCode

Gredio

83.77
9184.64
9285.33
93
TabPFN-TSHFCode

PriorLabs

85.42
94
TempoPFNHFCode

University of Freiburg

85.72
95
TinyCastHFCode

RAWS Labs

86.52
96
Metamorph1.0-4.5M

SRI International

87.38
97
Lingjiang

Alibaba Cloud

88.91
9889.06
99
Salesforce logo

Salesforce AI Research

90.19
100
Salesforce logo

Salesforce AI Research

91.29
101
Metamorph1.0

SRI International

91.67
102
LiteSpecFormerHFCode

FlowVortex (Xidian University)

91.82
10394.04
104
TimeTron-v2-33MHF

Corteri Intelligence

94.91
105
Chronos_baseHFCode

AWS AI Labs

94.97
106
IBM logo

Princeton University

97.82
107
FLAIRHFCode

Mellon Inc.

98.04
108
Chronos_smallHFCode

AWS AI Labs

98.34
109
i_transformer

Tsinghua University

99.93
110101.1
111
Super-LinearHFCode

Ben-Gurion University of the Negev

101.2
112
Google logo

Google Research

103.0
113
TFT

Google Research

103.3
114
VisionTSHF

Zhejiang University

105.5
115
Salesforce logo

Salesforce AI Research

105.8
116
N-BEATS

ServiceNow

106.3
117

IBM Research

106.9
118
Auto_Arima
108.2
119109.9
120
Auto_ETS
111.5
121
Auto_Theta
111.8
122112.0
123
Crossformer

Shanghai Jiao Tong University

112.0
124
Seasonal_NaiveHFCode
112.8
125
DLinear

The Chinese University of Hong Kong

114.1
126
DeepAR

Amazon Research

115.1
127
TIDE

Google Research

115.3
128
ServiceNow logo

Morgan Stanley & Service Now

119.2
129
NaiveHFCode
119.2
130
VISIT-1.0
119.4
130 models · Aggregated live from the official GIFT-Eval leaderboard. Refreshes every 12 hours.Last refreshed Sep 12, 2026, 9:20 AM

Model landscape

GIFT-Eval is not only a TSFM leaderboard — it includes classical statistical baselines (ARIMA, ETS, Theta), deep-learning architectures (PatchTST, iTransformer, TFT), and agentic systems, so foundation-model results can be compared against the full prior art.

Pretrained
35(27%)
Zero-shot
42(32%)
Fine-tuned
6(5%)
Agentic
31(24%)
Deep learning
10(8%)
Statistical
6(5%)

Why robustness across datasets matters

A model that tops one leaderboard can collapse on a different frequency or domain. GIFT-Eval forces models to prove themselves across 23 grouped dataset slices covering different frequencies, horizons, and series shapes. If a model ranks well here, you can be more confident it will not surprise you when your data does not look like the training distribution.

Understanding Weighted Quantile Loss

WQL penalizes both overconfident and underconfident prediction intervals. A model with a low WQL produces forecast distributions that are well-calibrated — the 90th percentile prediction actually lands above the true value about 90% of the time. This matters for capacity planning, inventory, and any decision that depends on reliable uncertainty estimates rather than just the median forecast.

How to interpret it

  • Lower Average Rank is better because the leaderboard aggregates placement across many slices.
  • WQL matters when forecast calibration and uncertainty quality matter to the business.
  • Use GIFT-Eval to sanity-check whether a model is robust beyond a single narrow domain.

Frequently asked questions

What is GIFT-Eval?
GIFT-Eval (General Time Series Forecasting Model Evaluation) is a probabilistic forecasting benchmark from Salesforce that evaluates time series foundation models across 23 diverse dataset groups spanning 7 domains and 10 frequencies. Models are ranked by Average Rank and Average Weighted Quantile Loss.
What does Average Rank mean in GIFT-Eval?
Average Rank is the mean placement of a model across every benchmark slice. A lower value means the model consistently finishes near the top across heterogeneous datasets rather than dominating only one slice.
When should I use GIFT-Eval instead of FEV Bench?
Use GIFT-Eval when you care about probabilistic forecast quality (calibrated uncertainty) and robustness across many different data domains, rather than just point-forecast accuracy on a single leaderboard.
Does GIFT-Eval test multivariate forecasting?
GIFT-Eval includes both univariate and multivariate slices. The leaderboard on TSFM.ai aggregates both variate types from the upstream grouped-by-univariate file, and you can filter per-slice rankings to inspect multivariate-only performance.
What is the 'test leak' column?
Some GIFT-Eval submissions are from models whose training data overlaps with the evaluation datasets. A 'Yes' in the test-leak column means the authors disclosed partial or full pretraining-data overlap; use the filter to compare only models with no known test-data leakage.
What model types appear on the leaderboard?
GIFT-Eval ranks pretrained foundation models, zero-shot models, fine-tuned models, agentic systems, classical deep-learning architectures (PatchTST, iTransformer, TFT, TiDE, N-BEATS), and statistical baselines (ARIMA, ETS, Theta). The page shows every type so you can benchmark a TSFM against classical prior art.
How often is the leaderboard refreshed?
TSFM.ai refetches the upstream Salesforce GIFT-Eval results every 12 hours via Next.js ISR. The 'last refreshed' timestamp at the bottom of the leaderboard reflects the most recent successful refresh.

Related reading

Compare with other TSFM benchmarks

FEV Bench

How well does a model generalize to unseen real-world forecasting tasks?

BOOM

How do models behave on observability telemetry instead of academic datasets?

Impermanent

Does model performance hold up as real time passes and the data distribution shifts?

ARFBench

Which multimodal models can reason about anomalies, timing, magnitude, and cross-series structure in production telemetry?

Sources