Moirai-1.0-R-Large
onlineSalesforce/moirai-1.0-R-large311M params | 512 context | $0.00025 per forecast
Moirai-1.0-R-Large is the highest-capacity dense checkpoint in Salesforce's original Moirai family, the top tier above Moirai-1.0-R-Small and Moirai-1.0-R-Base. At 311M parameters it is the first-generation Moirai design at its largest published scale, aimed at the strongest broad zero-shot performance the dense 1.0-R line offers.
It preserves the same masked-encoder transformer as the smaller checkpoints, with multi-patch projections, any-variate attention, and a mixture-distribution output head. The any-variate attention handles arbitrary numbers of target variables and dynamic covariates in one model, and the mixture output head yields probabilistic forecasts with calibrated quantiles. Like the rest of the family it is pretrained on the LOTSA corpus, the large open time-series archive behind the Moirai work, which underpins its zero-shot generalization.
On TSFM.ai choose it when accuracy matters more than serving cost and you want the original Moirai architecture at full capacity. If latency or throughput is the constraint, step down to Moirai-1.0-R-Base or Moirai-1.0-R-Small; if your data is dominated by low-frequency yearly and quarterly series, the updated 1.1-R checkpoints are the better fit.
Model Classification
Family
Moirai
Type
time series foundation model
Pretrained time-series model exposed on TSFM.ai for zero-shot or few-shot forecasting workloads.
Resources
Training Data
LOTSA, the Large-scale Open Time Series Archive, with roughly 27B observations across nine domains including energy, transport, finance, healthcare, sales, climate, web, and social data.
Recommended For
- • Multivariate forecasting across heterogeneous domains
- • Workloads that benefit from probabilistic outputs and arbitrary variate counts
Strengths
- • Strong multivariate coverage across the Moirai family
- • Well-suited to covariates and correlated series
Limitations
- • Model cards for some newer Moirai variants are still sparse on exact checkpoint details
- • Heavier family choices can be more expensive than tiny single-purpose baselines
Capabilities
Tags
Specifications
- Parameters
- 311M
- Architecture
- masked encoder transformer with multi-patch projections, any-variate attention, and mixture-distribution output
- Context length
- 512
- Max context
- 8,192
- Minimum history
- n/a
- Recommended history
- n/a
- Input step
- n/a
- Required target series
- 1
- Temperature
- Ignored
- Top P
- Ignored
- Max output
- 2,048
- Avg latency
- n/a
- Uptime
- n/a
- Plan limits
- 1,000 rpm free · 1,000,000 rpm with billing
- Accelerator
- L40S
- Regions
- Virginia, US
- License
- n/a
Pricing
- Per forecast
- $0.00025
Performance
- Average latency
- n/a
- Availability
- n/a
- Plan limits
- 1,000 rpm free · 1,000,000 rpm with billing