A

Kairos-23M

online
mldi-lab/Kairos_23m

23M params | 2K context | $0.00025 per forecast

Kairos-23M is the mid-size public Kairos checkpoint from the Ant Group and ShanghaiTech University Kairos release. It sits in the middle of the Kairos size ladder, offering more capacity than the 10M model for broader zero-shot coverage while staying well below the 50M checkpoint.

Like the rest of the family it is an encoder-decoder transformer built around adaptive tokenization and instance-adaptive rotary position encoding, so it varies both patching granularity and positional treatment to match each series' information density and temporal structure. It is pretrained on the PreSTS corpus of 300B+ time points, the same source that gives the family its broad cross-domain zero-shot reach.

On TSFM.ai reach for Kairos-23M when the smallest Kairos-10M checkpoint leaves accuracy on the table but you do not want the full serving footprint of Kairos-50M. It is the natural default in the family when you want the adaptive design without committing to either extreme of the size ladder.

Model Classification

Family

Kairos

Type

time series foundation model

Pretrained time-series model exposed on TSFM.ai for zero-shot or few-shot forecasting workloads.

Training Data

PreSTS corpus with 300B+ time points, as documented by the official Kairos model cards and project page.

Recommended For

  • Adaptive zero-shot forecasting across heterogeneous series
  • Teams that want a mid-size open family with modern tokenization ideas

Strengths

  • Adaptive tokenization handles changing information density well
  • Clear parameter-size ladder from small to larger public checkpoints

Limitations

  • Newer family with a smaller production footprint than the most established lines
  • Focused on forecasting rather than general multi-task time-series tooling

Capabilities

forecastingzero-shotadaptive-tokenization

Tags

kairosadaptivezero-shot

Specifications

Parameters
23M
Architecture
encoder-decoder transformer with adaptive patching and instance-adaptive rotary position encoding
Context length
2,048
Max context
2,048
Minimum history
n/a
Recommended history
n/a
Input step
n/a
Required target series
1
Temperature
Ignored
Top P
Ignored
Max output
1,024
Avg latency
n/a
Uptime
n/a
Plan limits
1,000 rpm free · 1,000,000 rpm with billing
Accelerator
T4
Regions
Virginia, US
License
n/a

Pricing

Per forecast
$0.00025

Performance

Average latency
n/a
Availability
n/a
Plan limits
1,000 rpm free · 1,000,000 rpm with billing

Related models