Financial AI Lab at Chicago Booth

Towards Financial World Modeling

A trillion-observation market dataset and systematic study of representation learning in financial markets.

Humzah MerchantAlec GuthrieSimon MahnsRandall BalestrieroBradford Levy

University of Chicago, Booth School of Business

A world model plans from a state representation, often for tasks it was never trained on. In markets, that state has to carry expected returns, liquidity, volatility, market-wide conditions and cross-asset relationships. Yet financial representation learning has mostly been judged on one forecasting task over one slice of time. We release Market-1T, nearly one trillion one-second observations of U.S. equities from 2008 to 2025, and use it to compare 18 encoder-training strategies across 32 months of market regimes, as the state layer for world models such as DINO-WM,[1] V-JEPA 2[2] and LeWM.[3]

Three augmentations, how each ranks among 14 SSL methods on clustering and forecasting, and supervised scaling curvesThree augmentations, how each ranks among 14 SSL methods on clustering and forecasting, and supervised scaling curves

Matching views of the same stock clusters latents by firm and hardly at all by day. At a fixed compute budget, ViT-Tiny beats larger ViTs.

Market-1T

A best-in-class academic dataset spanning 2008–2025 at 1 Hz frequency, released in partnership with Massive.

Monthly rank IC of one encoder recipe on return, volatility change and spread change, 2008 to 2025Monthly rank IC of one encoder recipe on return, volatility change and spread change, 2008 to 2025

One recipe, retrained on the six months before each month and evaluated on it, across the full span. The switch picks out our 32 evaluation months.

Rigorous Evaluation Protocol

Every method is evaluated on the same 32 months and trained only on the six months before each one. We report rank IC on 15-minute return, volatility change and spread change.

Annualized Sharpe of a long-short book versus EFQ for three encoders, exiting after 15 minutes or at the closeAnnualized Sharpe of a long-short book versus EFQ for three encoders, exiting after 15 minutes or at the close

A Sharpe ratio depends not only on the forecast but on how return forecasts are translated into portfolios and trades, and on the cost and execution model.

Rank IC for 10 seeds in each of 10 evaluation months, as box plots per monthRank IC for 10 seeds in each of 10 evaluation months, as box plots per month

Ten seeds in each of ten months. The month explains 97.2%, 98.9% and 99.9% of the variance; the seed almost none. A single evaluation period says little about a method.

Rank IC by month for the trained seeds beside the untrained Random ViTRank IC by month for the trained seeds beside the untrained Random ViT

The untrained Random ViT rises and falls with the same months, so most of a month's difficulty is shared by every encoder.

Rank IC minus the Random ViT mean, by month, for ten seeds in each monthRank IC minus the Random ViT mean, by month, for ten seeds in each month

Subtracting each month's Random ViT mean removes much of the month-to-month swing. Our standard errors are computed this way.

Rank IC relative to a retrained model versus months between training and evaluationRank IC relative to a retrained model versus months between training and evaluation

Each model is compared with the same model retrained just before the evaluation month, to control for the intrinsic difficulty of the tasks changing over time. Return holds up; volatility and spread lose about 20% and 10% of their IC over five years.

Rank IC versus training FLOPs for ViT-Tiny, ViT-Small and ViT-BaseRank IC versus training FLOPs for ViT-Tiny, ViT-Small and ViT-Base

At fixed training FLOPs, smaller encoders win on every task: ViT-Tiny beats a ViT-Base trained with 10× the compute. Every point is an annealed WSD[4] checkpoint.

Self-Supervised Learning

We test different data augmentation strategies and pretraining objectives.

Rank IC averaged over 32 months. A supervised arm's own head IC is in parentheses. Red fails to beat the Random ViT; green marks the best two non-supervised arms.
EncoderReturnVolatilitySpread
Supervised
Return0.0272 (0.0269)0.06950.1305 (below floor)
Volatility0.01970.0883 (0.0891)0.1480 (below floor)
Spread0.01780.07620.2517 (0.2519)
Multihead0.0311 (0.0317)0.0941 (0.0952)0.2455 (0.2441)
LeJEPA
Same Stock, Diff. View0.01820.07150.1672 (below floor)
Time Warping0.02040.07210.1691 (below floor)
Gaussian Noising0.01750.0688 (below floor)0.1625 (below floor)
Cross Stock0.0169 (below floor)0.0693 (below floor)0.1667 (below floor)
C-S, Same Industry0.01730.0692 (below floor)0.1660 (below floor)
SSL
DINO0.0097 (below floor)0.0310 (below floor)0.0626 (below floor)
BYOL0.0104 (below floor)0.0314 (below floor)0.0686 (below floor)
CPC0.01840.06970.1519 (below floor)
I-JEPA0.0103 (below floor)0.0256 (below floor)0.0646 (below floor)
MAE0.01810.0684 (below floor)0.1435 (below floor)
TS2Vec0.02050.07190.1500 (below floor)
CoST0.0100 (below floor)0.0339 (below floor)0.0302 (below floor)
TF-C0.0095 (below floor)0.0570 (below floor)0.1171 (below floor)
TimeMAE0.01990.07060.1408 (below floor)
Untrained floor
Random ViT0.01700.06940.1692

Organization

We embed views of six firms from three industries within one evaluation month and ask whether nearby embeddings share a partner, a day, a firm or an industry.

Four retrieval tasks: partner matched view, own-day centroid, own-firm centroid, partner firm centroidFour retrieval tasks: partner matched view, own-day centroid, own-firm centroid, partner firm centroid
Top-1 rate as a multiple of chance / mean percentile rank of the target (lower is better). Green marks the top three trained methods in a column; red is no better than the untrained Random ViT.
EncoderPartner Matched ViewOwn-Day CentroidOwn-Firm CentroidPartner Firm Centroid
Supervised
Return2.26× / 44.3%1.76× / 42.8%1.25× (below floor) / 45.6% (below floor)1.17× / 47.3% (below floor)
Vol2.13× / 44.8%1.70× / 43.2%1.91× / 33.9%1.27× / 45.6%
Spread1.65× (below floor) / 46.3% (below floor)1.45× (below floor) / 45.5% (below floor)2.41× / 26.1%1.15× (below floor) / 47.0% (below floor)
Multihead2.25× / 44.4%1.70× / 43.6%2.35× / 28.6%1.38× / 43.8%
LeJEPA
Same Stock, Diff. View1.65× (below floor) / 45.0%1.31× (below floor) / 46.0% (below floor)5.36× / 2.5%1.23× / 45.5%
Time Warping1.97× / 39.7%2.01× / 37.7%1.54× (below floor) / 40.9% (below floor)1.07× (below floor) / 47.5% (below floor)
Gaussian Noising2.33× / 41.1%2.03× / 39.2%1.59× (below floor) / 38.6% (below floor)1.12× (below floor) / 47.3% (below floor)
Cross Stock2.46× / 40.5%2.09× / 38.4%1.55× (below floor) / 39.3% (below floor)1.13× (below floor) / 47.6% (below floor)
C-S, Same Industry2.58× / 39.7%2.22× / 37.6%1.52× (below floor) / 39.9% (below floor)1.12× (below floor) / 47.3% (below floor)
SSL
DINO2.28× / 39.6%2.04× / 37.9%1.49× (below floor) / 40.5% (below floor)1.23× / 46.1%
BYOL2.23× / 41.0%1.94× / 39.1%1.72× (below floor) / 36.4% (below floor)1.19× / 46.4%
CPC2.07× / 46.2% (below floor)1.57× / 45.1% (below floor)1.55× (below floor) / 40.3% (below floor)1.15× (below floor) / 48.1% (below floor)
I-JEPA2.24× / 40.5%1.89× / 39.5%1.06× (below floor) / 49.0% (below floor)1.24× / 45.1%
MAE2.42× / 40.2%2.10× / 38.5%1.23× (below floor) / 45.4% (below floor)1.22× / 46.7% (below floor)
TS2Vec1.86× / 43.5%1.47× (below floor) / 44.4%3.99× / 11.0%1.37× / 43.2%
CoST2.13× / 40.6%2.00× / 38.1%1.03× (below floor) / 49.6% (below floor)1.17× / 47.4% (below floor)
TF-C2.15× / 45.6% (below floor)1.55× / 45.0% (below floor)4.14× / 10.6%1.39× / 44.7%
TimeMAE3.35× / 38.6%2.66× / 36.0%1.32× (below floor) / 43.9% (below floor)1.33× / 44.6%
Frozen time-series foundation models
Chronos-22.24× / 42.5%1.88× / 39.9%2.56× / 24.9%1.07× (below floor) / 48.4% (below floor)
Kronos2.90× / 41.1%2.27× / 38.6%1.43× (below floor) / 42.2% (below floor)1.40× / 44.2%
TimesFM 3.01.95× / 44.2%1.55× / 43.4%2.43× / 28.2%1.22× / 45.4%
Untrained floor
Random ViT1.82× / 45.6%1.48× / 44.4%1.81× / 36.2%1.16× / 46.6%

Factor Structure

We estimate statistical factors from the month's realized returns[5] and ask whether each firm's loadings can be decoded from its embedding, and whether the embedding's leading directions span the factor subspace.

Factor loadings estimated from realized correlations, then decoded from embeddingsFactor loadings estimated from realized correlations, then decoded from embeddings
Decoded loadings: mean out-of-sample correlation. Subspace alignment: share of the factor subspace captured. Higher is better. Green marks the top three trained methods in a column; red is no better than the untrained Random ViT.
EncoderDecoded LoadingsSubspace Alignment
Supervised
Return0.6330.202
Vol0.5810.207
Spread0.459 (below floor)0.161
Multihead0.5820.243
LeJEPA
Same Stock, Diff. View0.391 (below floor)0.199
Time Warping0.5180.218
Gaussian Noising0.381 (below floor)0.162
Cross Stock0.415 (below floor)0.167
C-S, Same Industry0.433 (below floor)0.173
SSL
DINO0.5440.190
BYOL0.444 (below floor)0.199
CPC0.475 (below floor)0.166
I-JEPA0.6080.203
MAE0.5470.178
TS2Vec0.464 (below floor)0.245
CoST0.7110.181
TF-C0.5340.221
TimeMAE0.6460.207
Frozen time-series foundation models
Chronos-20.6740.221
Kronos0.6200.184
TimesFM 3.00.7160.194
Untrained floor
Random ViT0.4950.151

Summary

How well an encoder forecasts says little about how it organizes the market. Methods that cluster by firm do not cluster by day, and the forecasting and organization rankings are unrelated.

Rank correlation between every pair of evaluation tasks across 19 encodersRank correlation between every pair of evaluation tasks across 19 encoders

Spearman correlation between the 19 encoders' rankings on every pair of tasks; bold is p < 0.05. Own-day and own-firm clustering pull in opposite directions.

Organization rank versus forecasting rank for 19 encodersOrganization rank versus forecasting rank for 19 encoders

Each encoder's mean rank on forecasting against its mean rank on organization. TimeMAE and the supervised multihead trace the frontier.

Citation

Correspondence: bll@uchicago.edu

\cite{merchant2026towards}
@inproceedings{merchant2026towards,
  title={Towards Financial World Modeling},
  author={Merchant, Humzah and Guthrie, Alec and Mahns, Simon and Balestriero, Randall and Levy, Bradford},
  booktitle={The Fortieth Annual Conference on Neural Information Processing Systems},
  year={2026},
  url={https://arxiv.org/abs/2610.09048}
}

References

  1. Gaoyue Zhou and Hengkai Pan and Yann LeCun and Lerrel Pinto. DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning. Forty-second International Conference on Machine Learning, 2025.
  2. Mido Assran and Adrien Bardes and David Fan and Quentin Garrido and others. V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning. arXiv:2506.09985, 2025.
  3. Lucas Maes and Quentin Le Lidec and Damien Scieur and Yann LeCun and Randall Balestriero. LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels. 2026.
  4. Shengding Hu and Yuge Tu and Xu Han and others. MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies. arXiv:2404.06395, 2024.
  5. Markus Pelger. Large-Dimensional Factor Modeling Based on High-Frequency Observations. SSRN, 2018.