# The Future is the Label.

Peer-reviewed research. Live benchmark wins. Government-cleared.

## [Forecasting supply chain disruptions with foresight learning](https://arxiv.org/abs/2604.01298)

Foresight learning trains LLMs to generate calibrated probability forecasts of rare supply chain disruptions, outperforming GPT-5 in accuracy, calibration, and precision — with structured probabilistic reasoning emerging from training alone.

## [Foresight-v3 becomes the \#1 AI forecaster](https://www.prophetarena.co/)

Foresight-v3 ranks first overall on ProphetArena — an independent AI forecasting benchmark from UChicago — by Brier score, outperforming GPT-5, Gemini 3 Pro, and every frontier model. Also #1 in Sports.

## [\#1 on ProphetArena Sports](https://www.prophetarena.co/leaderboard)

Foresight-32B beats every other model at predicting sports outcomes on ProphetArena, a live prediction market leaderboard — with 105.9% Market Return, ahead of GPT-5.2, Minimax M2, Gemini 3 Pro, and Qwen3-235B.

## [Foresight-32B outperforms frontier models on ForecastBench](https://forecastingresearch.substack.com/p/llms-are-closing-the-gap-on-human)

Top 5 on the ForecastBench tournament, outperforming Gemini 3 Pro, Claude Sonnet 4.5, and o3.

## [Foresight-tuned 32B model outperforms GPT-5 at predicting public company risks](https://arxiv.org/abs/2601.19189)

Foresight learning on raw SEC filings trains a 32B parameter model to beat GPT-5 in accuracy & calibration at predicting public company risks. Deployable on a single GPU for maximum data privacy.

## [Future-as-Label enables scalable RL](https://arxiv.org/abs/2601.06336)

We show that AI can learn directly from real-world outcomes at unlimited scale, no human annotation required. The future itself becomes the training signal. Improved Brier scores 27% and halved calibration error, outperforming Qwen3-235B with a 32B model.

## [Foresight-32B beats frontier LLMs on live Polymarket predictions](https://blog.lightningrod.ai/p/foresight-32b-beats-frontier-llms-on-live-polymarket-predictions)

On live Polymarket data, Foresight-32B defeated models 100x larger across every key metric — Brier score, calibration error, and simulated trading profit.

## [Defense & DARPA awardable](https://www.einpresswire.com/article/830532019/lightning-rod-labs-assessed-awardable-for-darpa-eris-marketplace)

Vetted and approved for immediate defense procurement. Government agencies can access our technology directly via the ERIS and CDAO Tradewinds federal innovation marketplaces.

## [Published in TMLR: Outcome-based RL achieves frontier accuracy with a 14B model](https://arxiv.org/abs/2505.17989)

Our 14B model matches OpenAI o1 in predictive accuracy and generates >10% profit in live trading simulations — published in Transactions on Machine Learning Research.

## [LLMs can teach themselves to predict the future](https://arxiv.org/abs/2502.05253)

Self-play and DPO yield 7–10% accuracy improvements on Phi-4 14B and DeepSeek-R1 14B — bringing smaller models to frontier-level forecasting performance without any human-annotated training data.
