Prediction Market Betting
A 32B model that beats frontier LLMs 100× its size on live markets
#1
ranked forecaster on ProphetArena, beating GPT-5, Gemini & Claude
69%
reduction in calibration error vs. base model on live Polymarket questions
10–100×
smaller than the frontier models it beats
What we did
- Collected timestamped news and market-style questions with later-resolved outcomes.
- Used the known future outcome at each question's resolution date as the training label — no annotators.
- Trained Foresight-32B with RL (GRPO-style) against a Brier-derived reward.
- Evaluated live on ProphetArena and Polymarket against frontier models.
Example datapoint
A sample training example — question, source, and outcome-derived label.
DATASET Market outcomes
RL
Question
Will the US Fed cut rates by more than 25bps before end of Q3?
Question source
Wall Street Journal Jun 12, 2025
Fed officials signal openness to larger cuts if labor market softens
Label
No.
Type
binary
Confidence
0.88
Label source
Federal Reserve Sep 30, 2025
FOMC statement: 25 basis point reduction in target range
DATASET Market outcomes
RL
Question
Will Bitcoin reach $120k before June 2026?
Question source
Bloomberg Aug 3, 2025
Options markets price in volatile path for crypto into next halving cycle
Label
Yes.
Type
binary
Confidence
0.79
Label source
CoinDesk May 28, 2026
Bitcoin trades above $120,000 for first time on spot volume surge
DATASET Market outcomes
RL
Question
Will OpenAI ship a new flagship reasoning model to the ChatGPT API before Anthropic ships a Claude 4.5 successor, as of Polymarket close on Mar 31, 2026?
Question source
The Information Jan 9, 2026
Labs race to refresh flagship APIs after holiday traffic spike
Label
Yes.
Type
binary
Confidence
0.76
Label source
OpenAI Mar 18, 2026
API changelog: new flagship reasoning tier in gradual rollout
DATASET Market outcomes
RL
Question
Will OPEC+ announce a combined voluntary cut of at least 1 million barrels per day at the Jan 5, 2026 ministerial videoconference?
Question source
Reuters Dec 12, 2025
Delegates expect voluntary cuts extension; deeper reduction on agenda
Label
No.
Type
binary
Confidence
0.83
Label source
OPEC Jan 5, 2026
Joint Ministerial Monitoring Committee rolls over existing quotas
Results
Benchmark comparisons against frontier models.
ProphetArena Overall Leaderboard
Foresight V3 holds the #1 spot on ProphetArena's live benchmark, ahead of Gemini 3 Pro and GPT-5.2 — while being 10–100× smaller than the frontier models it beats.
Live Polymarket Benchmark
On 251 live Polymarket questions, Foresight-32B achieved Brier score 0.199 vs. GPT-5's 0.207 — with 69% lower calibration error (ECE 6.0% vs. 16.1%) and positive simulated trading profit while frontier models lost money.
Ready to build your own expert?
Leverage your own raw data or use public sources. No labeling required.