# Prediction Market Betting

A 32B model that beats frontier LLMs 100× its size on live markets

#1

ranked forecaster on ProphetArena, beating GPT-5, Gemini & Claude

69%

reduction in calibration error vs. base model on live Polymarket questions

10–100×

smaller than the frontier models it beats

## What we did

- Collected timestamped news and market-style questions with later-resolved outcomes.
- Used the known future outcome at each question's resolution date as the training label — no annotators.
- Trained Foresight-32B with RL (GRPO-style) against a Brier-derived reward.
- Evaluated live on ProphetArena and Polymarket against frontier models.

---

## Example datapoint

A sample training example — question, source, and outcome-derived label.

DATASET **Market outcomes**

RL

Question

Will the US Fed cut rates by more than 25bps before end of Q3?

Question source

Wall Street Journal Jun 12, 2025

Fed officials signal openness to larger cuts if labor market softens

Label

No.

Type

binary

Confidence

0.88

Label source

Federal Reserve Sep 30, 2025

FOMC statement: 25 basis point reduction in target range

DATASET **Market outcomes**

RL

Question

Will Bitcoin reach $120k before June 2026?

Question source

Bloomberg Aug 3, 2025

Options markets price in volatile path for crypto into next halving cycle

Label

Yes.

Type

binary

Confidence

0.79

Label source

CoinDesk May 28, 2026

Bitcoin trades above $120,000 for first time on spot volume surge

DATASET **Market outcomes**

RL

Question

Will OpenAI ship a new flagship reasoning model to the ChatGPT API before Anthropic ships a Claude 4.5 successor, as of Polymarket close on Mar 31, 2026?

Question source

The Information Jan 9, 2026

Labs race to refresh flagship APIs after holiday traffic spike

Label

Yes.

Type

binary

Confidence

0.76

Label source

OpenAI Mar 18, 2026

API changelog: new flagship reasoning tier in gradual rollout

DATASET **Market outcomes**

RL

Question

Will OPEC+ announce a combined voluntary cut of at least 1 million barrels per day at the Jan 5, 2026 ministerial videoconference?

Question source

Reuters Dec 12, 2025

Delegates expect voluntary cuts extension; deeper reduction on agenda

Label

No.

Type

binary

Confidence

0.83

Label source

OPEC Jan 5, 2026

Joint Ministerial Monitoring Committee rolls over existing quotas

---

## Results

Benchmark comparisons against frontier models.

### ProphetArena Overall Leaderboard

Foresight V3 holds the #1 spot on ProphetArena's live benchmark, ahead of Gemini 3 Pro and GPT-5.2 — while being 10–100× smaller than the frontier models it beats.

### Live Polymarket Benchmark

On 251 live Polymarket questions, Foresight-32B achieved Brier score 0.199 vs. GPT-5's 0.207 — with 69% lower calibration error (ECE 6.0% vs. 16.1%) and positive simulated trading profit while frontier models lost money.

## Ready to build your own expert?

Leverage your own raw data or use public sources. No labeling required.
