# Economic Forecasting

A 32B model that beats GPT-5 on Fed Beige Book macro questions

22% lower Brier score than GPT-5 on Fed Beige Book forecasting questions

6× better calibration than GPT-5 (ECE 0.029 vs. 0.188)

6× fewer output tokens than GPT-5 — dramatically cheaper inference at higher accuracy

## What we did

- Built a time-indexed corpus of Beige Book narratives, CPI/payrolls/JOLTS releases, and survey data.
- Generated macro forecasting questions at each publication date using later prints as labels.
- Trained with RL against a calibration-aware reward.
- Benchmarked against GPT-5 on held-out Beige Book windows.

* * *

## Example datapoint

### DATASET **Macro prints**

- **RL**
- **Question**: Will next month’s CPI print come in above the consensus forecast?
  - **Question source**: ReutersApr 2, 2025
  - Economists pencil in +0.3% core CPI month-over-month
  - **Label**: Yes.
  - **Type**: binary
  - **Confidence**: 0.87
  - **Label source**: BLSMay 13, 2025
  - Consumer Price Index summary: core rose 0.4% in April

### DATASET **Macro prints**

- **RL**
- **Question**: Will payroll growth surprise to the downside in the next employment report?
  - **Question source**: Bloomberg surveyJun 4, 2025
  - Median forecast +185k nonfarm payrolls
  - **Label**: No.
  - **Type**: binary
  - **Confidence**: 0.82
  - **Label source**: BLS Employment SituationJul 3, 2025
  - Payrolls +272k; prior two months revised up

### DATASET **Macro prints**

- **RL**
- **Question**: What is the probability that the FOMC reduces the federal funds target range by 25 bps at the Jan 28–29, 2026 meeting, conditional on the Dec 2025 Summary of Economic Projections median dot at 3.6% end-2026 and fed funds futures pricing 1.2 cuts by midyear?
  - **Question source**: CME FedWatchJan 10, 2026
  - Implied probability of Jan cut rises to 38% after cooler payrolls revision
  - **Label**: 0.38
  - **Type**: continuous
  - **Confidence**: 0.86
  - **Label source**: FOMC statementJan 29, 2026
  - Committee holds target range unchanged; forward guidance softened

### DATASET **Macro prints**

- **RL**
- **Question**: Will the BLS JOLTS job openings level print below 7.5 million for Dec 2025 (release Feb 3, 2026) after three consecutive sub-8m reads and ISM Services Employment at 48.9 in the January survey?
  - **Question source**: Wall Street JournalJan 22, 2026
  - Economists split on whether JOLTS downtrend resumes after November bounce
  - **Label**: Yes.
  - **Type**: binary
  - **Confidence**: 0.80
  - **Label source**: BLS JOLTSFeb 3, 2026
  - December job openings 7.2 million; hires little changed

* * *

## Results

### Better Accuracy, Skill, and Calibration vs. GPT-5

Trained on Fed Beige Book narratives, Foresight posts a Brier score of 0.155 vs. 0.199 for GPT-5 and 0.211 for the base model — a 22% reduction in error. It is the only model to beat the base rate (Brier Skill Score +6.2% vs. −20.7% for GPT-5 and −27.7% for the base model), and cuts calibration error (ECE) by ~6× vs. GPT-5.

### Probabilities That Match Reality — Unlike GPT-5

Foresight (yellow) hugs the perfect-calibration diagonal — when it says 30%, roughly 30% of events materialize. GPT-5 and the base model are systematically overconfident, drifting well below the line at higher probabilities.
