# Supply Chain Disruptions

A forecaster that predicts supplier and logistics disruptions before they happen

34%

more accurate than GPT-5 on supply chain predictions

59%

better calibrated than GPT-5

4×

precision@10% vs. baseline on high-risk disruptions

## What we did

- Ingested supplier performance metrics, ERP shipment logs, and industry news with timestamps.
- Generated forward-looking disruption questions anchored at each point in history.
- Labeled each question using the actual downstream outcome (ASN slip, dwell spike, price move).
- Trained a 32B model with RL to output calibrated probabilities, then benchmarked against GPT-5.

---

## Example datapoint

A sample training example — question, source, and outcome-derived label.

DATASET **Disruption outlook**  
RL  
**Question**  
Will Acme Motors Tier-1 supplier FlexAlloy (plant Shenzhen) log a >3-day ASN slip on part HV-220 busbars in the next 30 days?  
**Question source**  
Supplier portal  
Jan 8, 2026  
On-time performance slipped to 94% vs. 98% prior quarter  
**Label**  
Yes.  
**Type**  
binary  
**Confidence**  
0.91  
**Label source**  
ERP shipment log  
Feb 4, 2026  
ASN delayed 11 days — weather and carrier capacity cited

DATASET **Disruption outlook**  
RL  
**Question**  
Will TSMC advanced-node (N3E) wafer starts allocated to our GPU line card fall short of the 12k wafers/quarter run rate implied by the Mar 14, 2025 foundry letter?  
**Question source**  
Financial Times  
Mar 14, 2025  
Foundries warn of tight advanced-node allocation through Q3  
**Label**  
No.  
**Type**  
binary  
**Confidence**  
0.86  
**Label source**  
Company ops review  
Jun 30, 2025  
Dual-sourcing plan held; no line-down events recorded

DATASET **Disruption outlook**  
RL  
**Question**  
Will median dwell time for our 40-foot import boxes at Port of Long Beach terminals exceed 4.5 days during the week of Jan 13–19, 2026 (per PIERS / terminal API)?  
**Question source**  
Journal of Commerce  
Jan 6, 2026  
ILWU local meetings raise risk of slower gate turns at San Pedro Bay  
**Label**  
No.  
**Type**  
binary  
**Confidence**  
0.85  
**Label source**  
TMS analytics  
Jan 20, 2026  
LB dwell 3.1 days avg; no demurrage events on monitored POs

DATASET **Disruption outlook**  
RL  
**Question**  
What is the probability that LME three-month nickel settles above $19,000/ton at least once in the 60 trading days after the Mar 1, 2026 Indonesian ore export quota announcement?  
**Question source**  
Fastmarkets  
Feb 18, 2026  
Nickel rally stalls as stainless restocking fades  
**Label**  
0.22  
**Type**  
continuous  
**Confidence**  
0.80  
**Label source**  
LME  
May 20, 2026  
Cash-settled month high $18,420/ton; no session above $19k in window

---

## Results

Benchmark comparisons against frontier models.

### Aggregate Performance vs. GPT-5 and gpt-oss-120b

The trained model beats GPT-5 and gpt-oss-120b on every headline metric: Brier 0.0791 vs. 0.1203 (GPT-5), Brier Skill Score +16.9% vs. −26.4%, and ECE 0.0525 vs. 0.1740 for the pretrained base — a ~70% reduction in calibration error.

### Calibration vs. GPT-5 on Supply Chain Predictions

The trained model closely tracks the perfect-calibration diagonal — when it says 30% probability, roughly 30% of disruptions materialize. GPT-5 and the base model severely under-predict high-risk events. Precision@10% improves 4× (35% vs. 9%).
