Track Record separates cautious forecasts from higher-variance research. Accuracy appears after recorded markets resolve.
Balanced Mode
A cautious view that favors stronger evidence and smaller, more defensible probability gaps.
◎73.5%Forecast Accuracy
▥34Total Forecasts
↗12.0%Avg Edge %
♨10WBest Streak
◎ Calibration Curve
Compare predicted probabilities with observed outcomes.
Expected 50% · Observed —
Ideal calibrationObserved accuracy: not enough data
Accuracy by Category — Balanced Mode
Category
Resolved
Accuracy
Avg Edge vs Market
No resolved forecasts yet.
Balanced Mode Resolved Forecasts
Date
Market
Category
AI Pred
Mkt Price
Edge %
Side
Outcome
No resolved forecasts yet.
▤ How We Measure Accuracy
◎
CRPS: Beyond Win/Loss
Unlike simple win/loss tracking, Continuous Ranked Probability Score (CRPS) measures how close our probability estimates are to the actual outcome. A forecast of 90% on an event that happens scores better than a forecast of 55% on the same event — even though both are technically "correct." This incentivizes precision, not just being on the right side.
◷
Nightly Recalibration
Every night, our system reviews all resolved forecasts and computes per-category calibration multipliers. If we're systematically overconfident in crypto forecasts (e.g., predicting 80% when the true rate is 75%), the multiplier corrects for this bias going forward. This process runs automatically and the results are reflected in the calibration chart above.
⚯
Model Disagreement Scoring
When all 6 models agree (6/6), our confidence is highest. When models disagree significantly (3/6 or 4/6), we flag lower confidence and adjust our calibrated output downward. Disagreement is itself a valuable signal — it often indicates genuinely uncertain markets where the outcome is harder to predict.