
Benchmarks
How good is Cassi?
ForecastBench and Metaculus are the two leading tournaments for testing AI forecasters against elite human forecasters and each other. Cassi is scored the same way everyone else is. No home advantage.
#1 AI
on the ForecastBench leaderboard
19.2%
better than the average tournament forecaster
7%
better than the wisdom of crowds
0.4%
better than the wisdom of the AI crowd
ForecastBench leaderboard
| Rank | Organization | Model | Overall | Brier |
|---|---|---|---|---|
| 1 | ![]() | Superforecaster median forecast | 69.2 | 0.095 |
| 2 | ![]() | Cassi-2026-05-10 | 68.8 | 0.097 |
| 3 | ![]() | Grok 4.20 (Beta, C) | 68.1 | 0.102 |
| 3 | ![]() | Grok 4.20 (Beta, D) | 68.1 | 0.102 |
| 5 | ![]() | green tree | 67.5 | 0.106 |
| 5 | ![]() | green-plant | 67.5 | 0.105 |
| 7 | ![]() | plastic-cactus | 67.2 | 0.107 |
| 8 | ![]() | yellow mouse | 67.1 | 0.108 |
| 8 | ![]() | blue-turtle | 67.1 | 0.108 |
| 10 | Artificial Judgement | aj-v1 | 67 | 0.109 |
| 11 | ![]() | Gemini | 66.7 | 0.111 |
| 12 | ![]() | Grok 4.20 (Beta, B) | 66.6 | 0.112 |
| 13 | ![]() | Cassi ensemble_2_crowdadj | 66.5 | 0.112 |
| 13 | ![]() | Grok 4.20 (Beta, A) | 66.5 | 0.112 |
| 13 | ![]() | red-lizard | 66.5 | 0.112 |
| 16 | Anonymous 8 | ensemble | 66.4 | 0.113 |
| 17 | ![]() | Grok 4.20 (Preview) | 66.2 | 0.114 |
| 18 | limeforecast | megaflash | 66.1 | 0.115 |
| 19 | OpenAI | gpt-5-2025-08-07† | 66 | 0.115 |
| 19 | Mantic | mantic-2026-01-16 | 66 | 0.115 |
| 19 | Lightning Rod Labs | Foresight-32B | 66 | 0.116 |
| 19 | ![]() | bronze-hat | 66 | 0.115 |
ForecastBench's elite-human-forecaster median is the combined score of top human forecasters, not any single person's best result. It sits beyond what any individual forecaster has achieved alone.
ForecastBench leaderboard last updated: 17/07/2026 at 07:31 am
Understanding the scoring
The percentage here comes from Metaculus's Peer Score (a log-loss scoring rule) and ForecastBench's Difficulty-Adjusted Brier Score. Both are technical, but the simple version is this: they show how much better or worse a forecast performed than the median forecaster on that platform.
These adjustments exist because some questions are much easier than others. Predicting whether the Earth still exists tomorrow is easier than predicting the FTSE 100 in ten years' time. That makes raw scoring comparisons misleading: what matters is how one forecaster performs against others answering the same set of questions, not comparing across entirely different tournaments. It's the difference between judging like for like and comparing apples to orangutans. So we calculate Cassi's score across the questions and tournaments it forecasts on, then compare that to the reference groups above. The percentage shown is how much better or worse that score is, normalised so every question across every platform and tournament carries equal weight.
The Science of Elite Forecasting · Overview last updated: 17/07/2026 at 10:55 am





