Benchmarks

How good is Cassi?

ForecastBench and Metaculus are the two leading tournaments for testing AI forecasters against elite human forecasters and each other. Cassi is scored the same way everyone else is. No home advantage.

#1 AI

on the ForecastBench leaderboard

19.2%

better than the average tournament forecaster

7%

better than the wisdom of crowds

0.4%

better than the wisdom of the AI crowd

ForecastBench leaderboard

RankOrganizationModelOverallBrier
1ForecastBenchSuperforecaster median forecast69.20.095
2CassiCassi-2026-05-1068.80.097
3xAIGrok 4.20 (Beta, C)68.10.102
3xAIGrok 4.20 (Beta, D)68.10.102
5Google DeepMindgreen tree67.50.106
5Google DeepMindgreen-plant67.50.105
7Google DeepMindplastic-cactus67.20.107
8Google DeepMindyellow mouse67.10.108
8Google DeepMindblue-turtle67.10.108
10Artificial Judgementaj-v1670.109
11Google DeepMindGemini66.70.111
12xAIGrok 4.20 (Beta, B)66.60.112
13CassiCassi ensemble_2_crowdadj66.50.112
13xAIGrok 4.20 (Beta, A)66.50.112
13Google DeepMindred-lizard66.50.112
16Anonymous 8ensemble66.40.113
17xAIGrok 4.20 (Preview)66.20.114
18limeforecastmegaflash66.10.115
19OpenAIgpt-5-2025-08-07†660.115
19Manticmantic-2026-01-16660.115
19Lightning Rod LabsForesight-32B660.116
19Google DeepMindbronze-hat660.115

ForecastBench's elite-human-forecaster median is the combined score of top human forecasters, not any single person's best result. It sits beyond what any individual forecaster has achieved alone.

ForecastBench leaderboard last updated: 17/07/2026 at 07:31 am

Understanding the scoring

The percentage here comes from Metaculus's Peer Score (a log-loss scoring rule) and ForecastBench's Difficulty-Adjusted Brier Score. Both are technical, but the simple version is this: they show how much better or worse a forecast performed than the median forecaster on that platform.

These adjustments exist because some questions are much easier than others. Predicting whether the Earth still exists tomorrow is easier than predicting the FTSE 100 in ten years' time. That makes raw scoring comparisons misleading: what matters is how one forecaster performs against others answering the same set of questions, not comparing across entirely different tournaments. It's the difference between judging like for like and comparing apples to orangutans. So we calculate Cassi's score across the questions and tournaments it forecasts on, then compare that to the reference groups above. The percentage shown is how much better or worse that score is, normalised so every question across every platform and tournament carries equal weight.

The Science of Elite Forecasting · Overview last updated: 17/07/2026 at 10:55 am

Tournament record

ForecastBenchLive1st of 269top 1% · among AI forecastersForecast Bench1421 questions · as of 17th Jul, 2026
Metaculus46th of 2867top 2%Bridgewater x Metaculus 202651 questions · ended 16th Mar, 2026
Metaculus18th of 539top 4%Metaculus Cup - Fall 202558 questions · ended 31st Dec, 2025
Metaculus2nd of 50top 4%MiniBench - 2026-03-3060 questions · ended 26th Apr, 2026
Metaculus50th of 1130top 5%Metaculus Cup - Spring 202647 questions · ended 5th May, 2026
Metaculus5th of 97top 6%Spring 2026 AI Forecasting Benchmark Tournament297 questions · ended 6th May, 2026
Metaculus6th of 54top 12%MiniBench - 2026-03-1659 questions · ended 11th Apr, 2026
Metaculus4th of 32top 13%MiniBench - 2025-11-2452 questions · ended 20th Dec, 2025
Metaculus6th of 43top 14%MiniBench - 2025-10-1359 questions · ended 8th Nov, 2025
Metaculus9th of 54top 17%MiniBench - 2026-03-0260 questions · ended 29th Mar, 2026
Metaculus10th of 53top 19%MiniBench - 2026-04-2057 questions · ended 10th May, 2026
Metaculus9th of 45top 20%MiniBench - 2025-09-2955 questions · ended 25th Oct, 2025
Metaculus10th of 48top 21%MiniBench - 2026-02-0249 questions · ended 28th Feb, 2026
Metaculus15th of 71top 22%Fall AIB - 2025377 questions · ended 31st Dec, 2025
Metaculus13th of 58top 23%MiniBench - 2026-05-1852 questions · ended 24th May, 2026
Metaculus15th of 57top 27%MiniBench - 2026-02-1658 questions · ended 15th Mar, 2026
Metaculus14th of 51top 28%MiniBench - 2026-05-0460 questions · ended 10th May, 2026
Metaculus16th of 43top 38%MiniBench - 2025-10-2760 questions · ended 22nd Nov, 2025
Metaculus17th of 37top 46%MiniBench - 2026-01-1959 questions · ended 4th Feb, 2026
edrith.co.ukLiveIn progressEdrith 2026 Forecasting Contest34 questions · ends 1st Jan, 2027
MetaculusLiveIn progressACX 2026 Prediction Contest34 questions · ends 1st Jan, 2027
MetaculusLiveIn progressSummer 2026 FutureEval Bot Tournamentends 6th Sep, 2026
Good Judgment OpenLiveIn progressThe Economist 2025-26 Supreme Court (SCOTUSBot) Forecasting Challenge16 questions · ends 1st Jul, 2026