We use privacy-friendly analytics to understand traffic. With your consent we also load Google Analytics. No analytics cookies are set until you accept. See our privacy policy.
What the rolling-strength model forecast before each kickoff — win/draw/loss probability and expected score — graded against what actually happened. Out-of-sample (the model never saw the result). What do these mean?
Every call the model made this season, graded against the result — and, more importantly, whether its probabilities held up.
Accuracy is the wrong lens for a three-way outcome: draws are rarely any side's single most-likely result, so ~25% of games are "wrong" by design. What matters is whether a stated 60% really happens ~60% of the time — that's calibration, and it's the number that actually matters.