Methodology

Model the goals,
derive the rest.

Everything starts from one thing: how many goals each team scores and concedes. From that goal intensity follow the result, the totals and the exact score. We open up this core — the method, the matrix, the calibration — so that you can judge us before you trust us.

Attacking strength · Serie A
Inter1.51
Atalanta1.42
Napoli1.27
Juventus1.18
Roma1.09
Milan1.02

Illustration — to be connected to the model.

Three steps

From strengths to proof.

01

Strengths

For each team, an attacking intensity and a defensive solidity, estimated over the season and adjusted for the quality of the opposition faced.

02

The scoreline matrix

From those strengths, the probability of every exact score. Every outcome follows from that grid: result, totals, margin.

03

Verification

Each probability band is checked against what actually happened. If we say 30%, it should happen three times in ten — and we show it.

Our models

One core,
and what we build on it.

The goals model is only the starting point. Around it, we replay entire seasons, simulate the European cups and measure what a player is worth — each building block with its own method, limits and evidence.

01

The team model

Available

An attacking strength and a defensive solidity for each team, estimated with a time-weighted Dixon-Coles model and recomputed after every completed matchday. From those strengths comes the full scoreline distribution for every match.

Model outputs →
02

Season projections

Available

The rest of the season replayed 50,000 times by Monte Carlo simulation, from the estimated strengths. For each club, the full distribution of its final position — title, Europe, survival, relegation — rather than a single “predicted” table.

Projected tables →
03

The European cups

Available

Probabilities for every upcoming fixture and, for the Champions League, a simulation of each club’s path all the way to the final.

European cups →
04

Player models

Advanced research area
In development

Dixon-Coles does not know who is playing. Our second model estimates what each player is worth at a given date — using only what came before that date, never what followed — and what his absence costs his team: the starting XI rebuilt with the actual replacement, the scoreline distribution recomputed, the difference quantified in percentage points of probability.

First result, on matches held out of training: adding what each player produces triples the predictive gain obtained by describing the team’s style alone. A solid signal, not yet a finished product: it is open as a pilot.

Predictive gain · on matches the model has never seen
Knowing the team’s style+1.18%
+ what the player actually produces+3.86%

Both models already know which teams are playing. Describing the team’s style improves the prediction slightly. Adding what each player produces triples the gain — and it holds out of sample, on matches held back during training. That is the evidence that the player signal is real, not a redescription of the team.

05

Public calibration

The safeguard for the whole family. What we forecast is checked, band by band, against what actually happened — and published openly, including where the model performs worst.

See the evidence →
Evidence, not claims

Forecast versus observed.

This is the only chart that proves a model tells the truth: for each probability band, what we forecast set against what actually happened. The closer the points sit to the diagonal, the more the probabilities can be trusted. We publish it; almost nobody else does.

Reliability · what the model forecasts against what happens
01000100%Forecast probabilityForecast 8.3% → observed 0.0% · 7 casesForecast 15.5% → observed 9.9% · 71 casesForecast 25.8% → observed 25.0% · 156 casesForecast 35.3% → observed 31.9% · 360 casesForecast 45.2% → observed 43.5% · 855 casesForecast 53.7% → observed 53.3% · 396 casesForecast 64.2% → observed 67.7% · 164 casesForecast 74.5% → observed 87.7% · 57 casesForecast 83.8% → observed 88.5% · 26 casesForecast 90.6% → observed 100.0% · 2 cases

Each point is a probability band; its size is the number of cases. On the diagonal, the model keeps its word. Here, across 2,094 matches (1X2 result), the mean gap between forecast and observed is 2.3%.

0Competitions modelled
0Seasons simulated per league
0Calibration matches
0Leagues projected to the end of the season
Open, and ours

What we show, and what we keep.

Open

Enough to judge us

The method, the scoreline matrix, the calibration curve. Enough to check that we tell the truth before you place your trust in us.

Ours

Our edge

The player effect, in-play modelling, conditional projections. The edge we are building on top of the core — what you deliver to your teams.

Want to go into the detail?

We are happy to walk you through the core, its limits and what it can do for your use case. Get in touch.