Predicted Edge

Methodology

The Dixon–Coles model explained

How two numbers per team become a probability for every scoreline, and why the correction for low-scoring games is the part that matters.

Goals as a Poisson process

Goals are rare, roughly independent events spread across ninety minutes, which is the situation the Poisson distribution describes. If a side is expected to score 1.45 goals, Poisson gives the chance of each exact total: 23.5% for none, 34.0% for one, 24.7% for two, and so on. The expected figure is the only input; the whole distribution follows from it.

A match needs two of these — one per side — and the expected goals come from the ratings. Each team has an attack strength and a defence strength, estimated from its results against the sides it actually played, and the home side carries a home-advantage term. Home expected goals are the home attack times the away defence times the home advantage; away expected goals are the away attack times the home defence. Multiply the two distributions and every scoreline has a probability.

What Dixon and Coles fixed

Plain Poisson treats the two sides’ goals as independent, and in the lowest-scoring games they are not quite. Tested against decades of results, the plain model gives too little weight to 0–0 and 1–1 and too much to 1–0 and 0–1. Mark Dixon and Stuart Coles, in a 1997 paper, added a single parameter — rho — that scales just those four scorelines: 0–0 and 1–1 up, 1–0 and 0–1 down, by amounts that depend on the two expected-goal figures. Everything from 2–0 upwards is untouched.

Their second change was time decay. A result from last season says less about a side than one from last week, so each match is weighted by its age when the strengths are fitted. This site uses a half-life of 180 days: a match that old counts half as much as one played yesterday.

A worked example

Home side expected to score 1.45, away side 1.1, rho of -0.05 — the value this site uses. The four adjusted scorelines, plain Poisson against Dixon–Coles:

The four low scorelines under plain Poisson and after the Dixon–Coles adjustment
ScorelinePlain PoissonAdjustmentDixon–Coles
0–07.81%× 1.0808.43%
1–011.32%× 0.94510.70%
0–18.59%× 0.9287.97%
1–112.45%× 1.05013.08%

Summed over every scoreline, the result market moves from 45.1% / 26.2% / 28.7% to 44.5% / 27.4% / 28.1%: a point onto the draw, taken from both sides. Over 2.5 goals is unchanged at 46.9%, because the adjustment only moves probability among scorelines with two goals or fewer. Both teams to score edges up from 51.1% to 51.7%, carried by the extra weight on 1–1.

Small numbers — and that is the point. The correction is worth about a percentage point on the result, but the draw is priced at around 3.40 and the under-2.5 line at around 1.90, and a point of probability at those prices is the difference between a fair price and a value one. On the correct-score market, where 0–0 and 1–1 are priced at 8 and 6, it is the whole edge.

How this site fits it

Attack and defence strengths are fitted by gradient ascent on the time-weighted likelihood of roughly the last 40 matches per team, with a ridge penalty that shrinks a side with little evidence — newly promoted, early in the season — toward the league average rather than letting three results define it. The fit is done twice, once on goals and once on expected goals, and the strengths are blended with xG weighted at 60%, because xG is the less noisy measure of what a side is doing. Where xG is unavailable the goals-only fit is used and the forecast records that it was.

Every market on a match page — result, totals, both teams to score, Asian handicap, correct score — is read off the one scoreline table, so the page cannot contradict itself, and the probabilities are published at least 24 hours before kick-off and kept on public record. What the model cannot see — injuries, motivation, weather, a manager’s first match — is set out honestly on how it works.

Try it

The Poisson football calculator takes two expected-goal figures and returns the result, totals, both-teams-to-score probabilities and the most likely scores with the same arithmetic. Today’s fixtures show the live output for every match, and model versus market shows where it disagrees most with the bookmakers.

Questions

What is the Dixon–Coles model?
A football prediction model published by Mark Dixon and Stuart Coles in 1997. It treats each side’s goals as a Poisson process driven by its attack strength, the opponent’s defence strength and home advantage, then corrects the plain Poisson model’s known tendency to under-rate 0–0 and 1–1 and over-rate 1–0 and 0–1, and weights recent matches more heavily than old ones.
How is it different from a plain Poisson model?
Two things: a dependence parameter (rho) that adjusts the four lowest scorelines, and a time decay so that form counts. On most matches the two models agree to within a percentage point on the result; the difference shows in the draw, the correct-score market and the under-2.5 line, which is where a plain Poisson model is most often wrong.
What does rho mean?
A small negative number — this site uses -0.05 — that says low-scoring games are slightly likelier to be draws than independence implies. A rho of zero gives the plain Poisson model. It is estimated from results, not chosen.
Does the model use expected goals (xG)?
Yes, where the data provider supplies it. The strengths are fitted twice, once on goals and once on xG, and blended with xG weighted at 60%. Where xG is missing the model falls back to goals alone and records that it did.
Can I run the model myself?
The Poisson calculator on this site takes two expected-goal figures and returns the result, totals, both-teams-to-score and the most likely scores using the same arithmetic as the worked example. The full fit — estimating each team’s strengths from hundreds of matches — is described on the how-it-works page and in the original paper.

Reference: Dixon, M. J. and Coles, S. G. (1997), “Modelling Association Football Scores and Inefficiencies in the Football Betting Market”, Journal of the Royal Statistical Society, Series C, 46(2), 265–280.