A match-outcome model built from a quarter-century of English top-flight data: engineered form features, a leakage-safe temporal split, four competing models and an honest comparison against the bookmakers' own prices. The trained model is embedded below — pick two teams and watch it price the game.
Two engines are loaded. The market engine uses 17 form features plus the bookmakers' margin-free implied probabilities. The no-market engine swaps those three odds columns for Elo ratings, so it can price a fixture that has no odds yet — anything in 2026-27, for instance.
Training stopped at 2022-23. Validation used 2023-24 and 2024-25. The whole of 2025-26 was held back entirely — so here is every match of it, priced by the model, next to the bookmakers' favourite and what actually happened.
| Date | Fixture | Score | Model H / D / A | Market H / D / A | Model | Market |
|---|
Probabilities come from the forest embedded in this page; the pick shown is the draw-aware hybrid.
Six steps, in order. Most of the work was not the modelling.
Every Premier League season from 2000-01 to 2025-26 from football-data.co.uk — one CSV per season, loaded live so nothing goes stale. Two files () are malformed at source and are skipped by the loader rather than silently half-parsed. That leaves clean matches.
The odds columns were renamed twice. For each outcome the loader takes the first column that actually has a
value — Avg* from 2019 onward, then BbAv*, then Bet365's B365*. Odds
coverage starts in 2002-03; earlier seasons predate odds archiving, so they are flagged and excluded from
odds-based models rather than imputed.
Walking through matches in date order: last-5 points, goals for and goals against for each side; season-to-date points per game, goals and win rate; and head-to-head win rate, draw rate and goal average over the last ten meetings. The ordering rule is the whole game — features are computed, then the match is appended to history. Nothing a model sees was knowable only after kick-off.
Bookmakers build a margin into their prices, so the three implied probabilities sum to about , not 1. Dividing each by that sum strips the vig out and leaves three numbers that add to one — the market's actual opinion, which is what the model gets to see.
Train on everything up to 2022-23 ( matches), validate on 2023-24 and 2024-25 (), hold 2025-26 back completely (). A random split would let the model learn from matches that happened after the ones it is scored on. The hybrid's draw threshold is tuned on the training set only, for the same reason.
Logistic regression; a random forest; two Ridge regressions predicting each side's goals; and a hybrid that uses the goal margin to call draws and the forest to pick a winner otherwise. Accuracy alone rewards a model that never predicts a draw — so draw-F1 is reported next to it everywhere.
Football features are a trap. The obvious way to describe a team is "points per game this season" — but if that average is computed over the whole season, it already contains the result of the match being predicted. The model looks sharp and has learned nothing.
Same models, same temporal split, same eight aggregate features. The only difference is whether they are computed over the whole season or over prior games only:
The inflation here is real but modest — of accuracy — because the market odds in the feature set already carry most of the signal, leaving the leak less room to help. That number is still unshippable: it describes a world in which you already know the result.
| Feature leak | Season aggregates that include the match itself. Fixed by building every feature from prior games only, then appending the match to history. |
| Split leak | A random train/test split mixes future matches into training. Fixed with a season-based temporal split. |
| Tuning leak | Choosing the draw threshold by whatever scores best on the test set. Fixed by tuning on training data only — the winning value was t = . |
Scored on 2023-24 and 2024-25 — matches the models never trained on.
| Model | Accuracy | Draw F1 | Macro F1 | Log loss | Read |
|---|
Log loss scores the probabilities rather than just the pick — lower is better. It is the number a pricing desk actually cares about. The logistic regression lands level with the market on it; the forest picks more winners but is the worst-calibrated model on the bench, which is exactly the trade a book has to notice.
The hybrid calls a draw when the two Ridge forecasts land within goals of each other. Widen it and the model catches more draws but gets fewer matches right overall. Here is that trade-off, measured on the validation seasons.
Group every match by the probability assigned to it, then check how often that outcome actually happened. A perfectly calibrated model sits on the dashed diagonal.
Would the model do better if it knew explicitly how strong each team is? An Elo rating was built across all 26 seasons — every team starts at 1500, the winner takes points from the loser, and the swing is larger when the result is a surprise. Then the same two models were retrained on four different feature sets:
| Feature set | Rows | LogReg | Forest |
|---|
Elo is strong on its own — it lifts the no-odds model by about a point. But once market odds are in the feature set it adds nothing, because the odds already price team strength in. So it earns its place only in the no-market engine, which is exactly where this page uses it.
Every team-season since 2000, clustered by playing style with K-means and projected onto two principal components. of the variance is simply "how good are they"; the second component separates open, high-scoring sides from tight, low-scoring ones.
| Tier | PPG | GF | GA | Win % | n |
|---|
Every finishing position has a characteristic profile of wins, draws and defeats. Stack all complete seasons on top of each other and the shape is remarkably stable: wins collapse as you go down the table, defeats climb to meet them, and draws stay almost flat throughout.
The draw line is the flat one. It does bend — a shallow hump peaking around — but position explains only of the variation in how often a club draws. Wins and defeats are the whole story of a league table; draws are close to a constant.
| # | W | D | L | Pts avg | Draw % | Draws, range |
|---|
Two independent Poisson draws under-count the scores where both sides stay low, which is where draws live. Dixon-Coles fixes that with a single parameter that reweights 0-0, 1-0, 0-1 and 1-1. Fitting it by maximum likelihood on the training seasons alone gives ρ = , and here is what it actually bought:
| Model | Log loss | Mean P(draw) | Draws picked |
|---|
| Scoreline | Poisson | Dixon-Coles | Actual |
|---|
An honest, small result. The correction pulls the average draw probability from to against an actual , and improves log loss in the fourth decimal. It does not make the model pick draws, because a draw is almost never the most likely single outcome — the problem was never the low-score cells, it is that P(draw) rarely clears 30%. The scoreline grid in the predictor uses the corrected version.
No fixtures or odds exist for next season yet, so this runs the no-market engine over a full double round-robin: 380 matches, each turned into two Poisson goal forecasts and sampled a thousand times — with a season-level form swing on top, because match-level randomness alone leaves the table far too stable.
Poisson noise alone only shuffles the table a little, because 380 matches average most of it away. Real clubs also get better and worse between seasons — signings, injuries, a new manager. Each simulated season now draws an attack and a defence multiplier per club, so a side can have a genuinely good or bad campaign rather than just a lucky one.
Team strength starts from end-of-2025-26 form and Elo. The three promoted clubs have no recent Premier League record, so they start from the average profile of every promoted club in the dataset. It is a projection of what the model believes today, not a forecast anyone should bet on.
It does not beat the market. The market's own implied probabilities are among its strongest inputs, so the honest claim is that it lands in line with bookmaker prices — close on accuracy, close on log loss — not ahead of them. The useful signal is where model and market disagree, which is a flag to review a price, not a reason to bet against it.
Football is genuinely high-variance. The goals regression explains of the variation in goals scored, and roughly one match in lands on the exact scoreline the model calls most likely. Anyone quoting much better than that on this data is measuring something wrong.
Draws stay hard. Every accuracy-maximising model here learns to almost never predict one. That is a rational response to draws being the minority class — and precisely why draw-F1 sits next to accuracy throughout.
Premier League only. Everything is trained on English top-flight matches. Another league needs a retrain, and in a league without reliable odds the no-market engine becomes the only option.