How Our Football Model Works: 1X2, Over/Under, BTTS, Asian Handicap
Most betting sites tell you a prediction. Almost none tell you how it was made. This post opens the hood on our football model — what data goes in, how the two underlying models differ, where they get things wrong, and how a probability turns into an actual bet on our model page. Nothing here is hype: we will show you real error metrics, and we will tell you honestly about the mistakes we found and fixed in production.
Two models, four market families
Football is not one prediction problem — it is at least two. Predicting who wins (1X2: home / draw / away) is a classification problem. Predicting how many goals get scored (Over/Under, Asian Handicap, BTTS, half-time markets) is a scoring problem. We use a separate model for each, because forcing one architecture to do both jobs tends to do neither well.
Skip the hand-calculation.
Get real value bets flagged for you — 7-day free trial- 1X2 (match result) — an XGBoost classifier trained on ~102,000 matches since 2006.
- Over/Under, Asian Handicap, BTTS, half-time — a Poisson goals model with a Dixon-Coles correction.
Both models are trained across the top European leagues — Premier League, La Liga, Serie A, Bundesliga, Ligue 1, Championship, Eredivisie, Primeira Liga — plus international tournaments. Coverage per competition is on the individual league hubs, for example the Premier League hub.
The XGBoost 1X2 model
The match-result model is a gradient-boosted tree classifier (XGBoost) fed by 110 input features per match. It does not just look at league position — it blends several signal families:
- Elo ratings — updated nightly for every team, with confederation-aware handling for national teams (see Elo ratings explained) and no artificial home-advantage bump at neutral tournament venues.
- Recent form — rolling windows of results, goals, and underlying performance.
- Head-to-head history between the two sides.
- An xG proxy built from shots on target — real Opta/StatsBomb-style expected-goals data is not freely available for every league we cover, so we fit `xG_hat = league coefficient × shots on target`, calibrated per league (R² ≈ 0.46 across leagues; for example the Premier League coefficient is 0.316). See what is xG (expected goals) for the underlying concept.
- News sentiment and weather as smaller supporting signals.
On raw historical data the model's Brier score — a standard measure of probabilistic accuracy where lower is better — is roughly 0.18. After calibration (below) it improves to around 0.177. For context, a naive coin-flip 3-way guess scores much worse; the real benchmark we compare against is the accuracy of Pinnacle's own de-vigged market, which is the sharpest publicly available estimate of true probability (more on that in devigging explained).
Confirmed line-ups move the number
Starting eleven data is checked every 15 minutes before kick-off. Once a line-up is confirmed, a missing forward costs roughly −5.5% of that team's expected goals, a missing midfielder −3.5%, a missing defender −2%, capped at a maximum −20% swing so one absence can't distort the whole prediction — and the 1X2 win probabilities get a proportionally scaled shift to match. If no confirmed line-up is available yet, a weaker injury-list adjustment fills in as a fallback until team news drops.
The Poisson goals engine
Over/Under, Asian Handicap, BTTS and half-time markets are all driven by the same underlying quantity: how many goals each team is expected to score. The model predicts an expected-goals value (lambda) for the home team and one for the away team, then uses the Poisson distribution to turn those two numbers into probabilities for every scoreline, total, and handicap outcome.
A plain Poisson model has a known flaw: it misprices low-scoring games. It systematically overestimates 0-0 draws and underestimates 1-0 / 0-1 / 1-1 results, because in reality the two teams' scores are not perfectly independent — cagey, low-scoring matches produce a small extra correlation that a naive model misses. We correct for this with the Dixon-Coles adjustment, using a correlation parameter (rho) fitted directly on our own match data via maximum likelihood rather than borrowed from old literature. The classic academic value is rho ≈ −0.13, taken from 1990s English football data — on our own dataset that turned out to be too strong; our fitted value is rho ≈ −0.06.
Accuracy on held-out data: mean absolute error of roughly 0.94 goals for the home team and 0.85 goals for the away team, and Over/Under 2.5 accuracy of 0.579 against a baseline of 0.549 for always picking the historical base rate. That is a real, if modest, edge — goals are inherently noisy, and no model closes that gap entirely.
Line choice and Asian quarter-balls
For Over/Under markets the engine does not automatically bet whichever line has the highest expected value — it bets the line closest to the model's expected total (home lambda + away lambda). Extreme lines, far from the expected total, are exactly where a goals model's calibration is weakest, so anchoring to the central estimate avoids chasing noisy edges on the tails. Over/Under totals explained covers the market mechanics in more depth.
Asian Handicap and quarter-ball lines (e.g. Over 2.75, handicap +0.25) need careful push-probability normalisation — part of the stake can be refunded rather than won or lost outright, and the EV formula has to account for that correctly. See Asian handicap explained for a worked example of how quarter lines split a bet across two half-lines.
Calibration: making probabilities honest
A model's raw output is rarely perfectly calibrated — when it says "40%", the true long-run frequency might actually be 35% or 45%. We correct this with isotonic calibration, a technique that learns a monotonic mapping from raw model probabilities to observed outcome frequencies.
One honest detail worth sharing: a single shared calibration mapping does not work well for football, because club football and international football behave differently. In our own data, club matches' raw draw probability needed to be shrunk toward reality, while international matches' raw draw probability was already close to correct — pushing both through one calibrator over-corrected international draws almost to zero, which was wrong. The fix was to split calibration by domain: the isotonic calibrator is fitted only on club matches and applied only to club matches; international (national-team) matches keep the model's raw, uncalibrated output. We found this in production, checked it against the market, and fixed it rather than papering over it with a display-level filter.
From probability to bet
A calibrated probability is not a bet by itself. It only becomes one when it is measured against a price. Our anchor is Pinnacle's de-vigged closing line — the sharpest publicly available estimate of true probability (see devigging explained for exactly how that stripping-the-margin math works). We compare our model's probability against what a book is actually offering and compute expected value:
Staking uses fractional Kelly at 25%, capped at 5% of bankroll on any single bet — a deliberately conservative fraction of the theoretical full-Kelly stake, because model probabilities carry uncertainty that full Kelly assumes away.
How we know it works — and where it doesn't
We track two different things, because they answer two different questions. Closing line value (CLV) — whether our bet consistently gets a better price than the market ends up settling at — is the reliable long-run signal for moneyline-type markets like 1X2. For Over/Under and Asian Handicap, CLV alone can be misleading (it is possible to have positive CLV and still lose money on a specific market shape), so we lean on realised ROI and calibration accuracy as the leading metrics there instead.
We publish every settled bet on our track record — wins, losses, stakes, and CLV, not a curated highlight reel. Small samples are noisy: a hundred bets tell you very little, five hundred start to be statistically meaningful. Yield in the 3–8% range over a large sample is realistic for a disciplined value-betting approach; anything dramatically higher over a short stretch is more likely variance (or overfitting) than a durable edge, and we say so even when the number would look better if we didn't.
The model is not, and will never be, a crystal ball. Draws are hard. Low-scoring games are noisy. International matches have thinner data than a title-winning Premier League team. What we can promise is a transparent, honestly validated process — not a guaranteed profit.